← Back to Website
📄 Download PDF

Himanshu Pal

Machine Learning Engineer-II

Professional Summary

Machine Learning Engineer with 5+ years across ML systems, distributed model training, and deep learning research. Builds and troubleshoots multi-GPU / multi-node training pipelines (PyTorch, SLURM) with reproducible tracking, checkpointing, and failure recovery. Delivered production ML at enterprise scale at ServiceNow, including an agentic ticket-resolution system with a neurosymbolic verification layer. Published at AAAI 2026, TMLR 2025, and KDD/ASONAM 2025.

Education

MS (Research), IIIT Hyderabad Jul 2023 – Dec 2025
Thesis: Adversarial and Privacy-preserving methods in Geometric Deep Learning
B.Tech, IIT Roorkee Jul 2016 – May 2020

Publications

AAAI 2026 LoReTTA: A Low Resource Framework to Poison Continuous Time Dynamic Graphs · Himanshu Pal*, Venkata Sai Pranav Bachina*, Ankit Gangwal, Charu Sharma
Low-resource adversarial poisoning framework for continuous-time dynamic graphs, exposing temporal-graph vulnerability under a minimal perturbation budget.
TMLR 2025 Federated Spectral Graph Transformers Meet Neural Ordinary Differential Equations for Non-IID Graphs · Kishan Gurumurthy*, Himanshu Pal*, Charu Sharma
Combined spectral graph transformers with neural ODEs for federated learning over non-IID graph distributions.
KDD(W)/ASONAM 2025 Predict Confidently, Predict Right: Abstention in Dynamic Graph Learning · Jayadratha Gayen, Himanshu Pal, Naresh Manwani, Charu Sharma
Selective prediction (abstention) for dynamic graph models, letting them decline uncertain predictions instead of forcing low-confidence outputs.

Professional Experience

ServiceNow Oct 2025 – Present
Machine Learning Engineer II
  • Built a production churn prediction system processing 25M daily customer action events, designed and delivered three pipeline stages (feature selection, training with hyperparameter tuning, prediction with SHAP explainability) across three internal services (Glide scheduler, Nagini training pipeline, Java prediction service); achieved 80% AUC on held-out evaluation.
  • Engineered parallel execution of SHAP-based explainability alongside batch prediction to meet latency requirements at enterprise scale.
  • Benchmarked tabular foundation models (TabPFN, TabDPT) and relational foundation models against tuned gradient boosting baselines; found foundation models achieved comparable accuracy without task-specific training.
  • Developing a neurosymbolic verification layer for an agentic ticket resolution system (Zero Touch Service Desk), the agent performs actions such as password resets and account modifications, and the verification layer checks privilege constraints before execution; improved end-to-end task closure from 0% to 33%.
Barclays PLC, Mumbai Feb 2021 – Jul 2023
Quantitative Analyst — Automated Pricing & Trade Analytics
  • Built an automated PnL reconciliation pipeline in Python and SQL that replaced a manual Excel-based workflow, reducing daily report generation from 10 minutes to 2 minutes and eliminating manual data-entry errors.
  • Designed an event-driven pricing service triggered by inbound email RFQs (requests for quote), generating structured option quotes in under 30 seconds versus a previous manual turnaround of several hours.
  • Integrated and maintained pricing and market-data APIs to support quote generation and desk-level analytics for Interest Rate Options products.
  • Worked on the Rates Options Structuring desk covering interest rate derivatives (swaptions, caps/floors, exotics), requiring working knowledge of stochastic pricing models, Greeks, and risk sensitivities.

Research Experience

IIIT Hyderabad Jul 2023 – Dec 2025
Graduate Researcher · Advisor: Dr. Charu Sharma | PyTorch, CUDA, HuggingFace, SLURM, W&B
  • Built and optimized distributed training pipelines for graph and language models across multi-GPU / multi-node SLURM clusters, with reproducible experiment tracking (W&B) and checkpointing for long-running jobs.
  • Diagnosed and resolved distributed training failures and implemented checkpoint/failure-recovery workflows to support publication-grade results in graph learning.

Technical Skills

Distributed Training: Multi-GPU / multi-node training, SLURM, PyTorch, CUDA, checkpointing, failure recovery
LLM / Fine-tuning: HuggingFace, Transformers; LoRA / QLoRA / PEFT, RAG (academic project)
MLOps / Infra: Docker, AWS (EC2/S3/Lambda), W&B, experiment tracking
Languages: Python, C++, SQL, Bash