About
I started in computer vision and medical imaging — uncertainty quantification, model fairness, and segmentation with teams at IIT Kharagpur and Bosch Research. Moving into language models, I kept running into the same problem: a model that looked aligned on benchmarks would fail in ways that weren’t visible from the outside. That gap — between what training instils and what survives deployment — became the question I keep returning to.
On the safety side, I study how fine-tuning and quantization silently erode alignment, and how to find and repair the specific weights responsible using mechanistic interpretability: circuit-level refusal (C-ΔΘ), quantization-permanent unlearning (Forgetting That Sticks), and safety drift auditing. On the structured data side, I work on tabular foundation models — how to train, fine-tune, and distil them down to something deployable (TabTune, Orion-MSP, Orion-BiX). Most of this ships as open-source tools. The next problems I care most about are safety in reasoning models and interpretability for agentic systems.
I lead the Model Science group at Lexsi Labs across India and Paris. I completed my B.Tech in Data Science from MIT Manipal in 2024, and before that worked at Mila Quebec AI Institute (with Prof. David Rolnick), Bosch Research India, and IIT Kharagpur. I’m an AAAI Undergraduate Consortium Scholar (2023 → Mentor, 2026).
Recent work: AlignTune · CircuitKIT · C-ΔΘ · CuratorKIT · SafeTune · Forgetting That Sticks · TabTune · Orion-MSP
Feel free to reach out or see my Resume.
Key Highlights
Published at ICML, ACL, WWW, MIDL, IJCNN, AAAI, and Nature Scientific Reports, plus workshops and shared tasks at CVPR, NeurIPS, ICLR, MICCAI, EurIPS, ACL, and SIGMOD.
- Research Focus: LLM safety post-training — circuit-level mechanistic interpretability, safety weight arithmetic, unlearning, post-training alignment, and a growing focus on AI agents & evaluation.
- 7 open-source libraries released — see Libraries & Toolkits.
- AAAI Undergraduate Consortium Scholar & Mentor (2023 → 2026); Spotlight Talk at EurIPS Workshop on Private AI Governance, Copenhagen (December 2025).
Lexsi Labs — Internships & Full-Time Roles: For internship and FTE applications at Lexsi Labs, please apply directly via lexsi.ai rather than reaching out for referrals.
Mentoring: I am open to mentoring early-stage and young researchers. If you’d like to connect, feel free to reach out via email — please be respectful of my time and include a brief note about your background and what you’re working on.
News
Recent Publications & Acceptances
- 2026.07New Pre-Print: CircuitKIT: Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability
- 2026.07New Pre-Print: Faithfulness to Refusal: A Causal Audit of Neuron Selectors in LLMs
- 2026.06New Pre-Print: CuratorKIT: Data Curation and Synthetic Data Generation for LLM Post-Training
- 2026.06New Pre-Print: ALIGNBEAM: Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing, accepted at AI for Good Workshop @ ICML 2026
- 2026.06New library release: SafeTune — a unified library for auditing and repairing safety drift in fine-tuned LLMs
- 2026.05Data Presentation over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models accepted as an Oral at FinDS Workshop @ ACM SIGMOD 2026
Earlier publications & acceptances
- 2026.05Distilling Tabular Foundation Models for Structured Health Data wins Best Paper Runner-Up (Spotlight) at SD4H Workshop @ ICML 2026
- 2026.05New Pre-Print: Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands
- 2026.05Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees accepted at FMSD Workshop @ ICML 2026
- 2026.05Ensembling Tabular Foundation Models: A Diversity Ceiling and a Calibration Trap accepted at FMSD Workshop @ ICML 2026
- 2026.02New Pre-Print: AlignTune: Modular Toolkit for Post-Training Alignment of Large Language Models
- 2026.02New Pre-Print: C-ΔΘ: Circuit-Restricted Weight Arithmetic for Selective Refusal
- 2026.01Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning accepted at WWW 2026
- 2026.01Exploring Fine-Tuning for Tabular Foundation Models accepted at WWW 2026
- 2026.01TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models (Demo) accepted at WWW 2026
- 2026.01Laplacian reconstructive network for guided thermal super-resolution accepted at Scientific Reports (Nature)
- 2025.11Interpretability as Alignment: Making Internal Understanding a Design Principle accepted at EurIPS Workshop on Private AI Governance
- 2025.11Bridging the gap in XAI-why reliable metrics matter for explainability and compliance accepted at EurIPS Workshop on Private AI Governance
- 2025.11EurIPS Workshop on Private AI Governance 2025 Spotlight Talk
- 2025.09Interpretability-aware pruning for efficient medical image analysis accepted at MICCAI Workshop 2025
- 2025.05SELF-PERCEPT: Mental Manipulation Detection accepted at ACL 2025
- 2025.05Alberta Wells Dataset accepted at ICML 2025 (grateful to the team for their efforts, and to Prof. David Rolnick)
Academic Service & Reviewing
- 2026.07Reviewer for Actionable Interpretability Workshop @ COLM 2026
- 2026.07Program Committee for AAAI 2027
- 2026.06Reviewer for ACM AIES 2026
- 2026.06Reviewer for System Demonstrations Track @ EMNLP 2026
- 2026.06Reviewer for NLP4PI Workshop @ EMNLP 2026
- 2026.06Reviewer for BlackboxNLP Workshop @ EMNLP 2026
Earlier service & reviewing
- 2026.05Reviewer for AI for Good Workshop @ ICML 2026
- 2026.05Reviewer for Mechanistic Interpretability Workshop @ ICML 2026
- 2026.05Reviewer for TAIGR Workshop @ ICML 2026
- 2026.05Reviewer for FMSD Workshop @ ICML 2026
- 2026.05Reviewer for FAIMI-BRIDGE-EPIMI Workshop @ MICCAI 2026
- 2026.05Reviewer for WACV 2027
- 2026.04Reviewer for NeurIPS 2026
- 2026.04Reviewer for BMVC 2026
- 2026.03Reviewer for ECCV 2026
- 2026.03Reviewer for FinDS Workshop @ ACM SIGMOD 2026
- 2026.02Reviewer for Advances in Financial AI Workshop (ICLR 2026)
- 2026.01Reviewer for CVPR 2026
- 2025.12Mentor at AAAI Undergraduate Consortium 2026