This is Swadesh Swain

PhD Student @ UMD ECE | IIT Roorkee Alum

profile_picture.png

PhD Student

Electrical and Computer Engineering

University of Maryland, College Park

Hey there! I’m a PhD student in Electrical and Computer Engineering at the University of Maryland, College Park. I did my B.Tech in Electronics and Communication Engineering at IIT Roorkee, where my thesis won the ECE department’s Best B.Tech Project Award. I spend most of my time thinking about how to make AI systems both powerful and safe. My research interests lie at the intersection of Mechanistic Interpretability, AI Safety, and Adversarial Robustness of Vision-Language Models and LLMs.

My first foray into interpretability research led to a paper on improving theoretical guarantees of Integrated Gradients attribution methods, accepted at the NeurIPS 2024 Interpretable AI Workshop. This was followed by an extensive study on adversarial vulnerabilities in VLMs – our reproducibility and enhancement work on Cross-Prompt Attacks was published in TMLR and received the Best Paper Award at MLRC 2025 (presented at Princeton). We further extended this into CroPA++, accepted at the NeurIPS 2025 Reliable ML Workshop, introducing three-fold enhancements that made attacks transferable across images and models.

Currently, I’m working with Dr. Sanghamitra Dutta at UMD on safety-critical features that interpretability tools usually overlook because they are inactive – suppressed features whose silencing can turn refusals into compliance. Our latest work introduces the Counterfactual Activation Potential (CAP) to discover these features at scale, and suggests that jailbreaks may work partly by suppressing them rather than only by activating harmful ones (under review at ICLR 2027). I’ve also studied refusal geometry in VLMs, showing that their safety is encoded in a high-dimensional subspace yet gated by a single direction (under review at TMLR). Previously, I collaborated with Dr. Koustuv Sinha at META AI (FAIR) on benchmarking world model understanding and anticipation mechanisms in Video Language Models, and with Dr. Nagender Aneja at Virginia Tech on user-intervenable LLM pipelines that apply circuit-tracing interpretability methods for real-time model steering.

On the generative AI front, I developed RIGS – a lightweight Riemannian-guided diffusion framework for synthetic signal data generation in collaboration with BOSCH India (accepted at CVIP 2026). During my internship at AuraML, I built text-to-3D scene generation frameworks using Graph Diffusion Models for industrial simulation, contributing directly to their product AuraSim. I’ve also worked on multi-task RL with Diffusion Models at IIIT Hyderabad’s Robotics Research Centre, and on medical AI pipelines at IIT Bombay’s Koita Centre for Digital Health as part of the BharatGen consortium.

Beyond research, I led the Data Science Group at IIT Roorkee as Joint Secretary, heading the research division. Under my tenure, our members published 15+ papers at venues like NeurIPS, CVPR, and ICLR – most led solely by undergraduate teams. I also serve as a reviewer for TMLR.

When I’m not debugging code or reading papers, you’ll probably find me at campus chai spots discussing the latest ML papers, or exploring College Park’s food scene. Feel free to reach out if you want to chat about research, collaborate on projects, or just grab a cup of chai!

news

Oct 05, 2026 Our paper “Riemannian-Guided Diffusion for Scalable Synthetic Signal Data Generation” has been accepted at the International Conference on Computer Vision and Image Processing (CVIP 2026)! RIGS, developed in collaboration with BOSCH India, is a lightweight diffusion framework that generates synthetic vibration signals for industrial fault detection from scarce data.
Sep 28, 2026
Excited to share our new paper “Don’t Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential” – the revised version of our CAP paper, showing that safety-critical features can hide among a model’s inactive components, invisible to activation-focused interpretability, and that jailbreaks may work in part by suppressing them. Submitted to ICLR 2027, currently under review. arXiv OpenReview
Aug 23, 2026 My B.Tech thesis, “From Riemannian Diffusion to Refusal Geometry in Vision Language Models: Studies in Generative Data and AI Safety,” won the Best B.Tech Project Award in the ECE department at IIT Roorkee! Thesis
Apr 19, 2026 Excited to share our new paper “The Encoding-Behavior Dissociation: How Distributed Safety Representations Yield Single-Direction Vulnerabilities in Vision-Language Models” – uncovering how VLM safety is encoded in high-dimensional representations yet gated by a single direction, exposing a limitation of current alignment. Submitted to TMLR 2026, currently under review. PDF
Mar 31, 2026 Our new paper “CAP: Counterfactual Activation Potential for Quantifying Suppressed Safety Features in Language Models” has been submitted to COLM 2026!
Mar 20, 2026 Received offers of admission for NYU Masters in CSE from both Courant and Tandon schools, each with a scholarship of $5000 per year!
Feb 27, 2026 Received offer of admission into the ECE PhD program at the University of Maryland, College Park!
Jan 19, 2026 Our new paper “Riemannian-Guided Diffusion for Scalable Synthetic Signal Data Generation” has been submitted to IJCAI 2026!
Jan 09, 2026 Received offer of admission in Northeastern University’s MSc programs of AI and CS, with 2 merit awards for scholarships!
Dec 02, 2025 I am here at NeurIPS 2025, meet up to talk all things AI Safety or just to hang about!
Nov 01, 2025 Excited to finally start my collaboration with Prof. Sanghamitra Dutta from the University of Maryland, College Park! Our work will focus on investigating reasoning mechanisms inside LLMs which can aid in guardrailing against jailbreaks.
Sep 30, 2025 My paper “CroPA++: Exposing Vulnerabilities in Vision Language Models and Enhancing Adversarial Transferability of Cross-Prompt Attacks” has been accepted at the NeurIPS Reliable ML Workshop, 2025!
Sep 15, 2025 Excited to start our research with Dr. Koustuv Sinha from META AI (FAIR), on evaluating world model understanding of VideoLMs!
Sep 09, 2025 Fortunate to be accepted by Professor Nagendra Aneja at Virginia Tech to pursue applied interpretability research under his guidance. Our work involves designing user-intervenable reasoning agents for on-the-fly steering of LLMs.
Aug 21, 2025 Attending MLRC at Princeton University! Super excited to present our first oral presentation!
Aug 11, 2025 Our work “Revisiting CroPA: A Reproducibility Study and Enhancements for Cross-Prompt Adversarial Transferability in Vision-Language Models” is the recipient of Best Paper at the Machine Learning Reproducibility Challenge at Princeton University! Catch the tweet here.
Jun 27, 2025 Revisiting CroPA is further accepted at the Machine Learning Reproducibility Challenge.
Jun 16, 2025 Our work “Revisiting CroPA: A Reproducibility Study and Enhancements for Cross-Prompt Adversarial Transferability in Vision-Language Models” has been accepted into TMLR journal!
Jun 02, 2025 I will be joining AuraML as a Research Intern to work on cutting edge Generative 3D Vision for industrial simulation applications.
Apr 15, 2025 I will be joining Robotics Research Centre, IIIT Hyderabad as undergraduate research intern for the summer.
Dec 10, 2024 Here at NeurIPS to attend my first ever in-person conference!
Oct 10, 2024 My debut paper “Riemann Sum Optimization for Accurate Integrated Gradients Computation” has been accepted to the NeurIPS 2024, Interpretable AI Workshop!
Apr 04, 2024 I will be joining the BharatGen Team at Indian Institute of Technology (IIT) Bombay as Machine Learning Research intern this summer.

selected publications

  1. Revisiting CroPA: A Reproducibility Study and Enhancements for Cross-Prompt Adversarial Transferability in Vision-Language Models
    Atharv Mittal*, Agam Pandey*, Swadesh Swain*, and 2 more authors
    Transactions on Machine Learning Research, Jun 2025