BS Data Science @ IIT Madras · AI/ML Research Intern @ IIT Mandi
I'm a Data Science undergraduate at IIT Madras interested in understanding how machine-learning systems behave — especially when they encounter noisy data, distribution shifts, unreliable retrieved context, or adversarial inputs.
My work currently spans:
- LLM & agent evaluation — instruction–data separation, tool-use reliability, critique agents, and adversarial evaluation
- Temporal ML — sequence modeling with TCNs, LSTMs, and MS-TCNs
- Retrieval systems — RAG, embeddings, vector databases, and evidence-grounded generation
- Reproducible experimentation — controlled evaluations, failure analysis, structured logging, and model comparison
- Open models — experimenting with open-weight models and transparent evaluation pipelines
I like building experiments that make model failures easier to reproduce, measure, and understand.
|
LangGraph · Python · LLM Evaluation A controlled evaluation environment for studying instruction–data separation failures in tool-using agents.
Focus: agent reliability · prompt injection · tool-use evaluation |
LangGraph · Retrieval · LLM Evaluation A generator–critic system for testing whether an independent critique stage can reduce unsupported LLM outputs.
Focus: hallucination evaluation · grounding · model reliability |
|
Python · FastAPI · pgvector · LiteLLM A retrieval-backed platform for structured analysis of software repositories.
Focus: retrieval systems · repository analysis · reliable generation |
PyTorch · TCN · LSTM · MS-TCN Research internship focused on temporal modeling of real-world human-manipulation demonstrations.
Focus: temporal ML · sequence modeling · generalization |
┌──────────────────────┐
│ Model Behaviour │
└──────────┬───────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Agent Reliability Retrieval Temporal ML
│ │ │
▼ ▼ ▼
Adversarial Eval Grounding OOD Evaluation
└─────────────┬─────────────┘
▼
Reproducible Experiments
I'm particularly interested in questions like:
- When do tool-using agents treat untrusted data as instructions?
- Can critique or verification stages reliably reduce unsupported generation?
- How should model failures be measured beyond aggregate accuracy?
- How do sequence models behave when demonstrations differ from their training distribution?
- How can evaluations be made easier for other researchers to inspect and reproduce?
Mar 2026 – May 2026
Worked on temporal modeling of human-manipulation demonstrations using TCN, LSTM, and MS-TCN architectures.
The experimental pipeline covered:
video data → preprocessing → kinematic features → temporal models → evaluation → error analysis
The dataset contained approximately 9,000 real-world videos across six action classes, including evaluation on out-of-distribution demonstrations.
What I worked on
- Built reproducible PyTorch training and evaluation pipelines
- Prepared noisy temporal data for sequence-model experiments
- Extracted kinematic features from demonstrations
- Ran model and hyperparameter comparisons
- Evaluated TCN, LSTM, and MS-TCN architectures
- Analyzed model errors on continuous temporal predictions
- Tested generalization on out-of-distribution physical demonstrations
- Automated parts of the data-ingestion and preprocessing workflow
flowchart LR
A[Question] --> B[Controlled Experiment]
B --> C[Structured Evaluation]
C --> D[Failure Analysis]
D --> E[Iterate]
E --> B
D --> F[Document Results]
I prefer experiments where the evaluation setup, assumptions, metrics, and failure cases are visible, rather than treating a single benchmark number as the entire result.
Other Projects & Earlier Work
My earlier projects span recommendation systems, retrieval applications, database-backed systems, and applied ML.
These projects helped me build experience with:
- End-to-end ML pipelines
- Embedding-based retrieval
- API and database integration
- Recommendation systems
- Dockerized applications
- Git/GitHub development workflows
You can explore the rest of my repositories from my GitHub profile.
I'm currently spending most of my project time on:
LLM evaluation → adversarial behavior, critique systems, grounding
Agent reliability → instruction–data separation and tool-use failures
Open models → experiments that can be inspected and reproduced
Retrieval → evidence-grounded generation and failure-aware pipelines
Temporal ML → sequence modeling and generalization
I'm interested in conversations around ML research, open models, evaluation, agent reliability, retrieval systems, and reproducible experimentation.