Ayush Singh

I am a final-year undergraduate at Indian Institute of Technology Roorkee, pursuing a B.Tech in Materials Engineering. My research interests primarily lie in the domain of Large Language Models, RLHF/Preference Optimization, AI Safety & Red-Teaming, and Multi-Agent Systems.

I have had the pleasure of working as a Research Intern at Adobe MDSR, where I built agentic pipelines for Core Web Vitals optimization, and as an AI Engineer Intern at Prem AI, where I built a multilingual instruction-following benchmark across all 24 EU languages. I've also worked at Repello AI on AI security and red-teaming, at Boston University on LLM alignment research, and at Lossfunk collaborating with Paras Chopra on novel RLHF techniques.

I am a member of the Vision and Language Group (VLG) at IIT Roorkee, where I contribute to research in deep learning and multi-modal AI.

Email  /  Google Scholar  /  LinkedIn  /  GitHub

profile photo
Experience
Adobe · Research Intern, MDSR
May - Jul 2025 & 2026

Engineered a production agentic pipeline (Playwright + LangGraph/CrewAI) for end-to-end Core Web Vitals optimization: issue detection, patch generation, and PR creation. Built SWE-WEB, the first benchmark for CWV optimization by AI coding agents (10,689 repos, 19 agent-model configs), and trained Gemma4 as a CWV planner via SFT + RLVR on mined GitHub PRs.

Prem AI · AI Engineer Intern
Mar 2026 - May 2026

Built Euro-IFBench, extending IFBench's verifiable instruction-following rules with localized prompts and rewritten verifiers across all 24 EU languages. Benchmarked 13 instruction-tuned models, exposing a persistent English-vs-non-English gap and a strong negative correlation between tokenizer fertility and instruction-following accuracy.

Repello AI · AI Security Engineer Intern
Aug 2025 - Dec 2025

Fine-tuned RoBERTa into a multilingual guardrail spanning 15+ languages, surpassing baseline safety classifiers on adversarial-prompt detection. Red-teamed production LLM agents with adaptive, multi-turn attack strategies, and authored 4 technical blogs on model scanning and diffusion/distilled-model safety.

Boston University · Research Intern
Feb 2025 - Apr 2025

Designed a controlled study evaluating self-awareness, generalization, and reasoning-output alignment across SFT-, DPO-, and GRPO-tuned LLMs on 5 behavioral tasks, manually verifying ~4,800 think-answer pairs. Co-authored the paper, accepted to ACL 2026 Findings.

Lossfunk · Research Intern
Dec 2024 - Feb 2025

Designed Implicit Preference Optimization (IPO), a logit-based self-rewarding RLHF method that eliminates the need for a separately trained reward model. Co-authored the resulting paper under Paras Chopra, reaching up to 78% accuracy on RewardBench across 5 model families. Accepted to ACL 2025 Main Conference.

Publications

My research focuses on Large Language Models, RLHF techniques, and efficient fine-tuning methods. Papers with highlighted backgrounds are selected/notable works.

CATPO: Critique-Augmented Tree Policy Optimization
Ayush Singh
arXiv Preprint, 2026
arXiv

Introduces a tree-informativeness score combining leaf-outcome diversity with policy-reward decorrelation to filter uninformative rollout trees in tree-based RLVR at zero extra compute, reaching 37.5% macro accuracy on Qwen2.5-Math-1.5B across four math benchmarks.

Thinking About Thinking: Evaluating Reasoning in Post-Trained LLMs
Ayush Singh
ACL 2026 Findings
arXiv

Empirical evaluation of LLM self-awareness across Base, SFT, DPO, and GRPO training methods on 5 behavioral tasks. Finds GRPO lifts bias-induction self-awareness from 3% to 50.5% accuracy but shows the weakest reasoning-output alignment on OOD prompts.

IPO: Your Language Model is Secretly a Preference Classifier
Ayush Singh, Paras Chopra
ACL 2025 Main Conference
arXiv

Novel RLHF alternative using LLMs as preference classifiers, reducing dependence on external reward models. Demonstrates self-improving capabilities on RewardBench across Mistral-7B and Llama-1B models with significant improvements.

LoRA-Mini: Adaptation Matrices Decomposition and Selective Training
Ayush Singh
AAAI 2025 CoLoRAI Workshop
arXiv

Efficient parameter reduction technique achieving 20x fewer trainable parameters than vanilla LoRA. Maintains competitive accuracy on GLUE and WMT16 benchmarks across RoBERTa, BERT, and T5 models.

Adaptive Urban Planning: A Hybrid Framework for Balanced City Development
Ayush Singh
AAAI 2025 AI4UP Workshop
arXiv

Multi-agent LLM framework for urban planning optimization integrated with genetic algorithms. Achieves substantial improvements in planning efficiency and balanced city development through collaborative coordination.

Selected Projects
RAVEN: Hypothesis Validation Copilot
GitHub

A retry-aware LangGraph pipeline across planning, financial data collection (yfinance), and hybrid LLM/Python-REPL analytics stages with human-in-the-loop review gates. Generates audit-ready PDF reports with cited evidence and visualizations, served through a FastAPI backend.

StructuralDesignEnv: RL Environment for LLM Structural Engineers
GitHub

An OpenEnv RL environment where LLM agents design steel building frames, scored by a 3D 6-DOF direct-stiffness solver and Eurocode 3 checks. Three graded tasks from a single-story warehouse to a seismic hospital, served via a Dockerized FastAPI server.

RealPDE: Long-Term Test-Time Adaptation (NeurIPS 2026 Competition)
GitHub

Adapting neural-operator forecasters (CNO, FNO, Transolver) for streaming real-world PIV airfoil-wake flow. A bounded adaptation controller (skip / recalibrate / anchored norm-layer update) with uncertainty bounds and LLM-driven action selection under a fail-safe rule fallback.

JED Red-Team: Multi-Step Tool Attacks on LLM Agents (Kaggle)
GitHub

Reverse-engineered a dataflow guardrail's predicate logic (exfiltration, destructive-write, confused-deputy, untrusted-to-action) to expose exploitable scoring gaps, and built a search harness replaying multi-step tool-call chains against guardrailed GPT-OSS/Gemma agents.

Are VLMs Really Blind?
GitHub / arXiv

Benchmarked Gemini, PaliGemma, and GPT-4o on 8 low-level geometric reasoning tasks where VLMs underperform despite strong OCR/VQA results. Proposed a keyword-guided captioning pipeline boosting zero-shot accuracy by up to 14pts without any fine-tuning.

Education
Indian Institute of Technology Roorkee
Bachelor of Technology in Materials Engineering
Aug 2023 - May 2027

Activities:

  • Head of Research | Vision and Language Group (VLG)
  • 1st Place - HiLabs Hackathon (50+ teams), 2024
  • 1st Place - Composio IITD Hackathon, national winner (100+ teams), 2025



Template adapted from Jon Barron.