|
Ayush Singh
I am a final-year undergraduate at Indian Institute of
Technology Roorkee, pursuing a B.Tech in Materials Engineering.
My research interests primarily lie in the domain of Large Language Models,
RLHF/Preference Optimization, AI Safety & Red-Teaming, and
Multi-Agent Systems.
I have had the pleasure of working as a Research Intern at Adobe MDSR, where I built agentic pipelines for Core Web
Vitals optimization, and as an AI Engineer Intern at Prem AI,
where I built a multilingual instruction-following benchmark across all 24 EU languages.
I've also worked at Repello AI on AI security and red-teaming,
at Boston University on LLM alignment research,
and at Lossfunk collaborating with Paras Chopra on novel RLHF techniques.
I am a member of the Vision and Language Group (VLG) at IIT
Roorkee, where I contribute to research in deep learning and multi-modal AI.
Email  / 
Google
Scholar  / 
LinkedIn  / 
GitHub
|
|
|
|
Adobe · Research Intern, MDSR
May - Jul 2025 & 2026
Engineered a production agentic pipeline (Playwright + LangGraph/CrewAI) for end-to-end Core Web
Vitals optimization: issue detection, patch generation, and PR creation. Built SWE-WEB, the first
benchmark for CWV optimization by AI coding agents (10,689 repos, 19 agent-model configs), and
trained Gemma4 as a CWV planner via SFT + RLVR on mined GitHub PRs.
|
|
|
Prem AI · AI Engineer Intern
Mar 2026 - May 2026
Built Euro-IFBench, extending IFBench's verifiable instruction-following rules with localized
prompts and rewritten verifiers across all 24 EU languages. Benchmarked 13 instruction-tuned
models, exposing a persistent English-vs-non-English gap and a strong negative correlation
between tokenizer fertility and instruction-following accuracy.
|
|
|
Repello AI · AI Security Engineer
Intern
Aug 2025 - Dec 2025
Fine-tuned RoBERTa into a multilingual guardrail spanning 15+ languages, surpassing baseline
safety classifiers on adversarial-prompt detection. Red-teamed production LLM agents with
adaptive, multi-turn attack strategies, and authored 4 technical blogs on model scanning and
diffusion/distilled-model safety.
|
|
|
Boston University · Research
Intern
Feb 2025 - Apr 2025
Designed a controlled study evaluating self-awareness, generalization, and reasoning-output
alignment across SFT-, DPO-, and GRPO-tuned LLMs on 5 behavioral tasks, manually verifying ~4,800
think-answer pairs. Co-authored the paper, accepted to ACL 2026 Findings.
|
|
|
Lossfunk · Research Intern
Dec 2024 - Feb 2025
Designed Implicit Preference Optimization (IPO), a logit-based self-rewarding RLHF method that
eliminates the need for a separately trained reward model. Co-authored the resulting paper under
Paras Chopra, reaching up to 78% accuracy on RewardBench across 5 model families.
Accepted to ACL 2025 Main Conference.
|
|
Publications
My research focuses on Large Language Models, RLHF techniques, and efficient fine-tuning methods.
Papers with highlighted backgrounds are selected/notable works.
|
|
|
CATPO: Critique-Augmented Tree Policy Optimization
Ayush Singh
arXiv Preprint, 2026
arXiv
Introduces a tree-informativeness score combining leaf-outcome diversity
with policy-reward decorrelation to filter uninformative rollout trees in tree-based RLVR at zero
extra compute, reaching 37.5% macro accuracy on Qwen2.5-Math-1.5B across four math benchmarks.
|
|
|
Thinking About Thinking: Evaluating Reasoning in Post-Trained LLMs
Ayush Singh
ACL 2026 Findings
arXiv
Empirical evaluation of LLM self-awareness across Base, SFT, DPO, and
GRPO training methods on 5 behavioral tasks. Finds GRPO lifts bias-induction self-awareness from
3% to 50.5% accuracy but shows the weakest reasoning-output alignment on OOD prompts.
|
|
|
IPO: Your Language Model is Secretly a Preference Classifier
Ayush Singh, Paras Chopra
ACL 2025 Main Conference
arXiv
Novel RLHF alternative using LLMs as preference classifiers, reducing
dependence on external reward models. Demonstrates self-improving capabilities on RewardBench across
Mistral-7B and Llama-1B models with significant improvements.
|
|
|
LoRA-Mini: Adaptation Matrices Decomposition and Selective Training
Ayush Singh
AAAI 2025 CoLoRAI Workshop
arXiv
Efficient parameter reduction technique achieving 20x fewer trainable
parameters than vanilla LoRA. Maintains competitive accuracy on GLUE and WMT16 benchmarks across
RoBERTa, BERT, and T5 models.
|
|
|
Adaptive Urban Planning: A Hybrid Framework for Balanced City Development
Ayush Singh
AAAI 2025 AI4UP Workshop
arXiv
Multi-agent LLM framework for urban planning optimization integrated with
genetic algorithms. Achieves substantial improvements in planning efficiency and balanced city
development through collaborative coordination.
|
|
|
RAVEN: Hypothesis Validation Copilot
GitHub
A retry-aware LangGraph pipeline across planning, financial data
collection (yfinance), and hybrid LLM/Python-REPL analytics stages with human-in-the-loop review
gates. Generates audit-ready PDF reports with cited evidence and visualizations, served through a
FastAPI backend.
|
|
|
StructuralDesignEnv: RL Environment for LLM Structural Engineers
GitHub
An OpenEnv RL environment where LLM agents design steel building frames,
scored by a 3D 6-DOF direct-stiffness solver and Eurocode 3 checks. Three graded tasks from a
single-story warehouse to a seismic hospital, served via a Dockerized FastAPI server.
|
|
|
RealPDE: Long-Term Test-Time Adaptation (NeurIPS 2026 Competition)
GitHub
Adapting neural-operator forecasters (CNO, FNO, Transolver) for streaming
real-world PIV airfoil-wake flow. A bounded adaptation controller (skip / recalibrate / anchored
norm-layer update) with uncertainty bounds and LLM-driven action selection under a fail-safe rule
fallback.
|
|
|
JED Red-Team: Multi-Step Tool Attacks on LLM Agents (Kaggle)
GitHub
Reverse-engineered a dataflow guardrail's predicate logic (exfiltration,
destructive-write, confused-deputy, untrusted-to-action) to expose exploitable scoring gaps, and
built a search harness replaying multi-step tool-call chains against guardrailed GPT-OSS/Gemma
agents.
|
|
|
Are VLMs Really Blind?
GitHub
/
arXiv
Benchmarked Gemini, PaliGemma, and GPT-4o on 8 low-level geometric
reasoning tasks where VLMs underperform despite strong OCR/VQA results. Proposed a keyword-guided
captioning pipeline boosting zero-shot accuracy by up to 14pts without any fine-tuning.
|
|
|
Indian Institute of Technology Roorkee
Bachelor of Technology in
Materials Engineering
Aug 2023 - May 2027
Activities:
- Head of Research | Vision and Language Group (VLG)
- 1st Place - HiLabs Hackathon (50+ teams), 2024
- 1st Place - Composio IITD Hackathon, national winner (100+ teams), 2025
|
|