- π¬ I do research on LLM alignment/RLHF, multi-agent systems, and AI red-teaming/security β previously at Adobe MDSR, Prem AI, Repello AI, Boston University, and Lossfunk.
- π My work on preference optimization (IPO) and reasoning alignment (ACL 2026 Findings) has been published at ACL Main Conference / Findings and AAAI workshops.
- π οΈ Currently building agentic pipelines, RL environments for LLM agents, and test-time adaptation methods β see pinned repos below.
- π¬ Ask me about LLMs, RLHF/DPO/GRPO, multi-agent systems, or AI security/red-teaming.
- π« Reach me at ayushsingh73920@gmail.com
- CATPO: Critique-Augmented Tree Policy Optimization β arXiv Preprint, 2026 β Paper
- Thinking About Thinking: Evaluating Reasoning in Post-Trained LLMs β ACL 2026 Findings β Paper
- IPO: Your Language Model is Secretly a Preference Classifier β ACL 2025 Main Conference β Paper
- LoRA-Mini: Adaptation Matrices Decomposition β AAAI CoLoRAI Workshop β Paper
- Adaptive Urban Planning: A Hybrid Framework β AAAI AI4UP Workshop β Paper
- π RAVEN β LangGraph hypothesis-validation copilot generating audit-ready financial research reports.
- ποΈ StructuralDesignEnv β OpenEnv RL environment where LLM agents design Eurocode-3-compliant steel frames.
- π RealPDE-LTTTA β Long-term test-time adaptation for streaming real-world PIV flow forecasting (NeurIPS 2026 competition).
- π‘οΈ JED Red-Team β Multi-step tool-attack search against guardrailed LLM agents (Kaggle).