Skip to content
View ayushsi42's full-sized avatar

Highlights

  • Pro

Organizations

@vlgiitr

Block or report ayushsi42

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ayushsi42/README.md

Hi, I'm Ayush Singh πŸ‘‹

Final-year B.Tech (Materials Engineering) @ IIT Roorkee Β· AI/ML Research

ayushsi42

  • πŸ”¬ I do research on LLM alignment/RLHF, multi-agent systems, and AI red-teaming/security β€” previously at Adobe MDSR, Prem AI, Repello AI, Boston University, and Lossfunk.
  • πŸ“ My work on preference optimization (IPO) and reasoning alignment (ACL 2026 Findings) has been published at ACL Main Conference / Findings and AAAI workshops.
  • πŸ› οΈ Currently building agentic pipelines, RL environments for LLM agents, and test-time adaptation methods β€” see pinned repos below.
  • πŸ’¬ Ask me about LLMs, RLHF/DPO/GRPO, multi-agent systems, or AI security/red-teaming.
  • πŸ“« Reach me at ayushsingh73920@gmail.com

πŸ“š Publications

  • CATPO: Critique-Augmented Tree Policy Optimization β€” arXiv Preprint, 2026 β€” Paper
  • Thinking About Thinking: Evaluating Reasoning in Post-Trained LLMs β€” ACL 2026 Findings β€” Paper
  • IPO: Your Language Model is Secretly a Preference Classifier β€” ACL 2025 Main Conference β€” Paper
  • LoRA-Mini: Adaptation Matrices Decomposition β€” AAAI CoLoRAI Workshop β€” Paper
  • Adaptive Urban Planning: A Hybrid Framework β€” AAAI AI4UP Workshop β€” Paper

πŸš€ Featured Projects

  • πŸ”­ RAVEN β€” LangGraph hypothesis-validation copilot generating audit-ready financial research reports.
  • πŸ—οΈ StructuralDesignEnv β€” OpenEnv RL environment where LLM agents design Eurocode-3-compliant steel frames.
  • 🌊 RealPDE-LTTTA β€” Long-term test-time adaptation for streaming real-world PIV flow forecasting (NeurIPS 2026 competition).
  • πŸ›‘οΈ JED Red-Team β€” Multi-step tool-attack search against guardrailed LLM agents (Kaggle).

Connect with me:

ayush_singh_iitr ayush-singh-iitr ayush-singh-iitr

Pinned Loading

  1. raven raven Public

    RAVEN is a LangGraph-powered copilot that helps equity analysts validate investment hypotheses fast. Paste a thesis and the agent assembles an end-to-end diligence loop: structured planning, data c…

    Python 1

  2. realpde-lttta realpde-lttta Public

    Python

  3. structural_design_env structural_design_env Public

    An RL environment where an agent simultaneously learns physics and uses it to design functional architecture.

    Python

  4. behavior-in-the-wild/web-experience-benchmark behavior-in-the-wild/web-experience-benchmark Public

    Benchmarking for Evaluating Web Experience (Core Web Vitals, etc)

    Python 1 1

  5. vlgiitr/Are-VLMs-Really-Blind vlgiitr/Are-VLMs-Really-Blind Public

    Python 4

  6. shivank21/Implicit_Preference_Optimization shivank21/Implicit_Preference_Optimization Public

    https://arxiv.org/pdf/2502.16182v2

    Python 8 1