I am a PhD student in Operations Research & Systems Engineering at the University of Miami, with an undergraduate background in Mechanical Engineering. I also work at the Center for Naval Analyses (CNA), the Navy’s federally funded research and development center (FFRDC), where I contribute to data driven research and operational analysis for defense applications. I am also a Research Resident at Prime Intellect, working on open source AI training infrastructure.
- Continual learning methods for large language models
- Reinforcement learning environments for complex optimization problems
Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks (Technical Blog and Toolbox, 2026)
Released a bounded-staleness asynchronous ES trainer for long-horizon, stateful language agents, with an Endless Terminals integration and initial results using Qwen2.5-7B-Instruct.
Blog and Code: GitHub
Introduced a long-horizon agent approach that compacts context at every turn, together with a GRPO and privileged full-history distillation training method that improves task performance while keeping active prompts compact.
Paper: arXiv
Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training (First Author, COLM 2026)
Compared Evolution Strategies and GRPO across four tasks and continual-learning settings, showing that similar task accuracy can arise from geometrically distinct model updates with different implications for forgetting and knowledge preservation.
Extended Prime Intellect's verifiers framework to support multi-agent setups enabling dueling LLMs. Includes environments for multi-agent poker and other games. PR GitHub
A recent write up exploring how Evolution Strategies (ES) can be applied to LoRA adapters instead of full parameter fine tuning, reducing compute cost while enabling exploration-based optimization.
Blog Post: Evolution Strategies + LoRA
Proposed a gated continual learning method for LLMs that applies LoRA based updates only when they meet stability criteria, reducing catastrophic forgetting while enabling efficient sequential adaptation.
Paper: arXiv
Developed YOLO based detection and segmentation models for surgical instrument tracking and designed temporal metrics for evaluating surgical skill.
Paper: ScienceDirect
Hands on exploration of POMDPs and latent representation learning for decision making under uncertainty.
Built reinforcement learning models to optimize routing and decision making for eVTOL operations under uncertainty.
Paper: ACM Digital Library
Project: FAA Data Challenge Finalists
- Reinforcement Learning for Weapon Optimization: Built RL environments to model weapon effectiveness, engagement sequencing, and operational decision making under uncertainty.
- Defense Modeling and Analytics: Conducted end to end modeling work including data engineering, feature development, statistical analysis, and predictive modeling for Navy datasets.

