Hello, I work at Stemuli training math tutoring models.
Research bench where you can find open source software and reports: keppylab.com
Collective research & writing: yokozunasan.com
titans-mini A streaming engine over a swappable test-time memory core (MLP-weights memory vs. generated-weights vector memory)
thomas thomas.train() — a training harness. Case→reward→train: take a Case set and a score function, get a baseline card (gonogo), run a training loop, compare before/after. Two paths: pretrain (Modal GPU) and post-train (Tinker LoRA RL or TRL GRPO + vLLM).
cotfaith CoT (un)faithfulness, study one: hint-following and confession rates on Qwen3-1.7B — pre-registered decision log, blind-labeled judge validation, byte-exact run artifacts
gonogo The eval harness from my implementations, open sourced. Scores an agent on your real cases and returns a deployment decision, including "not enough evidence yet."
describe• Speak systems into existence
cobol-reporter 🔭 A RAG and report generator for helping to understand COBOL code
Disease Lab 🧪 Knowledge graph informed AI for disease research
WorldEnder.ai 🌎 A text-adventure about extinction events with a RAG backend




