AGENT INFRASTRUCTURE | ROBOTICS & PHYSICAL AI
Sandboxes, sensors and decision ledgers for agents that act on real systems.
Most of my work comes back to one gap: an agent calls a tool, reads its own report of what happened, and carries on, so nothing outside the agent ever checks the result. stoat gives the agent a disposable machine to act in, midwire measures what actually changed, and docket keeps the decisions and their reasons so the next session starts from them.
Lately the same question has moved to hardware: how a robot confirms that its action landed, what it should remember, and which limits it cannot talk its way past.
01 // AGENT INFRASTRUCTURE
- STOAT: local QEMU VMs for people and agents. Disposable Alpine sessions and persistent Linux guests from plain TOML, driven through a TUI, a JSON CLI or MCP tools. One Go binary.
- MIDWIRE: closed-loop agency. Confirms that a tool call's side effect landed before the agent takes its next step.
- DOCKET: an append-only decision ledger for coding agents. When a conversation resumes or gets compacted, it briefs the agent on what was decided, why, and what is still open.
- OPENCODE-NATIVE-SWARMS: permission-scoped background-agent workflows for OpenCode. It can research, review and run narrow checks, and has no path to edits, pushes or external writes.
- SKILLOPT: trains reusable natural-language skills for frozen agents through trajectory-driven edits and validation-gated updates.
- MONEY-MESH (private): a leaderless mesh of self-replicating earning agents under an immutable core. A Cedar policy engine sits at the enforcement point, revenue counts only from ground truth, and the spend cap is enforced by arithmetic.
02 // CLAUDE CODE PLUGINS
- PALPATINE: a "strategic" advisor for Claude Code and Codex. It names the actual problem and the actions to take, in 50 words.
- CURT: editing guidance and a heuristic prose linter that cuts AI slop and keeps the author's voice.
- READABLE-RESPONSES: a readability hook and a de-slopify skill for Claude Code and Codex CLI, because Opus 5 yaps too much.
- POLISHED-ARTIFACTS: one design language for Claude artifacts, in Carbon and Apple looks, with a lint for rendered pages.
03 // EPISTEMIC MEMORY
Memory that separates what an agent observed from what it concluded and what it made up.
04 // OTHER THINGS
05 // RESEARCH LOG
- Can an agent get a disposable machine as easily as a shell? (stoat)
- Can decisions outlive the session that made them? (docket)
- Can skills be learned in text space, without touching weights? (SkillOpt)
- Can stratified, warrant-backed memory run locally? (LeAP, shipped as Veil)
- Can an agent confirm its side effects before it continues? (midwire, in progress)
- Generation got cheap and verification didn't. What closes that gap? (essay)
- Can autonomy be bounded by a policy engine and arithmetic, rather than by the prompt?
- What should a robot remember? Object states, failures and maps, retrieved under partial observability without surfacing stale entries.
- World action models: predicting the consequences of an action in latent space instead of rendering future frames.
- Closing the loop on hardware: a VLA policy that checks its action landed through sensors, not through its own prediction.
- Safety bounds a VLA policy cannot override: runtime envelopes enforced outside the model.
- Cross-embodiment transfer: data from one robot's kinematics and sensors that still teaches another.




