You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #68 (F2b: per-item execution and the runner's commit step). Independent review of #68 found five gaps. They open only once the real lane D workspace from #66 part 2 replaces the test fake: until then, no production code runs ItemExecutor. They were deferred from #68 on purpose. Related: #66, #22, #51.
1. Partial-output export before task storage is removed
docs/implementation/runner-lifecycle.md (task storage row): "settled → the partial-output export if the attempt did not complete → the terminal write, which stores diagnostic_ref → removeTaskFilesystems". plan-format.md, "After each run": for a safety violation, "preserve the rejected output separately for diagnosis".
Today:TaskWorkspace has no export operation. A writable attempt that does not complete, including a safety violation, has its storage removed with no export and no diagnostic_ref.
2. Audit the changes of an attempt whose agent exited non-zero
Today:finish() runs only on a clean exit. An agent that makes unsafe changes (for example under .git) and then exits non-zero never reaches needs human: its work is discarded, and the task can be retried.
Contract: "Any safety violation stops tests and commits and moves the task to needs human."
Needs: the same inspection as item 1, run on failed attempts too.
3. Durable safety findings
Today:SafetyFindings is in memory. A finding is lost when:
the terminal write fails (the attempt stays active, so the move to needs human is refused);
the process crashes between the terminal write and that move.
After a restart, startup recovery finalizes the attempt as "Interrupted", and the task can be retried.
Proposed: a durable mark on the attempt, written like the first reason, before the terminal write. The terminal write, startup recovery and admission (retry) would all respect it.
4. Act on a finding outside ItemExecutor
Today: an execute attempt started through runner.start or runner.retry (for example the /api/runner retry) records a finding that nothing takes. The task stays running, and the finding stays in memory.
Fix: item 3 covers this. The move to needs human must not depend on which caller started the attempt.
5. Take the runner's commit out of task storage before release
Today: the runner commit exists only inside task storage, and release removes it:
the next item's materialize(head) has no way to get that object;
the review screen's HEAD observation (ReviewService.load) compares config.repository HEAD with the snapshot head. It would then record the old HEAD as a new snapshot, which rolls back the base and makes the next attempt stale.
Plan (D: inspect task changes and make runner commits (needed by F2) #66 decision 1):exportTaskCommit(storage, {base, head}) bundles the commit into a runner-owned host repository, and F verifies the SHA. TaskWorkspace needs that step, run after the terminal write and before release.
Also needed: the review screen's HEAD observation must account for runner commits.
Follow-up to #68 (F2b: per-item execution and the runner's commit step). Independent review of #68 found five gaps. They open only once the real lane D workspace from #66 part 2 replaces the test fake: until then, no production code runs
ItemExecutor. They were deferred from #68 on purpose. Related: #66, #22, #51.1. Partial-output export before task storage is removed
docs/implementation/runner-lifecycle.md(task storage row): "settled→ the partial-output export if the attempt did not complete → the terminal write, which storesdiagnostic_ref→removeTaskFilesystems". plan-format.md, "After each run": for a safety violation, "preserve the rejected output separately for diagnosis".TaskWorkspacehas no export operation. A writable attempt that does not complete, including a safety violation, has its storage removed with no export and nodiagnostic_ref.2. Audit the changes of an attempt whose agent exited non-zero
finish()runs only on a clean exit. An agent that makes unsafe changes (for example under.git) and then exits non-zero never reaches needs human: its work is discarded, and the task can be retried.3. Durable safety findings
Today:
SafetyFindingsis in memory. A finding is lost when:After a restart, startup recovery finalizes the attempt as "Interrupted", and the task can be retried.
Proposed: a durable mark on the attempt, written like the first reason, before the terminal write. The terminal write, startup recovery and admission (retry) would all respect it.
4. Act on a finding outside
ItemExecutorrunner.startorrunner.retry(for example the/api/runnerretry) records a finding that nothing takes. The task staysrunning, and the finding stays in memory.5. Take the runner's commit out of task storage before release
releaseremoves it:materialize(head)has no way to get that object;ReviewService.load) comparesconfig.repositoryHEAD with the snapshot head. It would then record the old HEAD as a new snapshot, which rolls back the base and makes the next attempt stale.exportTaskCommit(storage, {base, head})bundles the commit into a runner-owned host repository, and F verifies the SHA.TaskWorkspaceneeds that step, run after the terminal write and beforerelease.🤖 Generated with Claude Code