Skip to content

F2b: connect the real workspace (partial export, failed-run audit, durable safety findings, commit export) #87

Description

@mchwang

Follow-up to #68 (F2b: per-item execution and the runner's commit step). Independent review of #68 found five gaps. They open only once the real lane D workspace from #66 part 2 replaces the test fake: until then, no production code runs ItemExecutor. They were deferred from #68 on purpose. Related: #66, #22, #51.

1. Partial-output export before task storage is removed

docs/implementation/runner-lifecycle.md (task storage row): "settled → the partial-output export if the attempt did not complete → the terminal write, which stores diagnostic_ref → removeTaskFilesystems". plan-format.md, "After each run": for a safety violation, "preserve the rejected output separately for diagnosis".

  • Today: TaskWorkspace has no export operation. A writable attempt that does not complete, including a safety violation, has its storage removed with no export and no diagnostic_ref.

2. Audit the changes of an attempt whose agent exited non-zero

  • Today: finish() runs only on a clean exit. An agent that makes unsafe changes (for example under .git) and then exits non-zero never reaches needs human: its work is discarded, and the task can be retried.
  • Contract: "Any safety violation stops tests and commits and moves the task to needs human."
  • Needs: the same inspection as item 1, run on failed attempts too.

3. Durable safety findings

  • Today: SafetyFindings is in memory. A finding is lost when:

    • the terminal write fails (the attempt stays active, so the move to needs human is refused);
    • the process crashes between the terminal write and that move.

    After a restart, startup recovery finalizes the attempt as "Interrupted", and the task can be retried.

  • Proposed: a durable mark on the attempt, written like the first reason, before the terminal write. The terminal write, startup recovery and admission (retry) would all respect it.

4. Act on a finding outside ItemExecutor

  • Today: an execute attempt started through runner.start or runner.retry (for example the /api/runner retry) records a finding that nothing takes. The task stays running, and the finding stays in memory.
  • Fix: item 3 covers this. The move to needs human must not depend on which caller started the attempt.

5. Take the runner's commit out of task storage before release

  • Today: the runner commit exists only inside task storage, and release removes it:
    • the next item's materialize(head) has no way to get that object;
    • the review screen's HEAD observation (ReviewService.load) compares config.repository HEAD with the snapshot head. It would then record the old HEAD as a new snapshot, which rolls back the base and makes the next attempt stale.
  • Plan (D: inspect task changes and make runner commits (needed by F2) #66 decision 1): exportTaskCommit(storage, {base, head}) bundles the commit into a runner-owned host repository, and F verifies the SHA. TaskWorkspace needs that step, run after the terminal write and before release.
  • Also needed: the review screen's HEAD observation must account for runner commits.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions