Skip to content

F2b: per-item execution and the runner's commit step - #68

Merged
mchwang merged 26 commits into
mainfrom
feat/f2b-item-execution
Oct 1, 2026
Merged

mchwang merged 26 commits into
mainfrom
feat/f2b-item-execution

Conversation

@mchwang

@mchwang mchwang commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Lane F, step F2, slice F2b: per-item execution and the runner's commit step. Based on main. F1 (#53, #74, #57, #59, #60) and F2a (#67) are merged. Related: #22, #66, #51.

Restacked after #67 merged: head fe4b427, on main 3ac388a. The three commits are this PR's own; the F2a copy dropped out because #67 is on main. The file tree is identical to 3a62df0, the head CI last passed on.

What this does

Implements "codeboost makes every commit" (design, "How codeboost runs a plan") against a fake D workspace. The four workspace operations are the ones proposed in #66 (materialize, snapshotDeclaredLinks, inspectChanges, commit, plus release).

  • runner/execution.ts: executionDeps is the RunnerDeps for execute attempts:
    • prepare: a fresh workspace from the recorded head of the attempt's snapshot; a no-follow snapshot of declared symlinks; the prompt and approvedArgv from prepareExecution (F2a).
    • finish: inspectChanges, then auditRun (F2a).
      • A safety violation throws FinishFailure("Safety violation: …"): no commit, and the attempt fails with that diagnostic.
      • No changes: "planned but unchanged".
      • Otherwise the runner's own commit, with Plan-Item: P<n> and Plan-Revision: r<k> trailers, of exactly the audited paths, matched to the audit's digest. It returns the owned ledger entry.
    • release: removes the workspace.
  • ItemExecutor.runTask runs the plan items in order, one attempt each:
    • out-of-scope files are committed with the item, then a checkpoint is recorded and the task moves to needs amendment;
    • a safety violation moves the task to needs human;
    • any other non-completed attempt stops the run and reports why.
  • Coordinator (F1b):
    • an optional asynchronous finish step replaces validate for writable attempts; FinishFailure keeps its own diagnostic;
    • an optional release step runs after the terminal write, on every path, including endings before launch;
    • a failed release keeps the slot under a marker.
  • Store (F1a): settleAttempt takes a history record and calls recordHistory inside the same transaction as completed, after its guards. This is the publication order in the F1 contract. A commit made in task storage but not published (a refused CAS, or a stop during finish) is discarded with the storage.

Validation (head f0957ad)

  • npm run typecheck: passes.
  • CI's unit set: 552 passed, 0 failed. 8 of those are new, in test/runner-execution.test.ts, using the real Store, coordinator and executor with a fake workspace and launcher:
    • two items in order, with trailers, ledger owners, and each item starting from the previous commit;
    • an unchanged item;
    • out-of-scope files, giving a checkpoint and needs amendment;
    • a violation, giving needs human with no commit;
    • an agent failure;
    • a refused commit, with no ledger entry;
    • a stop during the commit, which is discarded;
    • a release failure, giving a marker;
    • the release always after the terminal write.
  • Mutation check: six guards were broken one at a time, and each was caught:
    • history not recorded with completed;
    • release before the terminal write;
    • a violation leaving the task running;
    • a commit despite a violation;
    • no Plan-Revision trailer;
    • out-of-scope files not pausing.

Third rebase (after #82 and #83)

Head 2afca97, restacked on #60's 99cf5e2. All four commits applied without conflicts. F2a and F2b have no Git calls to harden.

  • Typecheck passes.
  • CI's unit set: 833 passed.
  • npm run test:browser: 65 passed.

Code review (head 2caf235)

This PR's own commits were reviewed after the lower slices' fixes. There were no findings. The branch is restacked on #60's 29ea84e.

  • CI's unit set: 829 passed.
  • npm run test:browser: 65 passed.

Second rebase (after #74 merged)

Head 7eab124, restacked on #60's 3ab2359.

  • runner/coordinator.ts. The merged F1b: runner coordinator for attempts, slots and shutdown #74 rewrote launch and settlement, so this PR's changes were applied to main's version:
    • prepared is passed to every #endBeforeLaunch after preparation;
    • finish() and the ledger history go into settlement;
    • task storage is released after a saved terminal write on every path. That includes main's new foreign-result path, which this PR predates.
    • #release no longer overwrites a preparation marker that main already set.
  • New commit: two tests. Task storage is released on the foreign-result path and on the launch-failure path. Before this, removing either release failed no test. The launch-failure gap was already in this PR.
  • Validation on 7eab124.
    • Typecheck passes.
    • CI's unit set: 824 passed.
    • npm run test:browser: 65 passed.

Rebase onto main (after #53 merged)

Restacked on the rebased #60. Head 7cf7f63. Both commits applied without conflicts.

  • New commit: pass the runner owner token into execution deps. F1b: runner coordinator for attempts, slots and shutdown #74 made RunnerDeps.runnerOwner required, because D checks it against the task storage owner.
    • executionDeps now takes a runnerOwner argument, meant to come from F1d's Store.runnerOwnerToken().
    • TaskWorkspace must allocate storage under that token. TaskWorkspace is D's contract (D: inspect task changes and make runner commits (needed by F2) #66), so this is documented rather than enforced.
    • The first execution test checks that D receives the token.
  • Validation on 7cf7f63.
    • Typecheck passes.
    • CI's unit set: 744 passed, 0 failed. The first full run had four git-heavy timeouts, which passed on the rerun.
    • npm run test:browser: 63 passed.

Independent review fixes (head 3b59a5b)

Two rounds of review by a fresh subagent, checked against runner-lifecycle.md and plan-format.md's "After each run". Each fix has a test that fails without it.

Round 1:

  • Safety findings are recorded by the runner (SafetyFindings). They are never read back from diagnostic text, so an agent's stderr can no longer move its task to needs human, and a violation survives a later stale or stop outcome.
  • auditRun treats any agent commit as a violation (D: inspect task changes and make runner commits (needed by F2) #66 decision 2: no undo path). This reverses F2a's test, which counted a manifest with only agent commits as "unchanged".
  • An inspection that refuses sends the task to needs human (Publishing step 2).
  • Storage allocated before preparation fails is removed after the terminal write (PreparationFailure).
  • The run stops before the next item when the plan gets a new revision.
  • The checkpoint and the move to needs amendment commit in one transaction (Store.pauseForAmendment), through the shutdown capability.
  • A rename stages both paths.
  • A storage-removal failure is marked storage-not-removed, not "result could not be saved".
  • runTask returns a stopped outcome with the completed items, instead of throwing, when admission is refused or the terminal write failed.

Round 2:

  • A scope finding always pauses the task. The checkpoint is bound to the revision the item ran against and the snapshot its own commit created, so a revision saved during release cannot drop the pause. The round-1 fix had caused this regression.
  • The stop signal and the context are re-checked before the runner commit, after the last await.
  • A malformed change report is a safety violation.
  • The foreign-result path removes preparation files after the terminal write.
  • snapshotDeclaredLinks gets every declared path.

Round 3 (head d8713cc):

  • auditRun validates every field it reads before reading any: the metadata flag, each change's kind, entry types, underGit, a rename's old path, link target and traversal flag, and total path bytes. A partial record now fails closed instead of reading as clean.
  • A scope pause that was never recorded is paused at the start of the next run. This covers a failed write, the write gate and a crash, so no run can go past it.
  • Between items the executor checks the whole context: the snapshot the previous commit created, the assignment and the referenced code.
  • An invalid commit ID from the workspace fails the attempt instead of breaking the terminal write.
  • Only the stop's own abort error counts as the stop. Any other inspection refusal stays a safety finding.
  • Pausing and escalating keep a status set during release.

Round 3 left one question open: whether "Irreversible actions" (re-read review_version and closing before each commit) covers the runner commit inside task storage. It is only published by the guarded settle transaction, so this PR does not add those re-reads.

Validation: typecheck passes; CI's unit set: 860 passed. Every fix has a test that fails when the fix is removed.

Independent review, rounds 4–23 (head 5a7c1c8)

Twenty more review rounds by a fresh subagent, checked against runner-lifecycle.md, plan-format.md and AGENTS.md. Each fix has a test that fails when the fix is removed, confirmed by mutation. Rounds 20–23 found no correctness bug, only test gaps and hardening.

The main changes:

  • Scope pauses.
  • Safety findings.
    • Findings are recorded by the runner, never parsed from text. A finding is settled only once acted on.
    • Human gates keep it owed. Review statuses are escalated, so a task cannot be merged past a finding.
    • A finding owed from an earlier run is paid without the shutdown capability.
  • One run per task.
    • An in-flight guard and the coordinator's active-job check stop a second run from paying another run's outcome.
    • Between items, the run stops on any status, plan revision or context change, including the context generation.
  • auditRun fails closed on partial or malformed reports:
    • every field, kind and entry type is checked before use;
    • duplicates are refused, except a split case-only rename;
    • paths must be canonical;
    • every .git spelling Git refuses is refused (NTFS and HFS);
    • link targets are checked as written and resolved (Windows forms, the root, ..);
    • the size limit is measured as the saved JSON;
    • agent text is quoted and bounded.
  • Diagnostics. Agent and D text is quoted (stderr, preparation, inspection and commit errors, log lines). The commit title is one line, with no control, bidi or format characters, so it cannot forge trailers.
  • Task storage. Removed only after a saved terminal write, on every path, including a preparation failure after allocation.

Deferred:

Open questions for the owner:

  1. Do the "Irreversible actions" re-reads apply to the runner commit inside task storage?
  2. Should changes to .gitmodules or .gitattributes be violations?
  3. Should this run's own pause and escalation keep using the shutdown capability after the gate closes?

Validation on 5a7c1c8:

  • typecheck passes;
  • CI's unit set: 937 passed.

Not in this slice

🤖 Generated with Claude Code

@mchwang
mchwang force-pushed the feat/f1e-planning-feedback branch from bfe30df to 14cb1e0 Compare September 29, 2026 17:33
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch from f0957ad to 7cf7f63 Compare September 29, 2026 17:33
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch from 7cf7f63 to 7eab124 Compare September 29, 2026 20:02
@mchwang
mchwang force-pushed the feat/f1e-planning-feedback branch from 14cb1e0 to 3ab2359 Compare September 29, 2026 20:02
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch from 7eab124 to 2caf235 Compare September 29, 2026 20:49
@mchwang
mchwang force-pushed the feat/f1e-planning-feedback branch from 29ea84e to 99cf5e2 Compare September 29, 2026 23:57
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch from 2caf235 to 2afca97 Compare September 29, 2026 23:57
@mchwang
mchwang force-pushed the feat/f1e-planning-feedback branch from 99cf5e2 to 86fc293 Compare September 30, 2026 01:12
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch 2 times, most recently from 99a6546 to 98e1c66 Compare September 30, 2026 02:51
@mchwang
mchwang force-pushed the feat/f1e-planning-feedback branch from 86fc293 to 6bb34b5 Compare September 30, 2026 02:51
@mchwang
mchwang changed the base branch from feat/f1e-planning-feedback to main September 30, 2026 02:58
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch from 98e1c66 to 3a62df0 Compare September 30, 2026 02:59
mchwang and others added 3 commits September 29, 2026 20:44
ItemExecutor runs a task's plan items in order as execute attempts.
executionDeps materializes a fresh workspace and snapshots declared links,
builds the prompt with prepareExecution, and in finish() inspects changes,
audits them with auditRun, and makes the runner's own commit with
Plan-Item/Plan-Revision trailers. The coordinator gains an async finish step
and a release step after the terminal write; settleAttempt records the
owned ledger entry with completed in the same transaction. A safety violation
moves the task to needs human; out-of-scope files are committed and pause it
in needs amendment with a checkpoint. The workspace is D's (#66), faked here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RunnerDeps now carries runnerOwner (#74), which D requires on every
invocation and checks against the task storage owner. executionDeps takes
the database's token and the workspace allocates storage under it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aths

After the rebase onto main, the coordinator releases task storage after
the terminal write on main's foreign-result path too. Neither that path
nor the launch-failure path had a test that failed without the release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mchwang
mchwang force-pushed the feat/f2b-item-execution branch from 3a62df0 to fe4b427 Compare September 30, 2026 03:44
mchwang and others added 2 commits September 30, 2026 09:15
…every path, stable plan, atomic pause

- Safety violations are recorded by the runner's own audit (SafetyFindings),
  never read back from diagnostic text, so agent stderr cannot move a task to
  needs human, and a violation survives a stale or stop outcome.
- auditRun treats any agent commit as a violation, per #66 decision 2
  (no undo path). The TaskWorkspace.commit comment now says so.
- An inspection that refuses sends the task to needs human (contract,
  Publishing step 2); an aborted one is the stop, not a finding.
- Preparation that fails after allocating task storage hands it to the
  coordinator (PreparationFailure), which removes it after the terminal write.
- The executor stops before the next item when the plan gets a new revision,
  and binds a checkpoint to the revision the item ran against.
- The checkpoint and the move to needs amendment commit in one transaction
  (Store.pauseForAmendment), through the shutdown capability; a refused pause
  or status change returns a stopped outcome instead of throwing.
- A rename stages both paths in the runner commit.
- A storage-removal failure holds the slot as 'storage-not-removed', not
  'result-not-saved'.
- runTask returns stopped, with the items it completed, when admission is
  refused or the terminal write failed.
- A test covers that storage is never removed when the terminal write fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…commit, fail closed on bad reports

- A scope finding always pauses the task: the checkpoint is bound to the
  revision the item ran against and the snapshot its own commit created
  (Store.snapshotWithHead), not to what is current, so a revision saved
  during release no longer drops the pause. The round-1 fix had turned
  that case into a stopped outcome, which a later run skipped past.
- pauseForAmendment validates against the item's own revision; only a
  refused status change (a closed task) returns stopped, other errors throw.
- Before the runner commit, after the last await, finish re-checks the stop
  signal and the captured context; a change makes no commit.
- A change report missing any list fails closed as a safety violation, and an
  audit that throws is a violation too.
- The foreign-result path removes host-side preparation files after the
  terminal write, like every other path.
- snapshotDeclaredLinks gets every path the item declares, so D checks the
  actual entries (an earlier item may have renamed or added links).
- Tests: the pause is atomic (a refused pause leaves no checkpoint), an
  aborted inspection is the stop, a stop or context change during the audit
  makes no commit, malformed reports, foreign-result cleanup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mchwang and others added 2 commits September 30, 2026 09:38
…nnot be skipped, whole-context checks

- auditRun validates every field it reads before reading any: the metadata
  flag, each change's kind, entry types, underGit, rename old path, link
  target and traversal flag, and the total path bytes. A partial record fails
  closed as a violation instead of reading as clean (AGENTS.md).
- A scope finding whose pause was never recorded (a failed write, the write
  gate, a crash) is paused at the start of the next run, before any item, so
  no run goes past it. A closed write gate is no longer taken for a refusal.
- Between items the executor checks the whole context: the snapshot must be
  the one the previous item's commit created, and the assignment and
  referenced code unchanged, not only the plan revision.
- An invalid commit ID from the workspace fails the attempt instead of
  breaking the terminal write.
- Only the stop's own abort error counts as the stop; another inspection
  refusal stays a safety finding even when a stop is pending.
- Pausing and escalating keep a status someone set during release: both
  require the task to still be running.
- Tests for each, and for the audit-throws, rename-snapshot and checkpoint
  snapshot claims that had none.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A pause owed from an earlier run is paid when the task is queued again,
  not only while it runs, so a task can no longer get stuck on it.
- A recorded pause holds until a person approves continuing on the current
  revision (continuationRevision); the continued run must name the next item,
  not one the checkpoint already completed.
- A checkpoint is found by its item and its commit's head, not by the latest
  snapshot, so a later snapshot with the same head cannot make an approved
  pause owed again.
- auditRun refuses a change path of "." or "..", and an old path on any
  kind but rename; the path-size limit counts only the saved paths.
- finish refuses a commit equal to the base for a changed item.
- A slot held under a marker is reported with that marker's real cause.
- Tests for each, and for the referenced-code-only context check, the
  AbortError stop, the per-kind entry rules and foreign-result cleanup only
  after a saved terminal write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mchwang and others added 5 commits September 30, 2026 09:59
…ical paths, JSON-sized limit

- A task with a scope checkpoint runs no further items. Continuing after a
  scope pause needs its own design (reconcile the prefix, validate the rest
  from the checkpoint, consume the approval), tracked in #88; round 4's
  partial approval gate could block an approved continuation forever or skip
  items. Before round 4 the executor ran past the pause.
- A checkpoint is found by its commit's head alone; two items cannot share
  the head of a scope commit.
- This run's own pause and escalation keep any status someone set during
  release, queued included; only a pause owed from an earlier run is paid
  from queued.
- auditRun refuses a path not in canonical form (./, a/../, //, a trailing
  /) and a .git part of any case at any depth.
- The path-size limit measures each saved path as the JSON that stores it.
- Tests for each, and for keeping task storage when a foreign result's
  terminal write fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…us, stricter audit findings

- A safety violation moves the task to needs human over any status someone
  set during release (queued or a human gate), since a run from there would
  launch the item again; only a closed task is left as it is.
- A declared link retargeted into .git in any case or at any depth is
  refused, as paths already are.
- A rename from an undeclared path records that source path as out of scope,
  not only the declared destination.
- A change report that lists one path twice fails closed.
- Agent-controlled paths in findings are quoted and lists cut short, and the
  finding text is bounded (AGENTS.md).
- Tests for each, for task storage removed after the terminal write when a
  stop, a stale context or a cancel task lands during preparation, and for
  pauseForAmendment's plan-prefix check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, undeclared link removals

- A safety finding is settled only once acted on. If moving the task to
  needs human fails, or the task sits at a human gate, the finding stays owed
  and the task's next run escalates it before launching anything.
- Escalation moves only a running or queued task: leaving a human-gated or
  review status needs its own user action (runner-lifecycle.md). Round 6 had
  escalated over any open status.
- Between items, a status someone changed during release stops the run.
- Removing a pre-existing symlink, or turning it into a file, at an
  undeclared path is a safety violation.
- The path-size limit counts a rename's old path, which round 6 made a saved
  scope finding when undeclared. Duplicate paths are found under the plan's
  path identity.
- A refused runner commit fails with a quoted message.
- Tests for each, and for the owed-pause status rule and the prefix check
  against the item's own revision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pellings, undeclared link renames

- A safety finding moves a review status (in review, approved but merge
  blocked) to needs human, so the task cannot be merged past it; only the
  human gates keep the finding owed. A task already in needs human settles
  the finding, so it is not escalated a second time.
- Renaming a pre-existing symlink from an undeclared path is a violation,
  like removing or replacing it.
- isDotGit refuses every spelling Git treats as .git (any case; NTFS
  trailing dots or spaces and git~1; HFS ignorable code points), in paths
  and link targets.
- A case-only rename counts once under a case-folding identity.
- runner-lifecycle.md lists every unresolved marker reason.
- Tests for each, and for escalation refused by a merge in progress (the
  finding stays owed), escalation through the capability after the gate, an
  owed finding settled on a closed task, and completed plus the ledger entry
  being one transaction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…S .git names, links to the root

- A scope pause is recorded over a review status someone set during release
  (in review, approved but merge blocked), in the run that found it and when
  owed, so a merge cannot go past the finding; round 8 did this only for
  safety findings. A queued status set during release is still kept, and the
  next run pays the pause.
- isDotGit ends a name where NTFS does (a stream separator or a backslash)
  before comparing it with .git.
- A declared link retargeted to the repository root (., ./, a/..) is refused:
  the root contains .git.
- A change report without a digest fails closed.
- Tests for each, and for a declared link renamed to an undeclared path, a
  directory entry, an absolute path, every ignorable code point range, and
  escalation over approved but merge blocked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mchwang and others added 14 commits September 30, 2026 10:57
- A case-only rename that Git reports as a delete and an add (the file also
  changed a lot) counts as one path under a case-folding identity, instead of
  a duplicate that sent the task to needs human. Undeclared, the same pair is
  a scope finding; two adds of one folded path are still refused.
- The runner commit's error handling drops a redundant abort special case: a
  stop records its first reason before it aborts, so the outcome is that stop.
- Tests for an owed pause paid from a review status, a pause never
  overwriting needs human, possibly already fixed as a gate, task storage kept
  before launch when the terminal write fails, checkpointAtHead, and the
  unknown-snapshot check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…and one add

- The duplicate rule keeps every entry per folded path and allows a second
  entry only for a case-only rename reported as exactly one delete and one
  add with different spellings. Round 10 compared each entry only with the
  last one, so add, delete, add (two adds of one path) got through.
- The runner commit message writes the plan title on one line, so it cannot
  open a trailer block that forges Plan-Item or Plan-Revision.
- Looking for .git reads a backslash as a directory separator (NTFS).
- The completed-list comment says what a thrown error carries.
- Tests for each, and for an assignment-only change between items, D's
  underGit flag on its own, task storage removed when the budget is spent at
  the launch check, and checkpointAtHead finding an older checkpoint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…errors, remaining audit tests

- A link target is checked for .git as written as well as resolved, so a
  .git part that a later .. cancels (.git/../src) is refused, as non-canonical
  changed paths already are.
- A preparation error from the workspace is quoted in the diagnostic, like the
  inspection and commit refusals.
- Tests for needs amendment as a human gate, snapshotWithHead picking the
  latest snapshot, .git behind a backslash in a link target, empty and NUL
  link targets, a NUL in a path, directory, other and gitlink old entries, the
  full HFS ignorable ranges, the 300-character quote cut, and an empty digest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hutdown between items

- The coordinator quotes every preparation error in the diagnostic, so a
  materialize error that names an agent-chosen path cannot forge a second
  line. Round 12 had quoted only the declared-link snapshot's errors.
- Tests for runTask returning stopped with the completed items when shutdown
  refuses the next item's admission, and for a pause's executed prefix being
  built from the plan the item ran against when a revision inserts an item
  before it. The existing revision-during-release test now asserts that the
  revision really changed, since a failed import in release is absorbed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ok tests

- The agent's stderr is quoted in the attempt diagnostic, so it cannot forge
  a runner line such as "Safety violation:" in the stopped reason. Round 13
  quoted preparation errors for the same threat.
- Tests that set up a change inside the release hook now also assert that
  no storage marker was left and the stop names the change, so a failing
  setup there cannot make them pass vacuously.
- Tests for a rename with both sides undeclared and for an AbortError from
  the inspection with no stop pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gets

A declared link retargeted with a backslash or a drive prefix (..\..\outside,
C:\Windows) passed the "leaves the repository" and "absolute" checks, which
use POSIX paths, although the audit already reads a backslash as a separator
when it looks for .git. Such link text is refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a run still finishing

- runTask refuses to start while an earlier run of the task still has an
  active job (its storage release), so a second run cannot pay the first
  run's scope pause or safety finding and change that run's outcome.
- Paying an owed finding or pause at the start of a run is that run's own
  decision, not settlement: it writes without the shutdown capability, and a
  closed write gate leaves it owed.
- A finding whose terminal write failed reports the needs-restart cause and
  stays owed, instead of "An attempt is still active".
- The coordinator's log lines quote D's error text.
- Tests for each, and for a stop that lands before a rejected commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e remaining claims

- The owed-pause check looks at every completed execute attempt, not only
  the latest, so a later clean attempt cannot hide an earlier finding whose
  pause was never recorded.
- Tests for an owed scope pause left owed (not paid through the shutdown
  capability) once the write gate closed, the ledger record's base, and the
  coordinator's log lines quoting D's error text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ms, stop after shutdown began

- Between items the executor compares the whole context with what the
  previous item left: its commit's snapshot and a context generation raised by
  exactly that commit. A change that only bumps the generation (for example a
  same-valued reassignment) now stops the run too.
- A run started after shutdown began pays nothing owed and starts nothing.
- The needs-human settle test checks that no extra write happened.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…omes report not started

- ItemExecutor holds a per-task in-flight flag from the start of runTask to
  its return, so a second run can never pay the first run's pause or finding,
  even in the microtasks between the coordinator dropping the job and the
  first run resuming (which the isActive check alone left open).
- Paying or keeping an owed finding or pause reports the earlier item with
  state "not started", not that attempt's old state.
- runner-lifecycle.md says when preparation-not-removed happens on each path.
- Tests: a second run started at every microtask offset in the first run's
  release never pays its pause; owed outcomes report not started.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tside-job check, line separators

Round 20 found no correctness bug. These tests close its gaps: an owed
finding at a human gate, on a closed task and with a failed terminal
write, and an owed pause refused at a gate, each reported as not started;
runTask refusing while a job started outside the executor is still
releasing its storage; and U+2028 in a plan title.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…wed finding

Round 21 found no correctness bug. Tests now cover a scope pause at P2
(the checkpoint names the whole executed prefix) and an older owed finding
escalated although a later attempt completed cleanly. The ItemExecutor
docs state that its one-run-per-task guard needs one executor per Store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…itle, test a link to ..

- The runner commit's title drops every control character, not only line
  breaks, so an agent-written plan title cannot put terminal escapes into
  git log. (A NUL is already refused earlier, by the prompt builder.)
- A test covers a declared link retargeted to "..", the directory that
  holds the repository.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ommit title, two more tests

Round 23 found no correctness bug.
- The runner commit's title also drops bidi and invisible format characters
  (U+200B-U+200F, U+202A-U+202E, U+2060-U+206F, U+FEFF), so an agent-written
  plan title cannot reorder how git log shows it.
- Tests cover the C1 control range in that filter, and this run's own
  escalation rethrowing a closed write gate's error when it has no capability.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mchwang
mchwang marked this pull request as ready for review October 1, 2026 03:08
@mchwang
mchwang merged commit 872b0d9 into main Oct 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant