diff --git a/docs/architecture.md b/docs/architecture.md index 40546caa..c4a494bb 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -1,6 +1,6 @@ # codeboost architecture: a high-level overview -Status: current as of `main` at `10c7c5e` (2026-10-04). Open work is listed in **Build status and roadmap**. +Status: current as of `main` at `1d81303` (2026-10-06). Open work is listed in **Build status and roadmap**. Writing standard: plain language, ISO 24495-1:2023 ## About this document @@ -39,11 +39,12 @@ Writing standard: plain language, ISO 24495-1:2023 5. `agents/` — the isolation boundary: containers, the vendor-only network, and the Claude and Codex adapters. - **The central idea.** A runner-owned **commit ledger** records which plan item each commit belongs to. The linking engine uses the ledger, not commit messages, to put every changed line on a plan item's row. Anything the ledger cannot explain goes to the Unplanned or Ambiguous row. - **The central safety rule.** Everything that comes from outside is untrusted: issue text, agent output, the browser, and the repository's own files. Only the runner decides. Agents never see your credentials, your other files, or GitHub. -- **Build state.** Three levels: +- **Build state:** - **In use:** plan and linking library, SQLite store, review screen, Ask (questions to a Claude agent), guarded merge with merge-queue support, ranked Issues screen, agent isolation boundary, and the single-runner lock. - **In use for a non-demo review with a `github` block:** planning. `/api/plan/suggestions` and `/api/plan/drafts` ask Claude, through lane D, for suggestion cards or a whole next plan revision (#117, #124). Nothing becomes a revision until you apply it. - - **In use with an opt-in `runner` block in `review.json`:** the whole loop after planning. `start` and `resume` on `/api/runner` run a task's plan item by item with a Claude agent (#91). They require current plan-item approvals, bind both the task state and review version, recheck approvals before each later item, and refuse when a completed prefix no longer ends at the task's current head (#107). When a run ends, codeboost publishes the task's pull request, a draft if the task needs a person (#103). Approve & merge merges that pull request (#121). Cancelling a task closes its pull requests (#111). These are API actions; the screen has no buttons for them yet. - - **Not built:** the "trust this issue" action and its runner guard (#108), screens for planning and the runner actions, rebasing, running `cmd:` checks, review rounds, and the Planning, Queue and Learning screens. + - **In use with an opt-in `runner` block in `review.json`:** the whole loop after planning. `start` and `resume` on `/api/runner` run a task's plan item by item with a Claude agent (#91). They require current plan-item approvals, bind both the task state and review version, recheck approvals before each later item, and refuse when a completed prefix no longer ends at the task's current head (#107). After an out-of-scope pause, an approved amended plan can reconcile the completed prefix and resume only the unfinished suffix (#88, PR #134). When a run ends, codeboost publishes the task's pull request, a draft if the task needs a person (#103). Approve & merge merges that pull request (#121). Cancelling a task closes its pull requests (#111). These are API actions; the screen has no buttons for them yet. + - **Partly built for pre-merge automation:** a trusted local rebase engine records one-to-one commit mappings, verifies the resulting checkout byte for byte, owns its subprocesses durably, and recovers interrupted work (#22, PR #136). It is not wired into the merge path yet. + - **Not built:** the "trust this issue" action and its runner guard (#108), screens for planning and the runner actions, the remaining pre-merge rebase/conflict/check orchestration, running `cmd:` checks, review rounds, and the Planning, Queue and Learning screens. ## Terms used @@ -409,7 +410,7 @@ These are the decisions that shape the structure. The full list, with the reason | C, K | Guarded merge, merge queue | Done | | D | Agent isolation boundary | Done. Since the isolation gate: bounded diff export (#76), task-change inspection and the runner's commit as a bounded `git bundle` (#66, PRs #85 and #90), owner-scoped recovery for Ask (#95), the plan schema for Claude (#97), gitlink mounts and the pre-launch tree check (#81, #99, PR #100), and Codex refused in every phase (#93, PR #106). The runner-contract follow-ups (#51) are closed: the last one tests export and removal through a recovery handle after a real restart (PR #109). | | E | Planning library and suggestions | Done. E4 (PR #45) added the acceptance fixtures and recorded Claude output that replays through validation and the Store. | -| F | Runner | Merged: F1 (lifecycle store, coordinator, shutdown wiring, startup recovery and the lock, planning API, feedback events, and stops that wait for their commit: #53, #74, #57, #59, #60, #96); F2a–F2b (execute prompts, audit, per-item execution, real workspace, durable safety findings: #67, #68, #92, #94, #128); F2d (already-fixed check, PR opening, branch pusher, production publishing, closing PRs on cancel: #84, #101, #110, #115, #116, #119, #123); production wiring and `start`/`resume` (#91, PRs #102, #105); planning through lane D (#117, #124, PRs #118, #120, #125); merging the task's own PR (#121, PR #126); start/resume approval, review-version, recovery and completed-prefix guards (#107, PR #130); part of F6 (#89). In review: continue after a scope pause (#88, PR #134). Next after #88: rebasing and `cmd:` checks (#22). | +| F | Runner | Merged: F1 (lifecycle store, coordinator, shutdown wiring, startup recovery and the lock, planning API, feedback events, and stops that wait for their commit: #53, #74, #57, #59, #60, #96); F2a–F2b (execute prompts, audit, per-item execution, real workspace, durable safety findings: #67, #68, #92, #94, #128); F2d (already-fixed check, PR opening, branch pusher, production publishing, closing PRs on cancel: #84, #101, #110, #115, #116, #119, #123); production wiring and `start`/`resume` (#91, PRs #102, #105); planning through lane D (#117, #124, PRs #118, #120, #125); merging the task's own PR (#121, PR #126); start/resume approval, review-version, recovery and completed-prefix guards (#107, PR #130); continuation after a scope pause (#88, PR #134); and the F3 durable trusted local-rebase foundation (#22, PR #136). Active in #22: foreign-conflict handling, refreshed attribution and approvals, head-bound `cmd:` checks, remote push/check refresh, and guarded merge handoff. | | Hardening | Allowlisted subprocess environments | Done: every Git command (#82, #83) and every `gh` runner (#84, #86). | | Tests | Test suite reliability | Done. The full suite runs ordinary tests first and Docker-backed files serially; CI and the real-Docker gate passed on PR #132 (#129). | | G | Planning screen | Not started. Unblocked: E4 is done and the planning API runs Claude in production. | @@ -418,7 +419,7 @@ These are the decisions that shape the structure. The full list, with the reason ## Known limits -- History must be linear. Merge commits are refused with a request to rebase. +- The trusted local rebase foundation is built, but the merge path does not invoke it yet. Until the rest of #22 lands, the current merge flow still refuses a moved base or non-linear history instead of rebasing automatically. - Linking is line-based. It cannot see an unrelated edit inside a declared file; only the review agent and you can catch that. - The container's guarantees are those of Docker. The agent can always reach its own vendor account. - `cmd:` acceptance checks do not run yet, so any plan item with a `cmd:` check blocks merging.