Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 5 additions & 6 deletions docs/architecture.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# codeboost architecture: a high-level overview

Status: current as of `main` at `1d81303` (2026-10-06). Open work is listed in **Build status and roadmap**.
Status: current as of `main` at `d463cda` (2026-10-07). Open work is listed in **Build status and roadmap**.
Writing standard: plain language, ISO 24495-1:2023

## About this document
Expand Down Expand Up @@ -44,7 +44,7 @@ Writing standard: plain language, ISO 24495-1:2023
- **In use for a non-demo review with a `github` block:** planning. `/api/plan/suggestions` and `/api/plan/drafts` ask Claude, through lane D, for suggestion cards or a whole next plan revision (#117, #124). Nothing becomes a revision until you apply it.
- **In use with an opt-in `runner` block in `review.json`:** the whole loop after planning. `start` and `resume` on `/api/runner` run a task's plan item by item with a Claude agent (#91). They require current plan-item approvals, bind both the task state and review version, recheck approvals before each later item, and refuse when a completed prefix no longer ends at the task's current head (#107). After an out-of-scope pause, an approved amended plan can reconcile the completed prefix and resume only the unfinished suffix (#88, PR #134). When a run ends, codeboost publishes the task's pull request, a draft if the task needs a person (#103). Approve & merge merges that pull request (#121). Cancelling a task closes its pull requests (#111). These are API actions; the screen has no buttons for them yet.
- **Partly built for pre-merge automation:** a trusted local rebase engine records one-to-one commit mappings, verifies the resulting checkout byte for byte, owns its subprocesses durably, and recovers interrupted work (#22, PR #136). It is not wired into the merge path yet.
- **Not built:** the "trust this issue" action and its runner guard (#108), screens for planning and the runner actions, the remaining pre-merge rebase/conflict/check orchestration, running `cmd:` checks, review rounds, and the Planning, Queue and Learning screens.
- **Not built:** screens for planning and the runner actions, the remaining pre-merge rebase/conflict/check orchestration, running `cmd:` checks, review rounds, and the Planning, Queue and Learning screens.

## Terms used

Expand Down Expand Up @@ -357,14 +357,14 @@ The diagram shows the designed transitions. With the `runner` block, running, ne

### Stored data

All state is in one SQLite file, opened with WAL and full synchronization. Each write takes an immediate write lock. The schema version is `PRAGMA user_version`, and migrations run in explicit steps (version 16 today); an unknown version fails.
All state is in one SQLite file, opened with WAL and full synchronization. Each write takes an immediate write lock. The schema version is `PRAGMA user_version`, and migrations run in explicit steps (version 17 today); an unknown version fails.

| Group | Tables | Notes |
|---|---|---|
| Plans | `plans`, `revisions`, `requests` | SQLite allocates revision numbers. Old revisions are never changed. Suggestion and draft requests (`mode`, #124) are bound to a revision and snapshot; continuation requests also retain their checkpoint, audited head and completed prefix. |
| Code history | `snapshots`, `ledger`, `rewrites` | Ledger entries are immutable. Rebase mappings record which old commit became which new one; foreign stays foreign. |
| Review | `approvals`, `choices`, `review_notes`, `checkpoints`, `continuations` | Stored approvals are claims about a past snapshot. Freshness is recomputed every time. |
| Runner | `tasks`, `attempts`, `user_actions`, `feedback_events`, `merge_attempts`, `app_settings` | User actions carry idempotency keys. Feedback events are append-only and feed the future learning feature. An attempt carries its safety finding and diagnostic reference. `app_settings` holds the runner, Ask and planning owner tokens. |
| Runner | `tasks`, `attempts`, `issue_trust`, `user_actions`, `feedback_events`, `merge_attempts`, `app_settings` | User actions carry idempotency keys. Feedback events are append-only and feed the future learning feature. An attempt carries its safety finding, diagnostic reference and the issue-comment evidence prepared for its prompt. `issue_trust` records trust and revocation for each repository issue. `app_settings` holds the runner, Ask and planning owner tokens. |
| Publishing | `task_pull_requests`, `already_fixed_checks`, `publish_outcomes` | An opening is recorded before the GitHub call, so a lost outcome is recovered, not repeated. Each task's last publish outcome is stamped with the state version that publish last saw (#114), so a later change still makes a new publish owed. |

## Deployment view
Expand Down Expand Up @@ -414,7 +414,7 @@ These are the decisions that shape the structure. The full list, with the reason
| Hardening | Allowlisted subprocess environments | Done: every Git command (#82, #83) and every `gh` runner (#84, #86). |
| Tests | Test suite reliability | Done. The full suite runs ordinary tests first and Docker-backed files serially; CI and the real-Docker gate passed on PR #132 (#129). |
| G | Planning screen | Not started. Unblocked: E4 is done and the planning API runs Claude in production. |
| H | Issue ranking and Issues screen | Done except H4b (#108): the "trust this issue" action, plus a runner guard that refuses issues from non-collaborators that nobody trusted. It is unblocked now that #130 has merged. |
| H | Issue ranking and Issues screen | Done. H4b (#108, PR #139) records author-bound trust, guards runner and publishing boundaries, controls which comments reach prompts, records prompt comment evidence, and provides trust/untrust controls on the Issues screen. |
| I, J | Queue and run windows; learning from feedback | Not started (after F) |

## Known limits
Expand All @@ -425,7 +425,6 @@ These are the decisions that shape the structure. The full list, with the reason
- `cmd:` acceptance checks do not run yet, so any plan item with a `cmd:` check blocks merging.
- One user, one machine, one runner per database. The lock refuses a second runner, and it refuses network filesystems, where OS locks are unreliable.
- The runner is opt-in: it needs a `runner` block and a `github` block in `review.json`, and it is always off in the demo. It runs through the API only; the screen has no start, resume, publish or planning controls yet. Publishing also needs `github.baseBranch`.
- The runner does not check who wrote an issue yet. `start` and `resume` run a plan whose issue came from a non-collaborator (#130 adds an approval check on the plan, not on the issue author), and no attempt records which comments its prompt carried. Only collaborators' comments reach the prompt. #108 adds the guard and the record. Until then, start runs only for issues you have checked yourself.
- codeboost is Claude-only. Codex is refused in every phase (#93, PR #106), because it can read files only through its shell, which every phase turns off. `agent-isolation.md` says when to revisit this.

## Test this document with a reader
Expand Down
4 changes: 2 additions & 2 deletions docs/designs/codeboost-plan-indexed-review.md
Original file line number Diff line number Diff line change
Expand Up @@ -569,7 +569,7 @@ The numbers below identify delivery milestones, not a requirement to implement t
5. **Running agents.** Per-task clones, containers, agent adapters, permissions, one invocation per plan item, review rounds, the "already fixed" check, and opening PRs (step 6). Isolation boundary merged (lane D, PRs #31, #40, #44, #47, #50); the runner (lane F) is next.
6. **Planning screen.** Writing plans with an agent, and approving plan changes (steps 2 and 3). Backend in progress (lane E); no screen yet.
7. **Queue, schedule, and recovery** (steps 4 and 5).
8. **Issue list, sorted by how critical each issue is** (step 1). Ranking backend (H1–H3) and the Issues screen (H4a, #55) merged; the "trust this issue" action (H4b) remains.
8. **Issue list, sorted by how critical each issue is** (step 1). *Done:* ranking backend (H1–H3), the Issues screen (H4a, #55), and the author-bound "trust this issue" action and runner guard (H4b, #108, PR #139) have merged.
9. **Learning from your feedback** (step 10): lessons, the Lessons inbox, and the Learning screen. It needs the reject loop from steps 4 to 7.

### Build step 4: scope and progress
Expand Down Expand Up @@ -1952,7 +1952,7 @@ Read each row left to right: finish and validate step 1 before step 2 within tha
|---|---|---|---|
| F — runner and pre-merge automation | C and D merged; invocation contract available | Build step 5 runner plus #22 / remaining build step 4; T3, T6, T11. Own `runner/`, rebase helpers and shared acceptance persistence during this wave. | Preserve ledger attribution through rebase; recompute approval staleness; execute and persist head-bound checks; cover timeout, cancellation, shutdown and collaborator-push races; complete #22 acceptance. |
| G — planning screen | E merged; C releases shared UI files | Build step 6 UI and T18 integration. Own `web/` and dedicated browser tests during this wave. Route persistence changes through F. UI work can use controlled provider fixtures until D is available. | Import, generation and Apply preserve user drafts and attachments and reject stale/replayed suggestions. Final completion requires real D-backed invocation and integration with F/store, not fixtures alone. |
| H — issue prioritization | Done: access contract inspected and ranking policy recorded in `docs/implementation/issue-prioritization.md` (H1); H2–H3 merged (#39, #42); H4a Issues screen merged (#55) | Build step 8: issue-fetch/normalization and ranking modules with dedicated tests. Shared shell/navigation integration waits for G. | Stable ranking with a visible reason per issue; unavailable/stale data has explicit states. Ranking policy decided in H1; Issues screen merged in #55 (H4a); H4b trust action remains. |
| H — issue prioritization | Done: access contract and ranking policy (H1); retrieval/ranking (#39, #42); Issues screen (#55); author-bound trust, runner/publish guards, prompt comment evidence, and trust controls (#108, PR #139) | Build step 8: issue retrieval, normalization, ranking, and trust enforcement with dedicated unit, integration, and browser tests. | Done. Stable ranking has visible reasons and explicit unavailable/stale states; H4b persists trust, fails closed at runner and publishing boundaries, controls comment inclusion, records delivered comment evidence, and works in demo mode. |

F, G and H can proceed together within these ownership boundaries. If F and G need an incompatible shared storage/API change, land that small prerequisite first; neither edits the other's files in parallel. Merge independent backend modules first, then their shared integration, and rerun checks on the combined head.

Expand Down
Loading