This was generated by AI during triage.
Problem
The Spec run on #427 completed 13 of 17 Tickets but failed when #544's conflict Repair hit a temporary service failure. Ticket PR #582 had passed CI, then needed another merge because the Spec branch moved. During Repair 1, Codex reported repeated HTTP 503 errors, fell back from WebSockets to HTTPS, and ended with “Selected model is at capacity. Please try a different model.” thirdshift propagated the session failure and ended the Ticket's Run.
Existing Resume behavior handles a session ending with background work still running. It does not recover from capacity failures. Failed work is already preserved, and Continuations already pick it up; this proposal builds on those contracts.
Proposed behavior
- Recognize a narrow set of temporary Harness service failures, including the observed capacity failure, and retry with increasing delays.
- Give service recovery its own attempt and elapsed-time budget, separate from the Repair cap.
- Preserve the checkout, edits, and logs across attempts. Continue the recorded session where the Harness supports it; any recovery without session continuation must account for existing work and partially completed operations.
- Keep the chosen Harness and Model by default. Any Model fallback is an explicit operator configuration.
- Make recovery attempts and budget exhaustion visible in progress, logs, and the final notification.
- When the budget is exhausted, take the existing Failed run preservation path with the underlying cause.
Decisions needed
- Which structured Harness errors qualify, and what narrowly defined fallback classification is needed where structured codes are unavailable?
- What are the default attempt, delay, and total-time limits?
- What is the recovery contract when a conversation cannot be resumed or an operation's completion is uncertain?
- Should optional Model fallback be included here or handled separately?
Validation targets
- A temporary failure followed by recovery can finish the same logical session and Delivery without losing prior work.
- Repeated temporary failures stop within the configured budget and retain the causal error and each attempt's log.
- Recovery consumes no additional Repair slots.
- Command interruption cancels waits and prevents further attempts.
- Refusals, invalid configuration, and ordinary implementation failures do not qualify for temporary-service retries.
- Retrying does not blindly repeat external mutations or treat an incomplete session as successful.
Scope
Recovery within an already authorized Command. No new daemon, automatic retries of failed issues on later scheduled passes, or changes to review and merge requirements.
Evidence
Problem
The Spec run on #427 completed 13 of 17 Tickets but failed when #544's conflict Repair hit a temporary service failure. Ticket PR #582 had passed CI, then needed another merge because the Spec branch moved. During Repair 1, Codex reported repeated HTTP 503 errors, fell back from WebSockets to HTTPS, and ended with “Selected model is at capacity. Please try a different model.” thirdshift propagated the session failure and ended the Ticket's Run.
Existing Resume behavior handles a session ending with background work still running. It does not recover from capacity failures. Failed work is already preserved, and Continuations already pick it up; this proposal builds on those contracts.
Proposed behavior
Decisions needed
Validation targets
Scope
Recovery within an already authorized Command. No new daemon, automatic retries of failed issues on later scheduled passes, or changes to review and merge requirements.
Evidence
/home/jacob/.thirdshift/logs/JacobStephens2/thirdshift/commands/issue/427-20261007T231713-0400.log, lines 3072–3130./home/jacob/.thirdshift/logs/JacobStephens2/thirdshift/sessions/544-20261007T231713-0400-repair-1.jsonl, lines 97–103.