Skip to content

Recover from temporary Harness service failures within a bounded retry budget #584

Description

@JacobStephens2

This was generated by AI during triage.

Problem

The Spec run on #427 completed 13 of 17 Tickets but failed when #544's conflict Repair hit a temporary service failure. Ticket PR #582 had passed CI, then needed another merge because the Spec branch moved. During Repair 1, Codex reported repeated HTTP 503 errors, fell back from WebSockets to HTTPS, and ended with “Selected model is at capacity. Please try a different model.” thirdshift propagated the session failure and ended the Ticket's Run.

Existing Resume behavior handles a session ending with background work still running. It does not recover from capacity failures. Failed work is already preserved, and Continuations already pick it up; this proposal builds on those contracts.

Proposed behavior

  • Recognize a narrow set of temporary Harness service failures, including the observed capacity failure, and retry with increasing delays.
  • Give service recovery its own attempt and elapsed-time budget, separate from the Repair cap.
  • Preserve the checkout, edits, and logs across attempts. Continue the recorded session where the Harness supports it; any recovery without session continuation must account for existing work and partially completed operations.
  • Keep the chosen Harness and Model by default. Any Model fallback is an explicit operator configuration.
  • Make recovery attempts and budget exhaustion visible in progress, logs, and the final notification.
  • When the budget is exhausted, take the existing Failed run preservation path with the underlying cause.

Decisions needed

  • Which structured Harness errors qualify, and what narrowly defined fallback classification is needed where structured codes are unavailable?
  • What are the default attempt, delay, and total-time limits?
  • What is the recovery contract when a conversation cannot be resumed or an operation's completion is uncertain?
  • Should optional Model fallback be included here or handled separately?

Validation targets

  • A temporary failure followed by recovery can finish the same logical session and Delivery without losing prior work.
  • Repeated temporary failures stop within the configured budget and retain the causal error and each attempt's log.
  • Recovery consumes no additional Repair slots.
  • Command interruption cancels waits and prevents further attempts.
  • Refusals, invalid configuration, and ordinary implementation failures do not qualify for temporary-service retries.
  • Retrying does not blindly repeat external mutations or treat an incomplete session as successful.

Scope

Recovery within an already authorized Command. No new daemon, automatic retries of failed issues on later scheduled passes, or changes to review and merge requirements.

Evidence

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestneeds-triageMaintainer needs to evaluate this issue

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions