Skip to content

Fail Security sessions on safeguard refusals and log answering Models - #577

Merged
JacobStephens2 merged 2 commits into
issue-427from
issue-547
Oct 8, 2026
Merged

JacobStephens2 merged 2 commits into
issue-427from
issue-547

Conversation

@JacobStephens2

@JacobStephens2 JacobStephens2 commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Summary

Security sessions now fail with a named cause when Claude Code or Codex refuses cybersecurity work. Model evidence reaches progress and the Command log across all six supported Harnesses, including Model switches and resumed sessions.

Security session
├─ safeguard refusal → failed Run → fixed refusal cause in notification
└─ Model evidence    → session progress → Command log

Claude's synthetic errors and sub-agent Models are excluded. Codex reports the requested Model explicitly; retained Muse/OpenCode records supply their answering Models. Raw refusal diagnostics remain in local records.

Closes #547

Evidence

  • Before: a_claude_cyber_refusal_fails_the_security_audit_with_its_cause failed with a generic missing-final-line cause; refusal notifications omitted the cause.
    After: Both refusal Harnesses fail with their named safeguard cause, and notification regressions verify that private diagnostics stay out of email.
  • Before: The Spec review's exact commands failed (exit 101):
    cargo test --test security_run claude_security_progress_ignores_synthetic_api_errors -- --exact --nocapture
    cargo test --test security_run grok_security_audits_log_the_model_that_answered -- --exact --nocapture
    After: Both exact commands pass (exit 0), and both regressions remain in the suite. Additional red/green regressions cover Antigravity, Muse and OpenCode Model evidence.
  • Validation: Full cargo test passed: 1,575 tests passed, one ignored, zero failed across 40 suites. cargo check --all-targets, cargo fmt --check, and cargo clippy --all-targets -- -D warnings also passed.

Testing seams

  • End-to-end scenario harness: the real binary and Git with scripted Harness, GitHub and Resend adapters; observe Security run outcomes, progress, Command logs, private records and notifications.
  • Harness interpretation interface: feed stream events and observe progress/completion to cover Security policy independently of subprocesses.

Review

Reviewed against the captured issue-427 tip, f397e261e4649a21915e9d3f7efaf930fb0949ac, with one coverage follow-up per axis. Standards: three P3 maintainability findings addressed. Spec: two P2 behavioral findings reproduced, fixed and retained as regressions. Both reviewers read all 14 changed files and hunks. Changed files left unread: none.

Unaddressed findings

Standards

None. All three findings are addressed.

Spec

None. Both findings are addressed; their exact reproductions failed before the fixes and pass afterward.

Merge Danger

Door: two-way

Reversible changes to session interpretation, logging and notification rendering; no migration.

Blast Radius: Sessions

Security-purpose sessions gain refusal classification and Model reporting. Ordinary-session interpretation remains covered by regression tests.

Built with codex · gpt-6.1-sol · xhigh

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant