Conversation
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
🔍 Devin Review: 1 flag
Not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
EvalRunResult.passedwas true only when no row failed, errored, or was pending. It ignoredpass_rate_threshold. A run with 28 passed, 1 failed, and 1 error rows (93%) returnedpassed=Falsewhen each criterion setpass_rate_threshold=0.9.Change
_run_passedinmodule.py.pass_rate_threshold, the run passes when no row is pending andpassed_rows / total_rowsis at least the highest threshold. Errored rows count as not passed.pass_rate_threshold, the strict rule stays. Any failed, errored, or pending row fails the run.Limit
The summary endpoint returns row counts for the whole run. It does not return counts for each criterion. The check uses the highest threshold on the whole-run pass rate. A per-criterion check needs per-criterion counts from the API.
Test
Add a parametrized test for
_run_passed. The full client suite, ruff, and mypy pass.🤖 Generated with Claude Code
Note
Overview
EvalRunResult.passednow respectspass_rate_thresholdon evaluation criteria instead of requiring zero failed rows.A new
_run_passedhelper drives the gate: when every criterion sets a pass rate, the run passes once nothing is pending andpassed_rows / total_rowsmeets the highest threshold (failed and error rows lower the rate). If any criterion omits a pass rate, behavior stays strict—any failed, errored, or pending row fails the run. Empty runs never pass under pass-rate mode.Parametrized unit tests cover mixed thresholds, partial pass-rate configs, and edge cases. Package versions in
uv.lockbump to 0.2.4 (and related packages).Reviewed by Cursor Bugbot for commit ea171a9. Bugbot is set up for automated code reviews on this repo. Configure here.