Goal
Make decision-pack quality measurable and claims reproducible.
Scope
- Labeled fixtures and deterministic evaluation runner.
- Accuracy, confusion counts, per-class metrics, Brier score, and calibration/ECE where probabilities are available.
- Separate mock-provider tests from optional live-Laya evaluation.
- Machine-readable and human-readable reports.
Acceptance criteria
- Dev and held-out fixtures are clearly separated.
- Reports include model/checkpoint, configuration, sample size, hardware, and methodology.
- Loading, routing, network, and inference latency are distinguished.
- Negative results and known limitations are retained.
- CI prevents unsupported benchmark claims from being presented as verified.
Dependency
Decision-pack pull requests should add fixtures compatible with this harness.
Goal
Make decision-pack quality measurable and claims reproducible.
Scope
Acceptance criteria
Dependency
Decision-pack pull requests should add fixtures compatible with this harness.