Skip to content

refactor(tooling): adopt shared benchmark primitives - #269

Merged
acgetchell merged 2 commits into
mainfrom
refactor/268-shared-benchmark-primitives
Oct 2, 2026
Merged

acgetchell merged 2 commits into
mainfrom
refactor/268-shared-benchmark-primitives

Conversation

@acgetchell

@acgetchell acgetchell commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Summary

Benchmark support scripts duplicate primitives already available in the pinned research-repo-tools 0.1.7 package. Adopt those primitives while keeping la-stack's scientific eligibility, historical evidence formats, and report layout local.

  • Use shared Criterion parsing and comparison validation, SHA-256 verification, safe archive extraction, README marker replacement, and file transactions. Render and validate complete output groups before publishing them; preserve existing artifacts on failure.
  • Keep legacy CSV/JSON framing and hashes, signed timing changes, complete confidence intervals, and the 100-sample/95%-interval measurement policy. Reject numeric strings and booleans in raw timing fields and report invalid comparisons with useful context.
  • Remove superseded parser/publication code and dead helpers. Remove two Semgrep rules already covered by Ruff plus the obsolete Git-input implementation rule; update remaining subprocess rules and fixtures. Refresh Ruff to 0.16.10. Keep Semgrep at 1.178.0 for its Windows wheel and disable source builds that omit the native engine.
  • Add an OSV exception only for RUSTSEC-2024-0436, expiring January 1, 2027. This accepts the documented maintenance risk of benchmark-only paste through faer/gemm/pulp; it does not patch the dependency or disable other advisories.

The remaining common-harness and immutable-run workflow migration stays open in #268, pending publication of acgetchell/research-repo-tools#64. The shared scanner's missing terminal diagnostics are tracked in acgetchell/research-repo-tools#65, with a live reproduction added there.

Validation

  • just ci passed during the Python review: 499 Python tests and 893 runnable Rust tests, plus the configured doctest, example, and benchmark-build checks.
  • just check and just security passed after the Semgrep cleanup and security exception. OSV scanned 165 Rust and 81 Python packages with no remaining findings; Gitleaks found no secrets in reachable history or the working inventory.
  • Wheel/sdist build and isolated installed-command checks passed. The retained 225-row evidence round trip preserved the CSV, provenance, and rendered Markdown hashes.
  • After correcting the Semgrep pin, just check passed again. Locked dependency selection for Windows succeeded in a dry run; a separate Windows dry run confirmed the wheel-only policy rejects Semgrep 1.179.0, which lacks a Windows wheel.
  • Local validation ran on macOS arm64 with Python 3.14. Native CI on the corrected commit passed the full check suite on Linux, macOS, and Windows.

Performance

No Rust numerical kernels or benchmark definitions changed, and no timing improvement is claimed. This reduces duplicated support infrastructure while retaining the existing scientific workflow contract.

- Delegate Criterion parsing and comparisons, digest verification, archive
  extraction, and README section replacement to research-repo-tools.
- Publish complete report, evidence, plot, and retained-summary output groups
  through shared transactions while preserving historical schemas and hashes.
- Report invalid timings and publication failures with context, preserving
  existing artifacts when an operation fails.
- Remove duplicate tooling and Semgrep rules, refresh subprocess policy,
  and update the Ruff and Semgrep pins.
- Accept the benchmark-only paste maintenance advisory until January 1, 2027,
  while retaining checks for all other advisories.

Refs #268
Refs acgetchell/research-repo-tools#64
@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: acgetchell/la-stack/.coderabbit.yaml

Review profile: CHILL

Plan: Essentials

Run ID: 7e730f80-8d7b-40a6-bad8-c22989678793

📥 Commits

Reviewing files that changed from the base of the PR and between 1c9c66d and 0e74cc3.

⛔ Files ignored due to path filters (1)
  • uv.lock is excluded by !**/*.lock
📒 Files selected for processing (1)
  • pyproject.toml

Included review availability: This review used your included allowance. 1 included review remains after this review. Your included PR review attempts over the past 7 days set your current allowance at 3 reviews per hour.


📝 Walkthrough

Walkthrough

Benchmark scripts adopt shared utilities for Criterion data, archive extraction, digest checks, document sections, and multi-file publication. Documentation and tests describe these changes and retained consumer-owned behavior. Security files record a temporary OSV exception for RUSTSEC-2024-0436.

Changes

Performance tooling consolidation

Layer / File(s) Summary
Criterion validation and ownership
docs/code_organization.md, scripts/README.md, scripts/bench_compare.py, scripts/benchmark_summaries.py, scripts/criterion_measurements.py, scripts/performance_artifacts.py, scripts/tests/test_bench_compare.py, scripts/tests/test_performance_artifacts.py, scripts/tests/test_release_baseline.py
Benchmark scripts delegate estimate parsing, timing validation, comparison arithmetic, and digest checks to shared utilities. Documentation describes shared and consumer-owned performance behavior. Tests cover validation context, comparison output, and estimate requirements.
Benchmark artifact and report publication
scripts/bench_compare.py, scripts/criterion_dim_plot.py, scripts/performance_artifacts.py, scripts/tests/test_bench_compare.py, scripts/tests/test_criterion_dim_plot.py, scripts/tests/test_performance_artifacts.py, pyproject.toml, semgrep.yaml
Benchmark and plotting scripts prepare candidate outputs and publish them through shared replacement utilities. README updates use byte-based section replacement. Tests cover output validation, README byte preservation, and publication failures. Ruff settings, the Semgrep build configuration, and static-analysis rules also change.
Archive extraction and report promotion
scripts/archive_performance.py, scripts/tests/test_archive_performance.py
Archive processing uses shared extraction and replacement utilities. Generated artifacts, summaries, reports, and archive indexes are assembled for multi-file publication. Tests simulate replacement failures.

Temporary OSV exception

Layer / File(s) Summary
Advisory exception and review policy
SECURITY.md, osv-scanner.toml, docs/code_organization.md
Security guidance records the checked dependency date, exception period, workspace dependency-tree command, and removal and reassessment steps. OSV configuration adds an exception for RUSTSEC-2024-0436 through January 1, 2027.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Refactor

Merge Risk: ⚪ Minimal · up to 0e74c

The reviewed tooling and configuration changes show no supported workflow regression, so no actionable merge risk is identified in the supplied context.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: adopting shared benchmark primitives in tooling.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.02%. Comparing base (b570d0c) to head (0e74cc3).
⚠️ Report is 1 commits behind head on main.
✅ All tests successful. No failed tests found.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #269   +/-   ##
=======================================
  Coverage   98.02%   98.02%           
=======================================
  Files          13       13           
  Lines        6726     6726           
=======================================
  Hits         6593     6593           
  Misses        133      133           
Flag Coverage Δ
unittests 98.02% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 2, 2026
- Restore Semgrep 1.178.0, whose Windows wheel includes semgrep-core.exe.
- Disable Semgrep source builds so missing platform wheels fail during
  installation instead of leaving an unusable scanner in the environment.

Refs #269
@acgetchell
acgetchell marked this pull request as ready for review October 2, 2026 05:17
@acgetchell
acgetchell enabled auto-merge October 2, 2026 05:17
@acgetchell
acgetchell merged commit 66679d3 into main Oct 2, 2026
20 checks passed
@acgetchell
acgetchell deleted the refactor/268-shared-benchmark-primitives branch October 2, 2026 05:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant