Skip to content

Fail a non-UTF-8 input cleanly instead of crashing the expanded run - #27

Merged
ecrum19 merged 1 commit into
mainfrom
claude/optimistic-elgamal-747851
Sep 25, 2026
Merged

ecrum19 merged 1 commit into
mainfrom
claude/optimistic-elgamal-747851

Conversation

@ecrum19

@ecrum19 ecrum19 commented Sep 25, 2026

Copy link
Copy Markdown
Owner

Problem

On vcf-bench-1 with the v3.1.0 image, a directory input containing macOS AppleDouble sidecars (._P001.vcf etc.) crashed --mode full --sample-representation expanded with a traceback:

File "vcf_rdfizer.py", line 2276, in __enter__
    self._header = next(self._reader, None) or []
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xa3 in position 130: invalid start byte

cohort_scale_refusal → read_records_tsv_sample_count caught OSError/csv.Error/StopIteration but not UnicodeDecodeError, and the guard is called outside any per-input try. Condensed never runs the guard, so it failed cleanly later.

Fix

  • Encoding probe: check_records_tsv_is_utf8 decodes the first 1 MiB of each records TSV and raises InputEncodingError. It runs before the guard for every representation, and the call site reports it through fail_current at stage input-encoding: not a UTF-8 text VCF (byte 0xa3 does not decode as UTF-8; …). The other inputs still convert; the run exits 1 with failed_inputs.csv.
  • Guard hardening: read_records_tsv_sample_count turns UnicodeDecodeError into InputEncodingError instead of leaking it.
  • AppleDouble skipping: list_vcfs_in_dir and src/vcf_as_tsv.sh skip ._* files; resolve_input_snapshot prints one notice naming them. Documented in docs/cli-reference.md.

Tests

In test/test_cohort_guard_unit.py:

  • The guard's reader raises InputEncodingError, not a decode error; a UTF-8 character split at the probe boundary is not flagged.
  • End to end through main() (Docker mocked): a directory with P001.vcf + ._P001.vcf (expanded) skips the sidecar and converts P001 (exit 0); a directory with P001.vcf + a binary x.vcf (expanded and condensed) fails x at input-encoding and converts P001 (exit 1).

Against the old vcf_rdfizer.py the new tests reproduce the vcf-bench-1 traceback exactly. Full unit suite: 906 tests OK (25 skipped).

Limits

  • The probe reads only the first 1 MiB. A text VCF with a stray bad byte deeper in still fails later with the raw codec message.
  • Not yet re-run on vcf-bench-1; that needs an image built from this branch.

🤖 Generated with Claude Code

The cohort-scale guard read the records TSV outside the per-input error
handling and did not catch UnicodeDecodeError, so a binary input (macOS
AppleDouble ._P001.vcf sidecars on vcf-bench-1, v3.1.0) killed a whole
--sample-representation expanded run with a traceback.

- Probe the first 1 MiB of each records TSV for UTF-8 before the guard, for
  every representation, and report failures as stage input-encoding
  ("not a UTF-8 text VCF"); the other inputs still convert.
- The guard's reader raises InputEncodingError instead of leaking the
  decode error.
- Skip ._* AppleDouble sidecars during directory enumeration (wrapper and
  vcf_as_tsv.sh), with a notice; documented in docs/cli-reference.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@ecrum19
ecrum19 merged commit ac2748e into main Sep 25, 2026
24 checks passed
@ecrum19
ecrum19 deleted the claude/optimistic-elgamal-747851 branch September 25, 2026 12:17
@ecrum19 ecrum19 mentioned this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants