Skip to content

feat: Bytes partial decode - #4458

Open
bendichter wants to merge 10 commits into
zarr-developers:mainfrom
bendichter:bytes-partial-decode
Open

bendichter wants to merge 10 commits into
zarr-developers:mainfrom
bendichter:bytes-partial-decode

Conversation

@bendichter

@bendichter bendichter commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

I am working with an application where data is stored uncompressed, with BytesCodec as the only codec, and I want to read small subregions of large chunks. Currently, zarr-python reads the entire chunk for every selection. That is necessary when a chunk is compressed, but not when BytesCodec is the only codec. This PR makes BytesCodec implement partial decoding, so it fetches only the range that corresponds to the rows of interest. Selections that touch every row, and chunks with any other codec, are read whole as before.

This helps any uncompressed array with large chunks, and especially virtual datasets, where the chunk layout comes from existing files. A contiguous HDF5 dataset, for example, becomes a single chunk that can be many GB. In an example I am working on, reading 100 samples from a 20 MB single-chunk array fetched 20 MB before this change and 2 KB after.

Author attestation

  • I am a human, these are my changes, and I have reviewed and understood every change and can explain why each is correct.

TODO

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.md
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

bendichter and others added 5 commits September 30, 2026 01:25
BytesCodec now implements partial decoding. An uncompressed chunk is
stored in C order, so each row along its first axis is a contiguous run
of bytes; for a selection that does not touch every row, the rows from
the first to the last one it touches are fetched with a single range
request and the selection is applied to them. The pipeline already
uses partial decoding when the array-to-bytes codec supports it and
there are no array-to-array or bytes-to-bytes codecs, so this applies to
uncompressed arrays only. Reading 100 rows of a single-chunk 20 MB
array now fetches the bytes of those rows instead of the whole chunk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… reads

A boolean mask on the first axis, as from arr.oindex[mask], was trimmed to
start at the first selected row but not to end at the last, so its length
no longer matched the rows fetched and numpy raised an IndexError. The row
window and the shifted selection are now computed together in _row_window,
which trims the mask at both ends. This also gives mypy the narrowing it
needs.

FsspecStore over HTTP returns the whole object when a server ignores the
Range header. The partial read then failed to reshape a chunk that the full
read would have decoded, so a read that worked before this branch raised an
error. A response longer than the requested range is now sliced to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the needs release notes Automatically applied to PRs which haven't added release notes label Sep 30, 2026
@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.73%. Comparing base (bd0dc84) to head (c84c41b).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4458      +/-   ##
==========================================
+ Coverage   94.71%   94.73%   +0.02%     
==========================================
  Files          94       94              
  Lines       13663    13726      +63     
==========================================
+ Hits        12941    13004      +63     
  Misses        722      722              
Files with missing lines Coverage Δ
src/zarr/codecs/bytes.py 99.31% <100.00%> (+0.52%) ⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Comment thread src/zarr/codecs/bytes.py Outdated
Comment on lines +174 to +177
if len(chunk_bytes) > (stop - first) * row_bytes:
# The store sent the whole chunk, as an HTTP server that ignores
# the Range header does.
chunk_bytes = chunk_bytes[first * row_bytes : stop * row_bytes]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this needs a test. if the store honored the byte range start but not the stop, then re-indexing from the start again will generate an invalid result.

because byte range handling is so important, we should probably set up stores in our test fixtures that span the range of byte range handling behavior, to ensure that branches like this get tested thoroughly

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, this was a bug. It's now fixed in 9c2b182.

@d-v-b

d-v-b commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

thank you for this! would you mind implementing the same behavior for a synchronous method (_decode_partial_sync)? This would allow the FusedCodecPipeline to use this feature.

@d-v-b d-v-b added the benchmark Code will be benchmarked in a CI job. label Oct 1, 2026
@codspeed

codspeed Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Merging this PR will improve performance by 23.42%

⚡ 24 improved benchmarks
✅ 89 untouched benchmarks
⏩ 37 skipped benchmarks1

Performance Changes

Benchmark BASE HEAD Efficiency
⚡ test_sharded_morton_indexing_large[(33, 33, 33)-memory] 10 s 7.3 s +35.95%
⚡ test_sharded_morton_indexing_large[(32, 32, 32)-memory] 9.1 s 6.7 s +35.58%
⚡ test_sharded_morton_indexing[(32, 32, 32)-memory] 1,140.6 ms 842.7 ms +35.36%
⚡ test_sharded_morton_indexing_large[(30, 30, 30)-memory] 7.5 s 5.5 s +35.05%
⚡ test_sharded_morton_indexing[(16, 16, 16)-memory] 143.3 ms 106.3 ms +34.84%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory] 351.2 ms 261.7 ms +34.21%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory_get_latency] 405.6 ms 314.1 ms +29.14%
⚡ test_slice_indexing[(50, 50, 50)-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory] 400.9 ms 312.8 ms +28.17%
⚡ test_slice_indexing[(50, 50, 50)-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory_get_latency] 403.7 ms 315.8 ms +27.81%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(0, 3, 2), slice(0, 10, None))-memory] 3.4 ms 2.7 ms +24.43%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-semi_random] 4.3 s 3.5 s +23.18%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(0, 3, 2), slice(0, 10, None))-memory_get_latency] 3.9 ms 3.2 ms +22.8%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-repeated] 4.3 s 3.5 s +21.81%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-repeated] 4.4 s 3.6 s +21.68%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(100000000,), chunks=(100000,), shards=(100000000,))-None-repeated] 504.7 ms 417.2 ms +20.96%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-semi_random] 4.4 s 3.7 s +20.33%
⚡ test_slice_indexing[None-(slice(None, 10, None), slice(None, 10, None), slice(None, 10, None))-memory] 847.5 µs 722 µs +17.38%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(100000000,), chunks=(100000,), shards=(100000000,))-None-repeated] 598 ms 517.2 ms +15.64%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(100000000,), chunks=(100000,), shards=(100000000,))-None-semi_random] 598 ms 518.3 ms +15.37%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(100000000,), chunks=(100000,), shards=None)-None-repeated] 658.2 ms 572.9 ms +14.9%
... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing bendichter:bytes-partial-decode (7687617) with main (25d390b)2

Open in CodSpeed

Footnotes

  1. 37 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (b996f03) during the generation of this report, so 25d390b was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

bendichter and others added 2 commits October 4, 2026 11:58
…reads

Add _decode_partial_sync so FusedCodecPipeline reads only the needed rows
of an uncompressed chunk, sharing the decode step with the async method.

Decide what a store sent for a range request by its exact length: the
requested range, the whole chunk, or everything from the range start.
Any other length raises instead of returning wrong data.

Run the partial read tests on both codec pipelines and add a changelog
entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot removed the needs release notes Automatically applied to PRs which haven't added release notes label Oct 4, 2026
@bendichter

Copy link
Copy Markdown
Contributor Author

Done in 9c2b182. BytesCodec now has _decode_partial_sync, so FusedCodecPipeline takes the partial path too. The range calculation and the decode step are shared. The partial-read tests now run on both pipelines, including the byte-count assertions.

@d-v-b

d-v-b commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

this is great, thank you. one wrinkle that needs to be fixed in documentation: prior to this change, broken chunks (e.g., truncated due to an aborted write) would cause an error at indexing time. now, array indexing will not necessarily error on a broken chunk, if the region requested is locally valid. That's fine, but we should document this. I'll handle that in a follow-up PR.

@d-v-b d-v-b added this to the 3.5.0 milestone Oct 9, 2026
@d-v-b d-v-b added benchmark Code will be benchmarked in a CI job. and removed benchmark Code will be benchmarked in a CI job. labels Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark Code will be benchmarked in a CI job.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants