Skip to content

feat(zarr-metadata)!: read metadata against definitions in a scope, and make every model a pair of document and scope - #4490

Open
d-v-b wants to merge 88 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-metadata-document-context-stack
Open

d-v-b wants to merge 88 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-metadata-document-context-stack

Conversation

@d-v-b

@d-v-b d-v-b commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

This is a big one. It adds a runtime type checker with some fancy bells and whistles for parsing zarr metadata documents correctly in different contexts.

why not use pydantic

because we don't want the dependency on pydantic. that invites conflicts with the packages that might use zarr-metadata with their own range supported pydantic versions. otherwise we totally would use pydantic-core, because I think it handles typeddicts correctly and it has a lot of eyes.

why is it so complicated

because parsing zarr metadata documents is complicated. I'm open to simplifications but if we want:

  • useful, comprehensive error messages
  • recursive parsing of recursive data types / codecs
  • semantics of interactions between fields (like chunk grid and codecs)
  • explicit scope for metadata types (core, vs extensions)
  • utilities for repairing broken metadata (zarr has created its fair share!)

then we have to pay for some complexity.

I will self-merge this when I'm happy with it. in parallel, I will open a draft PR that contains the integration of these new zarr-metadata features in zarr. The promise is much more correctness from zarr. the cost is (eventual) breakage in the metadata classes zarr uses today.

🤖 AI text below 🤖

Every metadata field is now read against a definition in a scope, and every model is the pair of its document and the scope it was read in. A field of a document (a v3 data type, chunk grid, chunk key encoding, codec or storage transformer; a v2 dtype or numcodecs codec) is read as Read by the definition in scope that claims its name, Unclaimed when nothing does, or Refused with every problem found. read_array_metadata_v3, read_group_metadata_v3 and read_array_metadata_v2 read a document once and return everything the read found, the model among it.

A model is built only from a document in a scope (CORE_AND_EXTENSIONS or CORE_V2 by default) and is never invalid. Two models are equal when their documents mean the same in their scopes. update reads new members in the model's own scope, with_context and refined_in read the document in another, and refines orders models by information. Scopes compare, join and report disagreements. JSON Schemas are exported for a TypedDict, for a field in a scope and for a zarr.json. Known writer bugs are undone below the strict model by read_repaired_node_metadata_v3 and read_repaired_consolidated_metadata_v2. The core depends on typing-extensions and annotated-types only; the README says why there is no validation framework, and the public surface is cut to what a consumer needs.

Breaking: models are no longer dataclasses built from typed members; construct, the ...Partial TypedDicts, ZarrV3NamedConfig/ZarrV3MetadataField and the required context on update are gone; a structured v2 dtype with fill_value: 0 is refused. The changelog fragments under packages/zarr-metadata/changes/ describe every change for users.

History

This branch is the zarr-metadata work from fork PRs d-v-b#370 through d-v-b#398, squashed into one pull request with upstream main merged in. Each fork PR carries its own description and review history. The changelog fragments are consolidated under this PR's number.

🤖 Generated with Claude Code

…ery reader takes

`to_json`, `to_key_value` and `canonicalize` wrote a field with nothing
to configure by its bare name: `"crc32c"`, `"default"`. The spec allows
that since v3.1, but a Zarr v3.0 reader takes no short-hand name in
`codecs` (core L585-L592), and no zarr-python release reads one there or
in `chunk_key_encoding`, so documents the package wrote could not be
opened by them. Every extension point but a data type is now written as
an object, `{"name": ...}`; a data type with nothing to configure keeps
its bare name, as core data types have been written since v3.0.
Re-writing zarr-python's own documents through the model, zarr-python
now opens all of them, where 14 to 22 of 80 failed.

Assisted-by: ClaudeCode:claude-opus-5-5
…ry says what the package decided

Four users of the new APIs found the same rough edges. A problem's
message showed the Python value -- `got None`, `got (1,)` -- and named
shapes in Python's words, "a sequence", "a mapping", a one-member
Literal as a one-element tuple, a union of two objects as "an object or
an object". Values now show as the JSON the document holds, shapes in
JSON's words, and a value of another JSON type than a closed set's is
`invalid_type`, so the kind tells a wrong type from a wrong value.

`help` of a validator printed its default scope in full, some 14,000
characters; a scope's repr now says how many definitions it holds.

The README's validation boundary says what the package decides where
the specs leave it open: non-finite numbers in attributes, written back
as bare tokens; `must_understand: false` refused at every extension
point; a chunk length of 0 along a dimension of length 0. The models'
I/O methods have docstrings, and no public docstring names a private
function.

Assisted-by: ClaudeCode:claude-opus-5-5
…de the problems

The v3 array validator read every extension point, the chunks the codecs
are handed and each codec's stage, then kept only the problems, so a
caller wanting any of it -- a core-only policy, each codec's stage,
what was left unjudged -- resolved every field again. The reading is
now returned as a ZarrV3ArrayMetadataReading, and
validate_array_metadata_v3 is its problems.

fields() gives each field with where it sits in the document, and after
it the fields it holds, as the new fields_of gives them for any field
resolve read. Resolved gains the kind a field was read as, which a
field nothing in scope claims keeps too: before, a nested field nothing
claimed told nothing of its kind, and a pipelines hook naming data
types nothing claims read them as codecs. Built by hand, a Resolved is
checked as resolve builds one: its kind one of the five, and its
definition of that kind. A field its rules refuse keeps the fields it
read inside it, whose problems it already reported, so fields() leaves
none of them out.

Assisted-by: ClaudeCode:claude-opus-5-5
…, and one read builds a document's reading and model

- `resolve` gives one of three frozen dataclasses, built with keywords:
  `Read`, by the definition in scope that claims the field's name, with
  its definition and configuration; `Unclaimed`, when nothing in scope
  claims it; or `Refused`, with the definition that refused it, if any.
  `Resolved` is their union, for `match`. Each answers `name`,
  `read_as`, `definition` and `nested`. They compare by what they read,
  not how they were spelled: a configuration holds each field in it as
  `to_json` writes it.
- Models hold those fields. `ZarrV3NamedConfig`, `ZarrV3MetadataField`
  (the model's and the pydantic one), `ZarrV3ArrayMetadataPartial` and
  `ZarrV3GroupMetadataPartial` are gone.
- `read_array_metadata_v3` returns one reading, holding every problem
  and, when there is none, the model built by the same read;
  `read_group_metadata_v3` does the same for a group, reading each
  document its consolidated metadata holds once and holding the model of
  each that has no problem. `from_json` is the model, or the problems
  raised; `validate_*` are the problems, and build no models.
- A model holds no scope, as a pydantic model holds no validation
  context. `update(context=..., **members)` reads only the members it is
  given, in that scope; the fields the model holds pass through as they
  were read; `UNSET` leaves a member out; a group keeps the documents its
  consolidated metadata holds unless it is given others. The pydantic
  field types read in the scope their validation context holds.
- `to_key_value` reads the document it writes by the model's own fields,
  and refuses one with a problem, so a model changed by hand is never
  written invalid.
- A scope pickles as its definitions and copies as itself; the core
  definitions compare equal once pickled; `configuration_of` takes an
  equal definition. A read copies no document first, reads each member
  once, and reads a field a configuration holds a frame deeper than it.
- The changelog describes what ships: #370's and #372's fragments fold
  into #373's, and the statements of 4434, 4436 and 4443 that this
  replaces are dropped from theirs.

Assisted-by: ClaudeCode:claude-opus-5-5
…ys, as a discriminated union reads its tag

- `read_node_metadata_v3(value, *, context)` reads a v3 `zarr.json` of
  either kind: as `read_array_metadata_v3` or `read_group_metadata_v3`
  reads it, as its `node_type` says -- the tag of a union, as pydantic's
  discriminator and zod's discriminated union read one -- or, when it
  says neither or is not an object, as `ZarrV3UnknownNodeReading`, with
  the problem, reading nothing else of it. `ZarrV3NodeMetadataReading`
  is the three; `validate_node_metadata_v3` gives the problems.
- `node_metadata_from_json_v3` and `node_metadata_from_key_value_v3`
  build the model of either kind, `ZarrV3NodeMetadata`, as the models'
  own `from_json` and `from_key_value` build one, and as pydantic's
  `TypeAdapter` validates a discriminated union from Python and JSON.
- A group's consolidated metadata reads each document it holds as a
  node, so one of no node type is in `reading.consolidated`, as unknown,
  rather than missing. Its problem is located as before, and reads as a
  missing key when the entry has no `node_type`, and at the entry when
  it is not an object.

Assisted-by: ClaudeCode:claude-opus-5-5
…ries what was found and what was expected

Two borrowings from pydantic and zod.

Bounds on the type. A number's type carries its bounds, in
annotated-types' vocabulary as pydantic reads it: a gzip `level` is
`Annotated[int, Interval(ge=0, le=9)]`. The checker holds a value to
`Gt`, `Ge`, `Lt`, `Le` and `Interval`, one bound from each side, at any
depth: one problem per value, whose message says what the type admits
and whose `ctx` holds the bounds. A bound is held as the number it
equals, so Python's `Annotated` cache, which takes `Ge(True)` for
`Ge(1)`, cannot make a verdict or a message depend on import order.
`Annotated` may also carry a note, a string or a `Doc`. Any other
metadata is a `TypeError` when the type is compiled, so the checker and
pydantic never read one type two ways. `typeddict_keys` keeps a
member's metadata. The bounds that were rules move onto the types: gzip,
zstd, blosc's `clevel` and `blocksize`, the regular and rectilinear
grids, the shard's inner chunk shape, the eight integer fill values,
byte values, and the numpy time types' `scale_factor` and ticks.
`_InRange`, `byte_value_problems` and the numpy time rules are gone.

Problems carry data. `ValidationProblem` gains `input`, the JSON the
value handed in holds at `loc` (`UNSET` where nothing is, or where what
is there is not JSON, so an error always pickles), and `ctx`, what was
expected: a type's bounds by pydantic's names, or `expected` for a
closed set. Neither takes part in equality or the repr, and `ctx` is
read-only. Every reader fills `input` from the value its caller handed
it, so a rule says only where a problem is. `UNSET` moves below the
problem record.

BREAKING CHANGE: `typeddict_keys(...).members` keeps `Annotated`
metadata, and `check` and `Definition.check` hold values to the bounds
a type carries. A definition's rules are asked only of a configuration
within its bounds, as pydantic's after-validators are, so a value out
of bounds hides a rule's problem until it is fixed: blosc's `typesize`
behind a bad `clevel`, an `r16` fill value's count behind a bad byte. A
definition that reuses a configuration TypedDict takes its bounds with
it. A v2 `order` or `dimension_separator`, and a consolidated
envelope's `must_understand`, are reported as values outside a closed
set, `expected one of ["C", "F"], got "Q"`, and a `must_understand`
that is not a boolean as `invalid_type`. A numpy time fill value out of
range is reported as out of its integer range, without naming "NaT".

Assisted-by: ClaudeCode:claude-opus-5-5
… and canonical_of spells a field read already

`with_problems(fields, problems)` gives each field of one read -- a
reading's `fields()` and `problems`, or `fields_of` a field and the
problems `resolve` gave with it -- with the problems located in it, in
the fields it holds too: those it was read with, and those the document
found with it where it stands, its place in the pipeline, the chunk it
is handed, the array's shape. A field with none is valid there. A
function of the problems, as zod's `flattenError` is of the issues,
grouping them at every depth; a problem in no field is in none's.

`canonical_of(resolved, problems)` spells a field a scope has read
already, given its problems, in its simplest equivalent spelling,
without reading it again: None for a field with a problem, as
`canonicalize` gives, so a key its TypedDict does not declare is never
erased. `canonicalize` is `canonical_of(*resolve(...))`.

Assisted-by: ClaudeCode:claude-opus-5-5
…s its specification says

"Chunk sizes must be greater than zero"
(https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L40),
along a dimension of length 0 too. `RegularChunkGridConfiguration`
bounded its lengths with `Ge(0)`, and its shape rules took a 0 on an
empty dimension, reading the core specification's "The chunk shape
elements are non-zero when the corresponding dimensions of the arrays
have non-zero length" as allowing it. That sentence says less than the
grid's own. The bound is now `Ge(1)`, so a 0 is an `invalid_value` at
its place wherever it is, and the shape rules check only that the grid
has the array's dimensionality. `create_default` already wrote 1 there.

BREAKING CHANGE: a regular grid with a chunk length of 0 on a dimension
of length 0, as zarr-python 3.0 and 3.1 wrote one, is now refused.

Assisted-by: ClaudeCode:claude-opus-5-5
… and a zarr.json

As pydantic's `TypeAdapter(...).json_schema()` and zod's `toJSONSchema`
write theirs, in draft 2020-12:

- `json_schema(shape)`, in `zarr_metadata.typed_json`, writes a
  TypedDict as `check` reads it. A bound is JSON Schema's keyword for
  it, the stricter where a type and its `NewType` both say one, and a
  `Doc` the `description`. Each TypedDict and type alias is written
  once, in `$defs`, under its name.
- `field_json_schema(kind, context)`, in `zarr_metadata.v3.definition`,
  writes a field of one kind as a scope reads it: each definition's
  field, and a name none of them is written with, with any
  configuration. Raw bits' name is matched to its end, so a validator
  that matches patterns as Python does takes no final newline for it.
- `node_metadata_json_schema_v3(context=...)`, in `zarr_metadata.model`,
  writes a `zarr.json`: its fill value held to the data type it names,
  and a group's consolidated metadata holding documents.

`Schemas` writes each shape, asking a caller's `SchemaLeaf` first, as
the checker asks a `Leaf`, and hands back a schema sharing nothing. A
schema says what the types say and not what the rules say, so every
JSON document the package finds nothing wrong with, the schema accepts.
Property tests hold the typed_json schema to the checker's reference,
bounds in layers to what `check` holds a value to, and fields and
documents changed in one or two places to soundness. JSON Schema takes
`1.0` for an integer. Only a `Doc` is a description, as in zod: the
package's docstrings are written for Python's readers.

The six field aliases move to `v3._common`, so `ZarrV3ArrayMetadataJSON`
says what each extension point is: `data_type: DataTypeField`,
`codecs: tuple[CodecField, ...]`. To a type checker and to `check` they
are the JSON they were, and the model's tables of extension points are
read off the TypedDict. `shape` holds integers of at least 0, which
`check` now holds it to. A data type whose `fill_value` holds a metadata
field is refused when it is built, since the checker reads a fill value
as a value and a schema would write it as a field.

Assisted-by: ClaudeCode:claude-opus-5-5
…key_value writes it as it is

As pydantic validates in `__init__`: a v3 or v2 model's constructor
reads its document by its own fields, the read the writer made before,
raises `MetadataValidationError` with every problem, an extra field
named as a member among them, and holds its members as that read
refines them, in containers of its own, as pydantic holds what its
`__init__` coerced. So a model built by hand, or changed by
`dataclasses.replace` or a v2 `update`, is refused at the change rather
than at the write; none is built invalid, none shares a container with
what it was built of, and each writes as JSON. Each field, a `Read` or
an `Unclaimed`, is held as the scope read it.

The reads build their models with a private `construct`, as pydantic's
`model_construct` builds one, so a model a read builds is not read
twice, and `to_key_value` serializes without reading again. A group
checks its own members, since each document its consolidated metadata
holds is a model that checked itself; a consolidated path that is not a
string is a `TypeError`.

`ZarrV2ArrayMetadata.create_default(chunks=...)` without a shape they
fit is refused, as the v3 model refuses a grid its shape does not take.
The clauses about the checking writer that this makes untrue are
dropped from 373's and 4420's fragments.

BREAKING CHANGE: building or replacing a model into one whose document
has a problem raises, and a model holds the containers a read would, not
the ones it was given; a container changed in place is not checked again.

Assisted-by: ClaudeCode:claude-opus-5-5
…nknown key wherever it sits

`ProblemKind` defines `unknown_key` as a key an object's type does not
declare where the type is closed, and a v3 configuration reported one
so. A v2 `.zgroup`, a `.zmetadata` envelope, a group's
`consolidated_metadata` and a metadata field's envelope reported the same
case as `invalid_value`, so a reader that tolerates what another writer
added -- netCDF-C's `_nczarr_*` keys (zarr-python#2296) -- could not
filter it by kind. Each is now `unknown_key`, with the checker's message,
"unexpected key 'x'". A definition's `check` and `judge` give a
configuration back when a field it holds has a stray member, reported
and left out, as the checker leaves out a key a closed TypedDict does not
declare; a `must_understand` of `false` still refuses it.

Assisted-by: ClaudeCode:claude-opus-5-5
…s in

`read_node_metadata_v3` reads a document by the node type it names, and
one that names none was reported only as missing its `node_type`. A
crawler meeting the root `zarr.json` of zarr-python 2's draft of v3
(zarr-python#2982), whose `zarr_format` is a URL, learned "no node type",
not "not v3". A document the dispatch cannot follow is now judged by its
`zarr_format` as well: missing, or other than 3, is a problem beside the
node type's.

Assisted-by: ClaudeCode:claude-opus-5-5
…below its group

A group's consolidated metadata holds the hierarchy below the group:
zarr-python keeps the document of the node at /a/b of that hierarchy,
the group its root, at the key a/b. Nothing judged the keys or the tree
they make. A key that is no node's path ("", "__a", "a/../b", "/a"), a
document below an array's, and a document whose group is missing were
read without a problem; zarr-python raises a bare KeyError on three of
them and keeps the rest. Each is now a problem at its entry, a missing
group a missing_key, and ZarrV3ConsolidatedMetadata refuses them when it
is built. A listed group's own listing lists what the group lists, at
the joined key and of the same node type: a node it lists alone is
dropped by zarr-python's reader, which keeps the flat listing.

The rules are a hierarchy's, not consolidated metadata's: a private
hierarchy_problems judges the node type of each node, by its node path,
as the spec's tree, so a model of a whole hierarchy can use it too. Every
problem is one per value, saying its reasons at once -- a path of a
thousand bad names is one problem, a chain of missing groups one
missing_key at the nearest -- so what is reported weighs what was read.
NodeName and NodePath, in zarr_metadata.v3, are the spec's node names
and paths, modelled on zarrs' types, and zarr_metadata.model judges a
string by them with validate/is/parse_node_name_v3 and their node_path
twins.

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-fable-5-1
…ment

A v3 array model compared its members with ==, so "NaN" and "0x7fc00000"
were two float32 fill values, 0.0 and -0.0 one, a blosc with and without
the typesize that noshuffle ignores two codecs though canonical_of spelled
them alike, and a model holding a NaN attribute never equalled its own
copy. Two models are now equal when they mean the same document: what
the package interprets compares by its canonical spelling, and what it
does not compares as JSON text, which tells true from 1 and -0.0 from
0.0 and takes NaN for itself.

A data type's definition says which spelling of a fill value is its
value's own, DataTypeDefinition.fill_value_canonical, and
canonical_fill_value spells one, or gives UNSET for a fill value with a
problem: the float types read a number as a float64, as JSON parsers and
numpy do, and spell a value by its shortest number, a named value, or a
NaN's bits; complex by part; the numpy time types' -2**63 as "NaT"; bytes
as base64; a struct field by field. A field, Read, Unclaimed or Refused,
compares by field_key: its definition and its canonical configuration as
text, with the fields it holds by their own keys; default, v2,
sharding_indexed and zstd gain a canonical folding their spec defaults.
Every model and field hashes as it compares.

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-fable-5-1
…of the readers

This commit fixes what six review passes found, beyond the fixes already
in #381.

Nesting depth. A reader now walks at most 256 levels (JSON_DEPTH) and
reports a deeper level as an invalid_value problem. Before, a small
document nested a thousand levels deep made every validator, reader,
parser and from_key_value raise RecursionError, and json.loads on store
bytes made of 100,000 "[" did the same (now an invalid_json problem). A
chain of groups, each holding the next in its consolidated metadata, was
also unbounded. Each document is now read from its position in the
outer document, so depth is counted from the outer root. is_json and the
is_* guards stop at the same depth. A v2 dtype of field records is
checked as JSON before its shape is read. Readers, writers and
comparisons use one stack frame per level; the v2 models no longer use
copy.deepcopy, which used two. validate_json takes the location of its
value, so the v2 validator counts depth from the document root as the v3
readers do. arrays_to_tuples uses one frame per level. A problem copied by
with_input or prefixed is not re-checked. A struct fill value is judged in
linear time. An integer too long to print is shown by its bit length, a
value too deeply nested is described instead of printed, and a non-string
key is shown with Python's repr.

Hand-built field objects. A Read or Unclaimed object placed inside a
document passed to a reader is not JSON and is now refused. Before, it
passed validation as if it had been read, so a model could write a
document that the same validator then refused. Only a model's own fields
are treated as already read.

Absent values. A reading's data_type, chunk_grid and chunk_key_encoding
are UNSET when the document has no such key, and Refused.json is UNSET
when the field was not JSON. Both used to be None, which is also what a
JSON null becomes. A field with no name is a missing_key problem. A
consolidated_metadata of null is an invalid_type problem; it is no longer
read as absent, and neither JSON Schema admits it.
ZarrV3ConsolidatedMetadata declares must_understand as False.

Extension names. An extension name must match ^[a-z][a-z0-9-_.]+$ or be a
URI in RFC 3986 characters, so Python's re and ECMA-262 read the schema
pattern the same way. Any other name is refused before a definition is
looked up; before, any string was read as an unknown extension.
well_named checks this, validate_metadata_field_v3 enforces it, the JSON
Schema's unclaimed branch carries the pattern, and a definition with an
invalid name raises TypeError when built. A raw-bits name with more than
100 digits is treated as an extension name, not a size. Definition
functions always receive a read-only view of the configuration, and an
error raised by a canonical function names its definition.

Pydantic. The pydantic types raise ValidationError with one line error
per problem, carrying type, loc, input and ctx. For a missing key the
input is the object that lacks it. A message containing a ctx placeholder
like {expected} is passed as the ctx's "message" entry, because pydantic
renders messages as templates with no escaping.

Tests. New tests pin each behavior the mutation review found unpinned:
the stricter of two bounds per keyword, the unclaimed envelope in the
schema, no empty attributes written for a group, value_at at index 0,
every integer type's range, union naming, layered Annotated metadata,
and construct's default factory. The float-bits property test draws
sign, exponent and mantissa separately so it reaches normal values. A
shard nested to the depth cap is read, written, hashed, pickled,
deep-copied and walked; one level deeper reports the depth problem. The
same holds for a v2 document and a v3 fill value at the cap. Docs: two
spec anchors fixed, the struct data type rendered, a note on the
JSONValue export, a typed TypeAdapter example, with_problems compared to
zod's treeifyError, and the float fill value rounding cited to zarrs and
tensorstore.

Assisted-by: ClaudeCode:claude-fable-5-1
…t calling repr

`shown` relied on `repr` raising RecursionError for a deeply nested
value. Python 3.14 limits recursion by the real C stack, so on Linux a
list nested 100,000 levels deep printed fine and the problem message was
the whole value. Now any value with a problem past `JSON_DEPTH` is
reported as "a value nested too deep to show", on every interpreter and
platform. The test nests one level past the limit, where repr always
works.

Assisted-by: ClaudeCode:claude-opus-5-5
…SON values

`shown_by_python` caught RecursionError so that a non-JSON value nested
deeper than repr could handle was described instead of raising. No
parsed document can contain such a value: sets and frozensets are not
JSON. Its test only passed where the platform's stack happened to
overflow at the chosen depth. Values a document can hold that nest past
`JSON_DEPTH` are already reported as too deep by `shown`, without repr.

Assisted-by: ClaudeCode:claude-opus-5-5
- `shown` reports "too deep" only for the depth problem itself. A
  non-JSON value sitting exactly on the last allowed level is now
  printed instead of being mislabeled.
- Complex fill-value rules keep each component problem's `input` and
  `ctx`; only the `loc` is moved under the component index.
- `Schemas.defined` restores every `$defs` entry, name and use count a
  failed write touched, not only its own name.
- The node JSON Schema docstring no longer says `consolidated_metadata`
  may be `null`. Both the schema and the validator refuse that.
- The 382 bugfix changelog fragment is rewritten for users without
  naming internals.
- "modelled" is spelled "modeled" in new text.

Assisted-by: ClaudeCode:claude-opus-5-5
…lish

Assisted-by: ClaudeCode:claude-fable-5-1
The strict readers refuse documents that some writers produced. Repair is
now a separate step that a caller asks for by name:
`read_repaired_node_metadata_v3` applies `repair_node_metadata_v3` to
undo each known writer bug, then reads the result with
`read_node_metadata_v3`, and returns the reading together with the list
of repairs made.

The set of repairable documents is closed. Each repair applies only when
the document's members match that repair's TypedDict:
`ZarrV3ZeroChunkArrayMetadataJSON` for the chunk length of 0 along an
empty dimension that zarr-python 3.0 and 3.1 wrote, and
`ZarrV3NullConsolidatedGroupMetadataJSON` for the
`consolidated_metadata: null` that zarr-python 3.0.x wrote. Documents
inside consolidated metadata are repaired too. Anything else is left for
the strict read to report.

Assisted-by: ClaudeCode:claude-opus-5-5
…lish

Assisted-by: ClaudeCode:claude-fable-5-1
Two scopes are equal when they hold the same definitions under the same
kinds and names, regardless of construction order. Equal scopes hash
alike. Before this, `Context` was not hashable at all because its tables
are mapping proxies.

Assisted-by: ClaudeCode:claude-fable-5-1
…Error

`Claims` maps `(kind, filed name)` to the definition that read it, or
None when nothing claimed it. `Conflict` records one disagreement: the
key, what was claimed, what was found, and where. `ScopeConflictError`
carries every conflict, as `MetadataValidationError` carries every
problem.

Assisted-by: ClaudeCode:claude-fable-5-1
…d for each name

`claims_of` takes the fields of one reading, nested fields included, and
returns the definition that read each name, keyed as the scope files it
(raw bits under `r*`). A name claimed two different ways raises
`ScopeConflictError`.

Assisted-by: ClaudeCode:claude-fable-5-1
…s of a field

`refines(field, other)` is True when `field` holds everything `other`
holds: the same definition and canonical configuration where both were
read, and a read field over an unclaimed one with the same name and JSON
(a gain). The reverse direction is a loss, two different definitions are
a conflict, and both return False. Two fields that refine each other are
equal.

Assisted-by: ClaudeCode:claude-fable-5-1
`Context.disagreements(claims)` reports, for each claim, whether this
scope reads it the same way, would gain a definition, or conflicts with it
(a different definition, or none where one was claimed). `Context.joined`
returns the union of several scopes and raises `ScopeConflictError` when
two of them file different definitions under one name.

Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
`refines` compared a read field's written JSON to the unclaimed field's
with Python `==`. That treated `"crc32c"` and `{"name": "crc32c"}` as
different fields and `true` as equal to `1`, so the order was not
transitive across spellings. A gain is now compared the way `Unclaimed`
equality works: by name and by the configuration as JSON text. A refused
field refines only itself, so a read field that holds one still refines
itself. `Claims` is a real type at run time, not a string, so it works in
signatures. The property test now checks reflexivity, transitivity,
mutual refinement equals equality, and substitutivity, over readings of
several documents in several scopes.

Assisted-by: ClaudeCode:claude-fable-5-1
d-v-b and others added 19 commits October 6, 2026 22:25
…und 1 found refused

Aligned structs name their padding '' and may repeat it. zlib and gzip take level -1, which numcodecs writes. An empty codec id is reported, not accepted silently. create_default with a dtype and no fill value takes null. field_json_schema refuses a kind of another format with a clear error. A nested value that is no field is shown in its message.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…it; narrow blosc clevel

create_default with a dtype keeps fill_value 0 for the numeric families and takes null only where 0 is no fill value. blosc clevel is 0 to 9 again; only zlib and gzip take -1. A typestr size is ASCII digits only. fields_of places a v2 struct record's type at the struct's own location. The changelog says 21 of the numcodecs ids and the create_default change.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A typestr with a listed type code and no size, such as '<f', was filed under itself and left unclaimed, so its fill value went unjudged. It is now a problem: the v2 spec requires the size.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
read_array_metadata_v2 returns a ZarrV2ArrayMetadataReading: dtype, compressor and filters as the scope read them, every problem, and the model when there is none. validate_array_metadata_v2, is_array_metadata_v2 and parse_array_metadata_v2 read in the scope given, CORE_V2 by default; the group readers take a scope too.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…scope

ZarrV2ArrayMetadata(document, context=None) reads the document in a scope, CORE_V2 by default. dtype, compressor and filters are the fields as the scope read them. Equality is by what the document means. update reads new members in the model's own scope; with_context and refined_in read the document in another; refines orders models. ZarrV2ArrayMetadataUpdate replaces ZarrV2ArrayMetadataPartial. A member the spec does not define is kept.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…scope

ZarrV2GroupMetadata(document, context=None) holds its document and scope like the array model. ZarrV2GroupMetadataUpdate replaces ZarrV2GroupMetadataPartial.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… each node as a model

ZarrV2ConsolidatedMetadata(document, context=None) keeps its entries as written and adds nodes: each .zarray or .zgroup entry, merged with its .zattrs, as a model of the consolidated scope, keyed by node path. Equality, refines, with_context and refined_in go through the nodes. A problem in an entry is reported under that entry.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… construct

The v2 pydantic field types read a document in the scope the validation context holds, CORE_V2 when it holds none. construct is gone: every model is built from its document.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…roup entries; a scope key per format

zarr-python 3.x writes a consolidated_metadata member into each .zgroup entry below the root of a v2 .zmetadata, which the strict read refuses. repair_consolidated_metadata_v2 removes it and read_repaired_consolidated_metadata_v2 reads the repaired document. The pydantic field types read the v3 scope under zarr_metadata_context and the v2 scope under zarr_metadata_context_v2; a bare Context is the scope of every field type. Two keys naming one file of one consolidated node are a problem. ZarrV2NodeMetadata is exported.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…; leave the root .zgroup to the strict read

A second key for one file of a v2 consolidated node no longer hides the other nodes' problems, and a leading slash names the same node. The repair leaves the root .zgroup alone, since no writer puts the member there. The fragments say how to read a zarr-python 3.x .zmetadata and which context key the v2 pydantic field types read.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…names the same node

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tion framework

The public doors keep what a user of a document, a model, a definition or a scope needs and no longer export the helpers behind them: the scope algebra's functions, the canonical-spelling and field-walking helpers, the writer-bug JSON shapes, the key sets and the JSON helpers. The README says the core depends on no validation framework and why. ruff knows a TypedDict's annotations are evaluated at run time, so the per-import noqa comments go.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… casts

is_object, is_json_object, is_list_or_tuple, is_tuple and is_array in _json say what isinstance leaves unknown to a type checker, so the readers no longer cast a mapping or sequence to its element types after checking it: 239 casts become 179. The TypedDict parsers report a non-string key as a problem instead of assuming one is a string. The two loops kept for frame depth say so instead of carrying a PERF noqa.

Assisted-by: ClaudeCode:claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… documents instead of casting

frozen keeps an object's type; refined_object and object_at check what a
read already established instead of asserting it; is_alias and is_field
are type guards; field_aliases and the kind registry hold TypeAliasType.
Casts in src go from 179 to 128.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…yed base

Equality and hashing live once, in Keyed; each model computes its key in
a method of its own class instead of a module function reading its
private members. Entries given as node models are unwrapped through
to_json. The remaining reportPrivateUsage sites are the readers building
models through _of, which the docstrings now say.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ting to it

The model's field properties go through held, which raises for a field
a read refused or never read; the canonical checks and the nested-field
walkers use the type guards; reading_of takes a Protocol. Casts in src
go from 128 to 100.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions github-actions Bot added the needs release notes Automatically applied to PRs which haven't added release notes label Oct 8, 2026
@read-the-docs-community

read-the-docs-community Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

…r upstream PR 4490

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@d-v-b
d-v-b marked this pull request as ready for review October 8, 2026 13:58
Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@codecov

codecov Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.69%. Comparing base (246fa20) to head (bc4eced).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #4490   +/-   ##
=======================================
  Coverage   94.69%   94.69%           
=======================================
  Files          94       94           
  Lines       13619    13619           
=======================================
  Hits        12897    12897           
  Misses        722      722           
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

d-v-b and others added 5 commits October 8, 2026 16:24
A kind is declared in its class header, class CodecDefinition(Definition[C], kind=True),
as typing.Protocol and SQLAlchemy's __abstract__ mark a class and not its
subclasses, instead of an is_kind ClassVar that a subclass inherits but that did
not count. The hook keeps the mark in the class namespace, which the slots
rebuild of a dataclass carries over while the keyword is not.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…eld, UnclaimedField, RefusedField

Read, Unclaimed, Refused and their union Resolved become AcceptedField,
UnclaimedField, RefusedField and ResolvedField: the noun says these are
metadata fields as a scope read them, and Accepted is the opposite of
Refused where Read read like a verb.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…n; its check step is private

A definition reads a configuration as the model readers read a document:
the pair of what was read and every problem, never raising. The type-check
step that only read_configuration called is _check_configuration; the
public shape check stays zarr_metadata.typed_json.check.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs release notes Automatically applied to PRs which haven't added release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant