Repository navigation
feat(zarr-metadata)!: read metadata against definitions in a scope, and make every model a pair of document and scope - #4490
Open
d-v-b wants to merge 88 commits into
Conversation
…ery reader takes
`to_json`, `to_key_value` and `canonicalize` wrote a field with nothing
to configure by its bare name: `"crc32c"`, `"default"`. The spec allows
that since v3.1, but a Zarr v3.0 reader takes no short-hand name in
`codecs` (core L585-L592), and no zarr-python release reads one there or
in `chunk_key_encoding`, so documents the package wrote could not be
opened by them. Every extension point but a data type is now written as
an object, `{"name": ...}`; a data type with nothing to configure keeps
its bare name, as core data types have been written since v3.0.
Re-writing zarr-python's own documents through the model, zarr-python
now opens all of them, where 14 to 22 of 80 failed.
Assisted-by: ClaudeCode:claude-opus-5-5
…ry says what the package decided Four users of the new APIs found the same rough edges. A problem's message showed the Python value -- `got None`, `got (1,)` -- and named shapes in Python's words, "a sequence", "a mapping", a one-member Literal as a one-element tuple, a union of two objects as "an object or an object". Values now show as the JSON the document holds, shapes in JSON's words, and a value of another JSON type than a closed set's is `invalid_type`, so the kind tells a wrong type from a wrong value. `help` of a validator printed its default scope in full, some 14,000 characters; a scope's repr now says how many definitions it holds. The README's validation boundary says what the package decides where the specs leave it open: non-finite numbers in attributes, written back as bare tokens; `must_understand: false` refused at every extension point; a chunk length of 0 along a dimension of length 0. The models' I/O methods have docstrings, and no public docstring names a private function. Assisted-by: ClaudeCode:claude-opus-5-5
…de the problems The v3 array validator read every extension point, the chunks the codecs are handed and each codec's stage, then kept only the problems, so a caller wanting any of it -- a core-only policy, each codec's stage, what was left unjudged -- resolved every field again. The reading is now returned as a ZarrV3ArrayMetadataReading, and validate_array_metadata_v3 is its problems. fields() gives each field with where it sits in the document, and after it the fields it holds, as the new fields_of gives them for any field resolve read. Resolved gains the kind a field was read as, which a field nothing in scope claims keeps too: before, a nested field nothing claimed told nothing of its kind, and a pipelines hook naming data types nothing claims read them as codecs. Built by hand, a Resolved is checked as resolve builds one: its kind one of the five, and its definition of that kind. A field its rules refuse keeps the fields it read inside it, whose problems it already reported, so fields() leaves none of them out. Assisted-by: ClaudeCode:claude-opus-5-5
…, and one read builds a document's reading and model - `resolve` gives one of three frozen dataclasses, built with keywords: `Read`, by the definition in scope that claims the field's name, with its definition and configuration; `Unclaimed`, when nothing in scope claims it; or `Refused`, with the definition that refused it, if any. `Resolved` is their union, for `match`. Each answers `name`, `read_as`, `definition` and `nested`. They compare by what they read, not how they were spelled: a configuration holds each field in it as `to_json` writes it. - Models hold those fields. `ZarrV3NamedConfig`, `ZarrV3MetadataField` (the model's and the pydantic one), `ZarrV3ArrayMetadataPartial` and `ZarrV3GroupMetadataPartial` are gone. - `read_array_metadata_v3` returns one reading, holding every problem and, when there is none, the model built by the same read; `read_group_metadata_v3` does the same for a group, reading each document its consolidated metadata holds once and holding the model of each that has no problem. `from_json` is the model, or the problems raised; `validate_*` are the problems, and build no models. - A model holds no scope, as a pydantic model holds no validation context. `update(context=..., **members)` reads only the members it is given, in that scope; the fields the model holds pass through as they were read; `UNSET` leaves a member out; a group keeps the documents its consolidated metadata holds unless it is given others. The pydantic field types read in the scope their validation context holds. - `to_key_value` reads the document it writes by the model's own fields, and refuses one with a problem, so a model changed by hand is never written invalid. - A scope pickles as its definitions and copies as itself; the core definitions compare equal once pickled; `configuration_of` takes an equal definition. A read copies no document first, reads each member once, and reads a field a configuration holds a frame deeper than it. - The changelog describes what ships: #370's and #372's fragments fold into #373's, and the statements of 4434, 4436 and 4443 that this replaces are dropped from theirs. Assisted-by: ClaudeCode:claude-opus-5-5
…ys, as a discriminated union reads its tag - `read_node_metadata_v3(value, *, context)` reads a v3 `zarr.json` of either kind: as `read_array_metadata_v3` or `read_group_metadata_v3` reads it, as its `node_type` says -- the tag of a union, as pydantic's discriminator and zod's discriminated union read one -- or, when it says neither or is not an object, as `ZarrV3UnknownNodeReading`, with the problem, reading nothing else of it. `ZarrV3NodeMetadataReading` is the three; `validate_node_metadata_v3` gives the problems. - `node_metadata_from_json_v3` and `node_metadata_from_key_value_v3` build the model of either kind, `ZarrV3NodeMetadata`, as the models' own `from_json` and `from_key_value` build one, and as pydantic's `TypeAdapter` validates a discriminated union from Python and JSON. - A group's consolidated metadata reads each document it holds as a node, so one of no node type is in `reading.consolidated`, as unknown, rather than missing. Its problem is located as before, and reads as a missing key when the entry has no `node_type`, and at the entry when it is not an object. Assisted-by: ClaudeCode:claude-opus-5-5
…ries what was found and what was expected Two borrowings from pydantic and zod. Bounds on the type. A number's type carries its bounds, in annotated-types' vocabulary as pydantic reads it: a gzip `level` is `Annotated[int, Interval(ge=0, le=9)]`. The checker holds a value to `Gt`, `Ge`, `Lt`, `Le` and `Interval`, one bound from each side, at any depth: one problem per value, whose message says what the type admits and whose `ctx` holds the bounds. A bound is held as the number it equals, so Python's `Annotated` cache, which takes `Ge(True)` for `Ge(1)`, cannot make a verdict or a message depend on import order. `Annotated` may also carry a note, a string or a `Doc`. Any other metadata is a `TypeError` when the type is compiled, so the checker and pydantic never read one type two ways. `typeddict_keys` keeps a member's metadata. The bounds that were rules move onto the types: gzip, zstd, blosc's `clevel` and `blocksize`, the regular and rectilinear grids, the shard's inner chunk shape, the eight integer fill values, byte values, and the numpy time types' `scale_factor` and ticks. `_InRange`, `byte_value_problems` and the numpy time rules are gone. Problems carry data. `ValidationProblem` gains `input`, the JSON the value handed in holds at `loc` (`UNSET` where nothing is, or where what is there is not JSON, so an error always pickles), and `ctx`, what was expected: a type's bounds by pydantic's names, or `expected` for a closed set. Neither takes part in equality or the repr, and `ctx` is read-only. Every reader fills `input` from the value its caller handed it, so a rule says only where a problem is. `UNSET` moves below the problem record. BREAKING CHANGE: `typeddict_keys(...).members` keeps `Annotated` metadata, and `check` and `Definition.check` hold values to the bounds a type carries. A definition's rules are asked only of a configuration within its bounds, as pydantic's after-validators are, so a value out of bounds hides a rule's problem until it is fixed: blosc's `typesize` behind a bad `clevel`, an `r16` fill value's count behind a bad byte. A definition that reuses a configuration TypedDict takes its bounds with it. A v2 `order` or `dimension_separator`, and a consolidated envelope's `must_understand`, are reported as values outside a closed set, `expected one of ["C", "F"], got "Q"`, and a `must_understand` that is not a boolean as `invalid_type`. A numpy time fill value out of range is reported as out of its integer range, without naming "NaT". Assisted-by: ClaudeCode:claude-opus-5-5
… and canonical_of spells a field read already `with_problems(fields, problems)` gives each field of one read -- a reading's `fields()` and `problems`, or `fields_of` a field and the problems `resolve` gave with it -- with the problems located in it, in the fields it holds too: those it was read with, and those the document found with it where it stands, its place in the pipeline, the chunk it is handed, the array's shape. A field with none is valid there. A function of the problems, as zod's `flattenError` is of the issues, grouping them at every depth; a problem in no field is in none's. `canonical_of(resolved, problems)` spells a field a scope has read already, given its problems, in its simplest equivalent spelling, without reading it again: None for a field with a problem, as `canonicalize` gives, so a key its TypedDict does not declare is never erased. `canonicalize` is `canonical_of(*resolve(...))`. Assisted-by: ClaudeCode:claude-opus-5-5
…s its specification says "Chunk sizes must be greater than zero" (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L40), along a dimension of length 0 too. `RegularChunkGridConfiguration` bounded its lengths with `Ge(0)`, and its shape rules took a 0 on an empty dimension, reading the core specification's "The chunk shape elements are non-zero when the corresponding dimensions of the arrays have non-zero length" as allowing it. That sentence says less than the grid's own. The bound is now `Ge(1)`, so a 0 is an `invalid_value` at its place wherever it is, and the shape rules check only that the grid has the array's dimensionality. `create_default` already wrote 1 there. BREAKING CHANGE: a regular grid with a chunk length of 0 on a dimension of length 0, as zarr-python 3.0 and 3.1 wrote one, is now refused. Assisted-by: ClaudeCode:claude-opus-5-5
… and a zarr.json As pydantic's `TypeAdapter(...).json_schema()` and zod's `toJSONSchema` write theirs, in draft 2020-12: - `json_schema(shape)`, in `zarr_metadata.typed_json`, writes a TypedDict as `check` reads it. A bound is JSON Schema's keyword for it, the stricter where a type and its `NewType` both say one, and a `Doc` the `description`. Each TypedDict and type alias is written once, in `$defs`, under its name. - `field_json_schema(kind, context)`, in `zarr_metadata.v3.definition`, writes a field of one kind as a scope reads it: each definition's field, and a name none of them is written with, with any configuration. Raw bits' name is matched to its end, so a validator that matches patterns as Python does takes no final newline for it. - `node_metadata_json_schema_v3(context=...)`, in `zarr_metadata.model`, writes a `zarr.json`: its fill value held to the data type it names, and a group's consolidated metadata holding documents. `Schemas` writes each shape, asking a caller's `SchemaLeaf` first, as the checker asks a `Leaf`, and hands back a schema sharing nothing. A schema says what the types say and not what the rules say, so every JSON document the package finds nothing wrong with, the schema accepts. Property tests hold the typed_json schema to the checker's reference, bounds in layers to what `check` holds a value to, and fields and documents changed in one or two places to soundness. JSON Schema takes `1.0` for an integer. Only a `Doc` is a description, as in zod: the package's docstrings are written for Python's readers. The six field aliases move to `v3._common`, so `ZarrV3ArrayMetadataJSON` says what each extension point is: `data_type: DataTypeField`, `codecs: tuple[CodecField, ...]`. To a type checker and to `check` they are the JSON they were, and the model's tables of extension points are read off the TypedDict. `shape` holds integers of at least 0, which `check` now holds it to. A data type whose `fill_value` holds a metadata field is refused when it is built, since the checker reads a fill value as a value and a schema would write it as a field. Assisted-by: ClaudeCode:claude-opus-5-5
…key_value writes it as it is As pydantic validates in `__init__`: a v3 or v2 model's constructor reads its document by its own fields, the read the writer made before, raises `MetadataValidationError` with every problem, an extra field named as a member among them, and holds its members as that read refines them, in containers of its own, as pydantic holds what its `__init__` coerced. So a model built by hand, or changed by `dataclasses.replace` or a v2 `update`, is refused at the change rather than at the write; none is built invalid, none shares a container with what it was built of, and each writes as JSON. Each field, a `Read` or an `Unclaimed`, is held as the scope read it. The reads build their models with a private `construct`, as pydantic's `model_construct` builds one, so a model a read builds is not read twice, and `to_key_value` serializes without reading again. A group checks its own members, since each document its consolidated metadata holds is a model that checked itself; a consolidated path that is not a string is a `TypeError`. `ZarrV2ArrayMetadata.create_default(chunks=...)` without a shape they fit is refused, as the v3 model refuses a grid its shape does not take. The clauses about the checking writer that this makes untrue are dropped from 373's and 4420's fragments. BREAKING CHANGE: building or replacing a model into one whose document has a problem raises, and a model holds the containers a read would, not the ones it was given; a container changed in place is not checked again. Assisted-by: ClaudeCode:claude-opus-5-5
…nknown key wherever it sits `ProblemKind` defines `unknown_key` as a key an object's type does not declare where the type is closed, and a v3 configuration reported one so. A v2 `.zgroup`, a `.zmetadata` envelope, a group's `consolidated_metadata` and a metadata field's envelope reported the same case as `invalid_value`, so a reader that tolerates what another writer added -- netCDF-C's `_nczarr_*` keys (zarr-python#2296) -- could not filter it by kind. Each is now `unknown_key`, with the checker's message, "unexpected key 'x'". A definition's `check` and `judge` give a configuration back when a field it holds has a stray member, reported and left out, as the checker leaves out a key a closed TypedDict does not declare; a `must_understand` of `false` still refuses it. Assisted-by: ClaudeCode:claude-opus-5-5
…s in `read_node_metadata_v3` reads a document by the node type it names, and one that names none was reported only as missing its `node_type`. A crawler meeting the root `zarr.json` of zarr-python 2's draft of v3 (zarr-python#2982), whose `zarr_format` is a URL, learned "no node type", not "not v3". A document the dispatch cannot follow is now judged by its `zarr_format` as well: missing, or other than 3, is a problem beside the node type's. Assisted-by: ClaudeCode:claude-opus-5-5
…below its group
A group's consolidated metadata holds the hierarchy below the group:
zarr-python keeps the document of the node at /a/b of that hierarchy,
the group its root, at the key a/b. Nothing judged the keys or the tree
they make. A key that is no node's path ("", "__a", "a/../b", "/a"), a
document below an array's, and a document whose group is missing were
read without a problem; zarr-python raises a bare KeyError on three of
them and keeps the rest. Each is now a problem at its entry, a missing
group a missing_key, and ZarrV3ConsolidatedMetadata refuses them when it
is built. A listed group's own listing lists what the group lists, at
the joined key and of the same node type: a node it lists alone is
dropped by zarr-python's reader, which keeps the flat listing.
The rules are a hierarchy's, not consolidated metadata's: a private
hierarchy_problems judges the node type of each node, by its node path,
as the spec's tree, so a model of a whole hierarchy can use it too. Every
problem is one per value, saying its reasons at once -- a path of a
thousand bad names is one problem, a chain of missing groups one
missing_key at the nearest -- so what is reported weighs what was read.
NodeName and NodePath, in zarr_metadata.v3, are the spec's node names
and paths, modelled on zarrs' types, and zarr_metadata.model judges a
string by them with validate/is/parse_node_name_v3 and their node_path
twins.
Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-fable-5-1
…ment A v3 array model compared its members with ==, so "NaN" and "0x7fc00000" were two float32 fill values, 0.0 and -0.0 one, a blosc with and without the typesize that noshuffle ignores two codecs though canonical_of spelled them alike, and a model holding a NaN attribute never equalled its own copy. Two models are now equal when they mean the same document: what the package interprets compares by its canonical spelling, and what it does not compares as JSON text, which tells true from 1 and -0.0 from 0.0 and takes NaN for itself. A data type's definition says which spelling of a fill value is its value's own, DataTypeDefinition.fill_value_canonical, and canonical_fill_value spells one, or gives UNSET for a fill value with a problem: the float types read a number as a float64, as JSON parsers and numpy do, and spell a value by its shortest number, a named value, or a NaN's bits; complex by part; the numpy time types' -2**63 as "NaT"; bytes as base64; a struct field by field. A field, Read, Unclaimed or Refused, compares by field_key: its definition and its canonical configuration as text, with the fields it holds by their own keys; default, v2, sharding_indexed and zstd gain a canonical folding their spec defaults. Every model and field hashes as it compares. Assisted-by: ClaudeCode:claude-opus-5-5 Assisted-by: ClaudeCode:claude-fable-5-1
…of the readers This commit fixes what six review passes found, beyond the fixes already in #381. Nesting depth. A reader now walks at most 256 levels (JSON_DEPTH) and reports a deeper level as an invalid_value problem. Before, a small document nested a thousand levels deep made every validator, reader, parser and from_key_value raise RecursionError, and json.loads on store bytes made of 100,000 "[" did the same (now an invalid_json problem). A chain of groups, each holding the next in its consolidated metadata, was also unbounded. Each document is now read from its position in the outer document, so depth is counted from the outer root. is_json and the is_* guards stop at the same depth. A v2 dtype of field records is checked as JSON before its shape is read. Readers, writers and comparisons use one stack frame per level; the v2 models no longer use copy.deepcopy, which used two. validate_json takes the location of its value, so the v2 validator counts depth from the document root as the v3 readers do. arrays_to_tuples uses one frame per level. A problem copied by with_input or prefixed is not re-checked. A struct fill value is judged in linear time. An integer too long to print is shown by its bit length, a value too deeply nested is described instead of printed, and a non-string key is shown with Python's repr. Hand-built field objects. A Read or Unclaimed object placed inside a document passed to a reader is not JSON and is now refused. Before, it passed validation as if it had been read, so a model could write a document that the same validator then refused. Only a model's own fields are treated as already read. Absent values. A reading's data_type, chunk_grid and chunk_key_encoding are UNSET when the document has no such key, and Refused.json is UNSET when the field was not JSON. Both used to be None, which is also what a JSON null becomes. A field with no name is a missing_key problem. A consolidated_metadata of null is an invalid_type problem; it is no longer read as absent, and neither JSON Schema admits it. ZarrV3ConsolidatedMetadata declares must_understand as False. Extension names. An extension name must match ^[a-z][a-z0-9-_.]+$ or be a URI in RFC 3986 characters, so Python's re and ECMA-262 read the schema pattern the same way. Any other name is refused before a definition is looked up; before, any string was read as an unknown extension. well_named checks this, validate_metadata_field_v3 enforces it, the JSON Schema's unclaimed branch carries the pattern, and a definition with an invalid name raises TypeError when built. A raw-bits name with more than 100 digits is treated as an extension name, not a size. Definition functions always receive a read-only view of the configuration, and an error raised by a canonical function names its definition. Pydantic. The pydantic types raise ValidationError with one line error per problem, carrying type, loc, input and ctx. For a missing key the input is the object that lacks it. A message containing a ctx placeholder like {expected} is passed as the ctx's "message" entry, because pydantic renders messages as templates with no escaping. Tests. New tests pin each behavior the mutation review found unpinned: the stricter of two bounds per keyword, the unclaimed envelope in the schema, no empty attributes written for a group, value_at at index 0, every integer type's range, union naming, layered Annotated metadata, and construct's default factory. The float-bits property test draws sign, exponent and mantissa separately so it reaches normal values. A shard nested to the depth cap is read, written, hashed, pickled, deep-copied and walked; one level deeper reports the depth problem. The same holds for a v2 document and a v3 fill value at the cap. Docs: two spec anchors fixed, the struct data type rendered, a note on the JSONValue export, a typed TypeAdapter example, with_problems compared to zod's treeifyError, and the float fill value rounding cited to zarrs and tensorstore. Assisted-by: ClaudeCode:claude-fable-5-1
…t calling repr `shown` relied on `repr` raising RecursionError for a deeply nested value. Python 3.14 limits recursion by the real C stack, so on Linux a list nested 100,000 levels deep printed fine and the problem message was the whole value. Now any value with a problem past `JSON_DEPTH` is reported as "a value nested too deep to show", on every interpreter and platform. The test nests one level past the limit, where repr always works. Assisted-by: ClaudeCode:claude-opus-5-5
…SON values `shown_by_python` caught RecursionError so that a non-JSON value nested deeper than repr could handle was described instead of raising. No parsed document can contain such a value: sets and frozensets are not JSON. Its test only passed where the platform's stack happened to overflow at the chosen depth. Values a document can hold that nest past `JSON_DEPTH` are already reported as too deep by `shown`, without repr. Assisted-by: ClaudeCode:claude-opus-5-5
- `shown` reports "too deep" only for the depth problem itself. A non-JSON value sitting exactly on the last allowed level is now printed instead of being mislabeled. - Complex fill-value rules keep each component problem's `input` and `ctx`; only the `loc` is moved under the component index. - `Schemas.defined` restores every `$defs` entry, name and use count a failed write touched, not only its own name. - The node JSON Schema docstring no longer says `consolidated_metadata` may be `null`. Both the schema and the validator refuse that. - The 382 bugfix changelog fragment is rewritten for users without naming internals. - "modelled" is spelled "modeled" in new text. Assisted-by: ClaudeCode:claude-opus-5-5
…lish Assisted-by: ClaudeCode:claude-fable-5-1
The strict readers refuse documents that some writers produced. Repair is now a separate step that a caller asks for by name: `read_repaired_node_metadata_v3` applies `repair_node_metadata_v3` to undo each known writer bug, then reads the result with `read_node_metadata_v3`, and returns the reading together with the list of repairs made. The set of repairable documents is closed. Each repair applies only when the document's members match that repair's TypedDict: `ZarrV3ZeroChunkArrayMetadataJSON` for the chunk length of 0 along an empty dimension that zarr-python 3.0 and 3.1 wrote, and `ZarrV3NullConsolidatedGroupMetadataJSON` for the `consolidated_metadata: null` that zarr-python 3.0.x wrote. Documents inside consolidated metadata are repaired too. Anything else is left for the strict read to report. Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-opus-5-5
…lish Assisted-by: ClaudeCode:claude-fable-5-1
Two scopes are equal when they hold the same definitions under the same kinds and names, regardless of construction order. Equal scopes hash alike. Before this, `Context` was not hashable at all because its tables are mapping proxies. Assisted-by: ClaudeCode:claude-fable-5-1
…Error `Claims` maps `(kind, filed name)` to the definition that read it, or None when nothing claimed it. `Conflict` records one disagreement: the key, what was claimed, what was found, and where. `ScopeConflictError` carries every conflict, as `MetadataValidationError` carries every problem. Assisted-by: ClaudeCode:claude-fable-5-1
…d for each name `claims_of` takes the fields of one reading, nested fields included, and returns the definition that read each name, keyed as the scope files it (raw bits under `r*`). A name claimed two different ways raises `ScopeConflictError`. Assisted-by: ClaudeCode:claude-fable-5-1
…s of a field `refines(field, other)` is True when `field` holds everything `other` holds: the same definition and canonical configuration where both were read, and a read field over an unclaimed one with the same name and JSON (a gain). The reverse direction is a loss, two different definitions are a conflict, and both return False. Two fields that refine each other are equal. Assisted-by: ClaudeCode:claude-fable-5-1
`Context.disagreements(claims)` reports, for each claim, whether this scope reads it the same way, would gain a definition, or conflicts with it (a different definition, or none where one was claimed). `Context.joined` returns the union of several scopes and raises `ScopeConflictError` when two of them file different definitions under one name. Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
`refines` compared a read field's written JSON to the unclaimed field's
with Python `==`. That treated `"crc32c"` and `{"name": "crc32c"}` as
different fields and `true` as equal to `1`, so the order was not
transitive across spellings. A gain is now compared the way `Unclaimed`
equality works: by name and by the configuration as JSON text. A refused
field refines only itself, so a read field that holds one still refines
itself. `Claims` is a real type at run time, not a string, so it works in
signatures. The property test now checks reflexivity, transitivity,
mutual refinement equals equality, and substitutivity, over readings of
several documents in several scopes.
Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
…und 1 found refused Aligned structs name their padding '' and may repeat it. zlib and gzip take level -1, which numcodecs writes. An empty codec id is reported, not accepted silently. create_default with a dtype and no fill value takes null. field_json_schema refuses a kind of another format with a clear error. A nested value that is no field is shown in its message. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…it; narrow blosc clevel create_default with a dtype keeps fill_value 0 for the numeric families and takes null only where 0 is no fill value. blosc clevel is 0 to 9 again; only zlib and gzip take -1. A typestr size is ASCII digits only. fields_of places a v2 struct record's type at the struct's own location. The changelog says 21 of the numcodecs ids and the create_default change. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A typestr with a listed type code and no size, such as '<f', was filed under itself and left unclaimed, so its fill value went unjudged. It is now a problem: the v2 spec requires the size. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
read_array_metadata_v2 returns a ZarrV2ArrayMetadataReading: dtype, compressor and filters as the scope read them, every problem, and the model when there is none. validate_array_metadata_v2, is_array_metadata_v2 and parse_array_metadata_v2 read in the scope given, CORE_V2 by default; the group readers take a scope too. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…scope ZarrV2ArrayMetadata(document, context=None) reads the document in a scope, CORE_V2 by default. dtype, compressor and filters are the fields as the scope read them. Equality is by what the document means. update reads new members in the model's own scope; with_context and refined_in read the document in another; refines orders models. ZarrV2ArrayMetadataUpdate replaces ZarrV2ArrayMetadataPartial. A member the spec does not define is kept. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…scope ZarrV2GroupMetadata(document, context=None) holds its document and scope like the array model. ZarrV2GroupMetadataUpdate replaces ZarrV2GroupMetadataPartial. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… each node as a model ZarrV2ConsolidatedMetadata(document, context=None) keeps its entries as written and adds nodes: each .zarray or .zgroup entry, merged with its .zattrs, as a model of the consolidated scope, keyed by node path. Equality, refines, with_context and refined_in go through the nodes. A problem in an entry is reported under that entry. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… construct The v2 pydantic field types read a document in the scope the validation context holds, CORE_V2 when it holds none. construct is gone: every model is built from its document. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…roup entries; a scope key per format zarr-python 3.x writes a consolidated_metadata member into each .zgroup entry below the root of a v2 .zmetadata, which the strict read refuses. repair_consolidated_metadata_v2 removes it and read_repaired_consolidated_metadata_v2 reads the repaired document. The pydantic field types read the v3 scope under zarr_metadata_context and the v2 scope under zarr_metadata_context_v2; a bare Context is the scope of every field type. Two keys naming one file of one consolidated node are a problem. ZarrV2NodeMetadata is exported. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…; leave the root .zgroup to the strict read A second key for one file of a v2 consolidated node no longer hides the other nodes' problems, and a leading slash names the same node. The repair leaves the root .zgroup alone, since no writer puts the member there. The fragments say how to read a zarr-python 3.x .zmetadata and which context key the v2 pydantic field types read. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…names the same node Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…document-context-stack
…tion framework The public doors keep what a user of a document, a model, a definition or a scope needs and no longer export the helpers behind them: the scope algebra's functions, the canonical-spelling and field-walking helpers, the writer-bug JSON shapes, the key sets and the JSON helpers. The README says the core depends on no validation framework and why. ruff knows a TypedDict's annotations are evaluated at run time, so the per-import noqa comments go. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… casts is_object, is_json_object, is_list_or_tuple, is_tuple and is_array in _json say what isinstance leaves unknown to a type checker, so the readers no longer cast a mapping or sequence to its element types after checking it: 239 casts become 179. The TypedDict parsers report a non-string key as a problem instead of assuming one is a string. The two loops kept for frame depth say so instead of carrying a PERF noqa. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… documents instead of casting frozen keeps an object's type; refined_object and object_at check what a read already established instead of asserting it; is_alias and is_field are type guards; field_aliases and the kind registry hold TypeAliasType. Casts in src go from 179 to 128. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…yed base Equality and hashing live once, in Keyed; each model computes its key in a method of its own class instead of a module function reading its private members. Entries given as node models are unwrapped through to_json. The remaining reportPrivateUsage sites are the readers building models through _of, which the docstrings now say. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ting to it The model's field properties go through held, which raises for a field a read refused or never read; the canonical checks and the nested-field walkers use the type guards; reading_of takes a Protocol. Casts in src go from 128 to 100. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Documentation build overview
12 files changed ·
|
…r upstream PR 4490 Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b
marked this pull request as ready for review
October 8, 2026 13:58
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #4490 +/- ##
=======================================
Coverage 94.69% 94.69%
=======================================
Files 94 94
Lines 13619 13619
=======================================
Hits 12897 12897
Misses 722 722 🚀 New features to boost your workflow:
|
A kind is declared in its class header, class CodecDefinition(Definition[C], kind=True), as typing.Protocol and SQLAlchemy's __abstract__ mark a class and not its subclasses, instead of an is_kind ClassVar that a subclass inherits but that did not count. The hook keeps the mark in the class namespace, which the slots rebuild of a dataclass carries over while the keyword is not. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…eld, UnclaimedField, RefusedField Read, Unclaimed, Refused and their union Resolved become AcceptedField, UnclaimedField, RefusedField and ResolvedField: the noun says these are metadata fields as a scope read them, and Accepted is the opposite of Refused where Read read like a verb. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…n; its check step is private A definition reads a configuration as the model readers read a document: the pair of what was read and every problem, never raising. The type-check step that only read_configuration called is _check_configuration; the public shape check stays zarr_metadata.typed_json.check. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is a big one. It adds a runtime type checker with some fancy bells and whistles for parsing zarr metadata documents correctly in different contexts.
why not use pydantic
because we don't want the dependency on pydantic. that invites conflicts with the packages that might use
zarr-metadatawith their own range supported pydantic versions. otherwise we totally would usepydantic-core, because I think it handles typeddicts correctly and it has a lot of eyes.why is it so complicated
because parsing zarr metadata documents is complicated. I'm open to simplifications but if we want:
zarrhas created its fair share!)then we have to pay for some complexity.
I will self-merge this when I'm happy with it. in parallel, I will open a draft PR that contains the integration of these new
zarr-metadatafeatures inzarr. The promise is much more correctness fromzarr. the cost is (eventual) breakage in the metadata classeszarruses today.🤖 AI text below 🤖
Every metadata field is now read against a definition in a scope, and every model is the pair of its document and the scope it was read in. A field of a document (a v3 data type, chunk grid, chunk key encoding, codec or storage transformer; a v2 dtype or numcodecs codec) is read as
Readby the definition in scope that claims its name,Unclaimedwhen nothing does, orRefusedwith every problem found.read_array_metadata_v3,read_group_metadata_v3andread_array_metadata_v2read a document once and return everything the read found, the model among it.A model is built only from a document in a scope (
CORE_AND_EXTENSIONSorCORE_V2by default) and is never invalid. Two models are equal when their documents mean the same in their scopes.updatereads new members in the model's own scope,with_contextandrefined_inread the document in another, andrefinesorders models by information. Scopes compare, join and report disagreements. JSON Schemas are exported for a TypedDict, for a field in a scope and for azarr.json. Known writer bugs are undone below the strict model byread_repaired_node_metadata_v3andread_repaired_consolidated_metadata_v2. The core depends ontyping-extensionsandannotated-typesonly; the README says why there is no validation framework, and the public surface is cut to what a consumer needs.Breaking: models are no longer dataclasses built from typed members;
construct, the...PartialTypedDicts,ZarrV3NamedConfig/ZarrV3MetadataFieldand the requiredcontextonupdateare gone; a structured v2 dtype withfill_value: 0is refused. The changelog fragments underpackages/zarr-metadata/changes/describe every change for users.History
This branch is the zarr-metadata work from fork PRs d-v-b#370 through d-v-b#398, squashed into one pull request with upstream
mainmerged in. Each fork PR carries its own description and review history. The changelog fragments are consolidated under this PR's number.🤖 Generated with Claude Code