Skip to content

pydantic dataclasses: forward references leave classes unbuilt at import; the schema is built on the first construction, inside the caller's request #256

Description

@tboser

Summary

With pydantic_dataclasses on, a generated module leaves every message whose fields name a class defined later in the same module with an unbuilt pydantic schema (__pydantic_complete__ == False, __pydantic_validator__ a MockValSer). pydantic then builds the core schema and validator on the first construction of that class, in whatever code constructs it. For a wide message graph that is 10+ ms, paid inside the first request that builds the message rather than at import.

pydantic's documented pattern for a module of forward references is to call pydantic.dataclasses.rebuild_dataclass(cls) once the module is complete. The generated module never does, so every consumer has to.

Observed

envoy-data-plane 2.2.0 (generated with betterproto2-compiler 0.9.x; betterproto2 0.9.1, pydantic 2.13.4, Python 3.14; the template on main today has the same shape):

import envoy_data_plane.envoy.service.ext_proc.v3 as m
[n for n, c in vars(m).items() if isinstance(c, type) and getattr(c, "__pydantic_complete__", True) is False]
# ['BodyMutation', 'BodyResponse', 'CommonResponse', 'HeaderMutation', 'HeadersResponse', 'HttpHeaders',
#  'HttpTrailers', 'ImmediateResponse', 'ProcessingRequest', 'ProcessingResponse', 'ProtocolConfiguration',
#  'StreamedImmediateResponse', 'TrailersResponse']

First construction of ProcessingResponse(request_headers=HeadersResponse(response=CommonResponse(header_mutation=HeaderMutation(...)))): 14.6 ms; every later one 0.03 ms. In our Envoy ext_proc server that was a 10–13 ms stall on the first request of every process (measured from spans; py-spy puts the sample in pydantic/_internal/_dataclasses.py __init__).

A second, smaller lazy cost sits in the runtime: Message._betterproto (the ProtoClassMetadata) is built on a class's first serialisation / is_set — 0.4–3 ms across the same graph.

Suggestion

When pydantic_dataclasses is set, emit at the end of each generated module a rebuild pass over its message classes, e.g.

for _cls in (BodyMutation, BodyResponse, CommonResponse, ...):
    pydantic.dataclasses.rebuild_dataclass(_cls)

(or a single rebuild_dataclass call per class, or resolve the ordering so that no forward reference is left unresolved). Importing the module would then build everything; the runtime's field table could be built in the same pass (_cls._betterproto), which would also remove the second cost.

Workaround we use

At service startup:

for module in (ext_proc_v3, core_v3, type_v3):
    for cls in vars(module).values():
        if pydantic.dataclasses.is_pydantic_dataclass(cls):
            pydantic.dataclasses.rebuild_dataclass(cls)
            cls().SerializeToString()  # betterproto2's field table

Happy to open a PR against the template if the approach is welcome.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions