You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Following #8406 / #8405, which made parser.envelopes lazy in 3.35.0, we profiled the import cost that remains in production functions using utilities.batch and utilities.parser. Powertools still adds ~325–365 ms to every cold start at 512 MB, which is 10–17% of our handlers' total import time. Almost all of it comes from four places:
utilities/parser/models/__init__.py eagerly imports all 27 model modules. Importing one model, for example SqsModel, builds pydantic schemas for ALB, Cognito, SES, Kafka, Bedrock Agent, VPC Lattice and the rest. A submodule import such as parser.models.sqs still runs the package __init__, so users can't avoid this by importing narrowly.
utilities/batch/types.py imports parser.models whenever pydantic is already loaded. This happens through has_pydantic = "pydantic" in sys.modules, and the imports are only needed to build the BatchTypeModels / BatchSqsTypeModel aliases. In any application that already uses pydantic, importing the plain PartialItemFailureResponseTypedDict, or BatchProcessor without any parser usage, loads all 27 model modules.
utilities/data_classes/__init__.py eagerly imports ~30 event modules.utilities.batch.base imports four of them (SQS, DynamoDB, Kinesis, Kafka). Through the utilities.batch package __init__, even from aws_lambda_powertools.utilities.batch.types import PartialItemFailureResponse loads 32 data_classes modules.
Tracer(disabled=True) still imports aws_xray_sdk.core.__build_config() calls _patch_xray_provider() before disabled is applied, and _disable_tracer_provider() touches aws_xray_sdk again.
Reproduction
This uses a clean install of aws-lambda-powertools[parser,tracer]==3.35.0 in public.ecr.aws/lambda/python:3.13 (arm64), with --memory=512m --cpus=0.29 to approximate Lambda's 512 MB CPU share. Each statement ran in a fresh interpreter, 20 runs each. The full script is at the end of this issue.
Statement
Median
Min
Powertools modules loaded
parser.models.*
data_classes.*
from aws_lambda_powertools.utilities.batch.types import PartialItemFailureResponse
285 ms
196 ms
84
0
32
same, in a process that already defined a pydantic model
498 ms
409 ms
118
27
32
from aws_lambda_powertools.utilities.batch import BatchProcessor (pydantic in use)
495 ms
388 ms
118
27
32
from aws_lambda_powertools.utilities.parser.models import SqsModel
390 ms
288 ms
78
27
0
from ...parser.models.sqs import SqsModel with the package __init__ bypassed (to show the potential)
16 ms
10 ms
42
1
0
Tracer(service="svc", disabled=True)
113 ms
99 ms
—
—
—
The last row imports aws_xray_sdk even though tracing is disabled. At 0.29 vCPU, CFS throttling makes medians vary by up to ~25% between runs; module counts are deterministic.
Minimal snippet:
importsys, timeimportpydanticclass_AppModel(pydantic.BaseModel): # any application that already uses pydantica: intt0=time.perf_counter()
fromaws_lambda_powertools.utilities.batch.typesimportPartialItemFailureResponseprint(f'{(time.perf_counter() -t0) *1000:.0f} ms')
print(sorted(mforminsys.modulesifm.startswith('aws_lambda_powertools.utilities.parser.models.')))
# -> 27 parser model modules loaded for a TypedDict import
The figures outside this table, including the per-cold-start totals above and the savings estimates below, come from python -X importtime on our production handlers rather than from the script. They show the same picture. utilities.parser.models accounts for ~205 ms of Powertools self time, of which ~20 ms is the four models actually used (SQS, API Gateway, DynamoDB, S3). utilities.data_classes adds another 25–50 ms.
Full reproduction script (repro.py)
"""Reproduce Powertools for AWS Lambda (Python) import-time overhead on a clean install.Run inside the official Lambda image with Lambda's 512 MB CPU share (512 / 1769 vCPU): docker run --rm --platform linux/arm64 --memory=512m --memory-swap=512m --cpus=0.29 \ -v "$PWD":/w --entrypoint /bin/bash public.ecr.aws/lambda/python:3.13 -c ' pip install -q --target /tmp/site "aws-lambda-powertools[parser,tracer]==3.35.0" && PYTHONPATH=/tmp/site python /w/repro.py'Each case runs in a fresh interpreter. The setup code runs first and is not timed; only the statement is."""importstatisticsimportsubprocessimportsysfromtypingimportFinal, NamedTupleRUNS: Final[int] =20APP_USES_PYDANTIC: Final[str] ='import pydantic\nclass AppModel(pydantic.BaseModel):\n a: int'# Not a supported usage: registers an empty `parser.models` package so only `models/sqs.py` is executed.# It shows what a single-model import would cost if `parser/models/__init__.py` were lazy.BYPASS_MODELS_INIT: Final[str] = (
'import importlib.util, types\n''spec = importlib.util.find_spec("aws_lambda_powertools.utilities.parser")\n''pkg = types.ModuleType("aws_lambda_powertools.utilities.parser.models")\n''pkg.__path__ = [spec.submodule_search_locations[0] + "/models"]\n''sys.modules[pkg.__name__] = pkg\n'
)
MEASURE_TEMPLATE: Final[str] ="""import sys, time{setup}before = set(sys.modules)t0 = time.perf_counter(){statement}elapsed_ms = (time.perf_counter() - t0) * 1000loaded = set(sys.modules) - beforeprint( elapsed_ms, sum(m.startswith("aws_lambda_powertools") for m in loaded), sum(m.startswith("aws_lambda_powertools.utilities.parser.models.") for m in loaded), sum(m.startswith("aws_lambda_powertools.utilities.data_classes.") for m in loaded), any(m.startswith("aws_xray_sdk") for m in loaded),)"""classCase(NamedTuple):
name: strsetup: strstatement: strCASES: Final[tuple[Case, ...]] = (
Case(
name='batch.types PartialItemFailureResponse, no pydantic',
setup='',
statement='from aws_lambda_powertools.utilities.batch.types import PartialItemFailureResponse',
),
Case(
name='batch.types PartialItemFailureResponse, pydantic in use',
setup=APP_USES_PYDANTIC,
statement='from aws_lambda_powertools.utilities.batch.types import PartialItemFailureResponse',
),
Case(
name='batch BatchProcessor, pydantic in use',
setup=APP_USES_PYDANTIC,
statement='from aws_lambda_powertools.utilities.batch import BatchProcessor',
),
Case(
name='parser.models SqsModel',
setup=APP_USES_PYDANTIC,
statement='from aws_lambda_powertools.utilities.parser.models import SqsModel',
),
Case(
name='parser.models.sqs SqsModel, models __init__ bypassed',
setup=APP_USES_PYDANTIC,
statement=BYPASS_MODELS_INIT+'from aws_lambda_powertools.utilities.parser.models.sqs import SqsModel',
),
Case(
name='Tracer(disabled=True)',
setup='from aws_lambda_powertools import Tracer',
statement='Tracer(service="svc", disabled=True)',
),
)
defmeasure(case: Case) ->list[str]:
"""Run one case in a fresh interpreter and return its raw measurement fields."""code: str=MEASURE_TEMPLATE.format(setup=case.setup, statement=case.statement)
result=subprocess.run([sys.executable, '-c', code], capture_output=True, text=True, check=True)
returnresult.stdout.split()
defmain() ->None:
"""Print median import time and loaded-module counts for every case."""print(f'{"case":55}{"median":>8}{"min":>7}{"pt mods":>8}{"models":>7}{"data_cls":>8}{"xray":>5}')
forcaseinCASES:
samples: list[list[str]] = [measure(case) for_inrange(RUNS)]
times: list[float] = [float(sample[0]) forsampleinsamples]
_, pt_modules, models, data_classes, xray=samples[0]
print(
f'{case.name:55}{statistics.median(times):7.0f}ms {min(times):6.0f}ms {pt_modules:>8}{models:>7}{data_classes:>8}{xray:>5}',
)
if__name__=='__main__':
main()
Which area does this relate to?
Parser, Batch processing, Event Source Data Classes, Tracer
batch/types.py: stop importing parser.models at runtime just to build type aliases. BatchTypeModels is exported from utilities.batch and used in a module-level Union in batch/base.py, so a plain TYPE_CHECKING move isn't enough. Possible approaches:
Resolve BatchTypeModels / BatchSqsTypeModel lazily through a module __getattr__, in both batch.types and batch/__init__.py.
Define BatchEventTypes under TYPE_CHECKING, since base.py already uses from __future__ import annotations.
Keep the TypedDict response types in a module that the batch package __init__ doesn't pull through base.py.
Estimated saving: ~200 ms for batch users who don't use the parser.
data_classes/__init__.py: same lazy-export pattern. Estimated saving: ~25–50 ms for every BatchProcessor user.
Tracer: when tracing is disabled, either explicitly or through POWERTOOLS_TRACE_DISABLED / environment detection, avoid resolving the X-Ray provider and skip _disable_tracer_provider(). A no-op provider would do. Estimated saving: ~50–100 ms.
Optional: a CI guard that imports utilities.batch and utilities.parser.models.SqsModel in a subprocess and asserts an upper bound on the number of aws_lambda_powertools modules loaded. This would stop the import surface from growing again.
Together, items 1–4 would save roughly 250–300 ms per cold start at 512 MB in our functions. The absolute saving is larger at lower memory sizes, where Lambda allocates less CPU.
Why is this needed?
Following #8406 / #8405, which made
parser.envelopeslazy in 3.35.0, we profiled the import cost that remains in production functions usingutilities.batchandutilities.parser. Powertools still adds ~325–365 ms to every cold start at 512 MB, which is 10–17% of our handlers' total import time. Almost all of it comes from four places:utilities/parser/models/__init__.pyeagerly imports all 27 model modules. Importing one model, for exampleSqsModel, builds pydantic schemas for ALB, Cognito, SES, Kafka, Bedrock Agent, VPC Lattice and the rest. A submodule import such asparser.models.sqsstill runs the package__init__, so users can't avoid this by importing narrowly.utilities/batch/types.pyimportsparser.modelswhenever pydantic is already loaded. This happens throughhas_pydantic = "pydantic" in sys.modules, and the imports are only needed to build theBatchTypeModels/BatchSqsTypeModelaliases. In any application that already uses pydantic, importing the plainPartialItemFailureResponseTypedDict, orBatchProcessorwithout any parser usage, loads all 27 model modules.utilities/data_classes/__init__.pyeagerly imports ~30 event modules.utilities.batch.baseimports four of them (SQS, DynamoDB, Kinesis, Kafka). Through theutilities.batchpackage__init__, evenfrom aws_lambda_powertools.utilities.batch.types import PartialItemFailureResponseloads 32data_classesmodules.Tracer(disabled=True)still importsaws_xray_sdk.core.__build_config()calls_patch_xray_provider()beforedisabledis applied, and_disable_tracer_provider()touchesaws_xray_sdkagain.Reproduction
This uses a clean install of
aws-lambda-powertools[parser,tracer]==3.35.0inpublic.ecr.aws/lambda/python:3.13(arm64), with--memory=512m --cpus=0.29to approximate Lambda's 512 MB CPU share. Each statement ran in a fresh interpreter, 20 runs each. The full script is at the end of this issue.parser.models.*data_classes.*from aws_lambda_powertools.utilities.batch.types import PartialItemFailureResponsefrom aws_lambda_powertools.utilities.batch import BatchProcessor(pydantic in use)from aws_lambda_powertools.utilities.parser.models import SqsModelfrom ...parser.models.sqs import SqsModelwith the package__init__bypassed (to show the potential)Tracer(service="svc", disabled=True)The last row imports
aws_xray_sdkeven though tracing is disabled. At 0.29 vCPU, CFS throttling makes medians vary by up to ~25% between runs; module counts are deterministic.Minimal snippet:
The figures outside this table, including the per-cold-start totals above and the savings estimates below, come from
python -X importtimeon our production handlers rather than from the script. They show the same picture.utilities.parser.modelsaccounts for ~205 ms of Powertools self time, of which ~20 ms is the four models actually used (SQS, API Gateway, DynamoDB, S3).utilities.data_classesadds another 25–50 ms.Full reproduction script (
repro.py)Which area does this relate to?
Parser, Batch processing, Event Source Data Classes, Tracer
Suggestion
parser/models/__init__.py: use the same lazy module-level__getattr__pattern as perf(parser): lazy-load envelopes to reduce Lambda cold start latency by ~900ms #8405 (importlib.import_module+globals()caching), driven by a name-to-submodule map. Keepif TYPE_CHECKING:imports so IDEs and type checkers still work, and keep__all__unchanged. Subprocess-isolated tests like the ones in perf(parser): lazy-load envelopes to reduce Lambda cold start latency by ~900ms #8405 would guard against regressions. Estimated saving: ~200–290 ms for parser users.batch/types.py: stop importingparser.modelsat runtime just to build type aliases.BatchTypeModelsis exported fromutilities.batchand used in a module-levelUnioninbatch/base.py, so a plainTYPE_CHECKINGmove isn't enough. Possible approaches:BatchTypeModels/BatchSqsTypeModellazily through a module__getattr__, in bothbatch.typesandbatch/__init__.py.BatchEventTypesunderTYPE_CHECKING, sincebase.pyalready usesfrom __future__ import annotations.TypedDictresponse types in a module that the batch package__init__doesn't pull throughbase.py.Estimated saving: ~200 ms for batch users who don't use the parser.
data_classes/__init__.py: same lazy-export pattern. Estimated saving: ~25–50 ms for everyBatchProcessoruser.Tracer: when tracing is disabled, either explicitly or throughPOWERTOOLS_TRACE_DISABLED/ environment detection, avoid resolving the X-Ray provider and skip_disable_tracer_provider(). A no-op provider would do. Estimated saving: ~50–100 ms.Optional: a CI guard that imports
utilities.batchandutilities.parser.models.SqsModelin a subprocess and asserts an upper bound on the number ofaws_lambda_powertoolsmodules loaded. This would stop the import surface from growing again.Together, items 1–4 would save roughly 250–300 ms per cold start at 512 MB in our functions. The absolute saving is larger at lower memory sizes, where Lambda allocates less CPU.
Acknowledgment