Please read this first
- Have you read the docs? Yes.
- Have you searched for related issues? Yes. Ten open pull requests, none touching
_invoke_function_tool_with_metadata; no open or closed issue naming the late-timeout return.
Describe the bug
_invoke_function_tool_with_metadata in src/agents/tool.py returns a _FunctionToolInvocationResult from three places. Two pass the marker:
# :2240, no timeout configured
# :2250, wait_for returned normally
return _FunctionToolInvocationResult(
output,
is_sdk_generated_error=_consume_function_tool_default_failure(context),
)
The third, at :2256, does not:
except asyncio.TimeoutError as exc:
if tool_task.done() and not tool_task.cancelled():
tool_exception = tool_task.exception()
if tool_exception is None:
return _FunctionToolInvocationResult(tool_task.result())
asyncio.wait_for can raise TimeoutError after the tool task has already finished. This branch then returns that finished result with is_sdk_generated_error at its default False, for a string the SDK itself generated.
run_internal/tool_execution.py:2085 reads that marker as bypass_output_schema. So for a tool that carries both a timeout and a declared output_type, the SDK's own "An error occurred while running the tool." message goes through the user's schema, and the run aborts.
A fourth return at :2279 also omits the marker. That one carries the user-supplied timeout_error_function result, which is correctly not SDK-generated, so it should stay as it is.
Debug information
- Agents SDK version: repository
main at 588826c
- Related library versions:
openai 3.0.0, pydantic as pinned
- Python version: 3.12
- Operating system: Ubuntu 24.04
- Model and model provider: any. The defect is in tool invocation, before the model is involved.
- Does the issue reproduce with the latest release? Yes, on
main at 588826c.
- Does it occur consistently or intermittently? It depends on a race between the
wait_for deadline and the tool task finishing. So it is intermittent in general, and consistent at a very small timeout.
Repro steps
A failing tool, invoked at four timeout settings:
| timeout |
output is |
is_sdk_generated_error |
correct |
| None |
SDK message |
True |
yes |
| 1e-09 |
SDK message |
False |
no |
| 1e-05 |
SDK message |
False |
no |
| 0.001 |
SDK message |
True |
yes |
A timeout whose deadline has passed by the time the task finishes lands in the branch. A larger one takes the normal wait_for return, and is correct today.
End to end through Runner.run, with output_type=Weather passed explicitly:
BEFORE: RUN ABORTED -> UserError: Function tool output does not match its declared output schema.
AFTER: RUN COMPLETED -> tool output reaching the model:
'An error occurred while running the tool. Please try again.'
One thing to know before reproducing this. A return annotation alone does not declare an output schema for a non-direct tool. @function_tool(timeout=1e-09) on async def get_weather(...) -> Weather leaves output_json_schema and _output_type_adapter at None, so nothing aborts. output_type=Weather has to be passed explicitly, which matches what tool.py's own docstring says about inference applying to programmatic tools only.
Expected behavior
A late timeout that returns the tool task's own finished result carries the same is_sdk_generated_error value the other two returns carry, so an SDK-generated error message is not validated against the user's schema and the run continues.
Offer
CONTRIBUTING.md limits pull requests to collaborators and asks non-collaborators to open an issue, so this is an issue. The fix is written and verified locally. The marker on that one return, plus one test with the other timeout tests.
The test is deterministic rather than race-dependent, because tests/README.md asks for exactly that. It says "Tests should wait for observable state transitions rather than elapsed wall-clock time", and "Use events, deterministic fakes, immediate exceptions, and narrowly scoped mocks instead of real sleeps". It patches asyncio.wait_for to await the future and then raise TimeoutError. The tool task genuinely settles, and every production decision runs for real.
20 consecutive runs: pass=20 fail=0
uv run ruff format --check → 994 files already formatted
uv run ruff check → All checks passed!
uv run mypy src → Success: no issues found in 317 source files
pyright → 0 errors, 0 warnings, 0 informations
pytest -n 8 -m "not serial" → 11208 passed, 29 skipped
run_serial_tests.py → 77 passed, 4 skipped, 77 deselected
.agents/skills/code-change-verification/scripts/run.sh → all commands passed
Falsification, production file reverted and the test kept:
assert result.is_sdk_generated_error is True
AssertionError: assert False is True
+ where False = _FunctionToolInvocationResult(output='An error occurred while running
the tool. Please try again.', is_sdk_generated_error=False).is_sdk_generated_error
Say the word and I will open the pull request, or take the change as described.
Please read this first
_invoke_function_tool_with_metadata; no open or closed issue naming the late-timeout return.Describe the bug
_invoke_function_tool_with_metadatainsrc/agents/tool.pyreturns a_FunctionToolInvocationResultfrom three places. Two pass the marker:The third, at
:2256, does not:asyncio.wait_forcan raiseTimeoutErrorafter the tool task has already finished. This branch then returns that finished result withis_sdk_generated_errorat its defaultFalse, for a string the SDK itself generated.run_internal/tool_execution.py:2085reads that marker asbypass_output_schema. So for a tool that carries both atimeoutand a declaredoutput_type, the SDK's own "An error occurred while running the tool." message goes through the user's schema, and the run aborts.A fourth return at
:2279also omits the marker. That one carries the user-suppliedtimeout_error_functionresult, which is correctly not SDK-generated, so it should stay as it is.Debug information
mainat588826copenai3.0.0,pydanticas pinnedmainat588826c.wait_fordeadline and the tool task finishing. So it is intermittent in general, and consistent at a very small timeout.Repro steps
A failing tool, invoked at four timeout settings:
is_sdk_generated_errorA timeout whose deadline has passed by the time the task finishes lands in the branch. A larger one takes the normal
wait_forreturn, and is correct today.End to end through
Runner.run, withoutput_type=Weatherpassed explicitly:One thing to know before reproducing this. A return annotation alone does not declare an output schema for a non-direct tool.
@function_tool(timeout=1e-09)onasync def get_weather(...) -> Weatherleavesoutput_json_schemaand_output_type_adapteratNone, so nothing aborts.output_type=Weatherhas to be passed explicitly, which matches whattool.py's own docstring says about inference applying to programmatic tools only.Expected behavior
A late timeout that returns the tool task's own finished result carries the same
is_sdk_generated_errorvalue the other two returns carry, so an SDK-generated error message is not validated against the user's schema and the run continues.Offer
CONTRIBUTING.mdlimits pull requests to collaborators and asks non-collaborators to open an issue, so this is an issue. The fix is written and verified locally. The marker on that one return, plus one test with the other timeout tests.The test is deterministic rather than race-dependent, because
tests/README.mdasks for exactly that. It says "Tests should wait for observable state transitions rather than elapsed wall-clock time", and "Use events, deterministic fakes, immediate exceptions, and narrowly scoped mocks instead of real sleeps". It patchesasyncio.wait_forto await the future and then raiseTimeoutError. The tool task genuinely settles, and every production decision runs for real.Falsification, production file reverted and the test kept:
Say the word and I will open the pull request, or take the change as described.