Skip to content

A tool timeout that arrives after the task finished loses is_sdk_generated_error, so the SDK's own error message fails the output schema #5182

Description

@chrikrah

Please read this first

  • Have you read the docs? Yes.
  • Have you searched for related issues? Yes. Ten open pull requests, none touching _invoke_function_tool_with_metadata; no open or closed issue naming the late-timeout return.

Describe the bug

_invoke_function_tool_with_metadata in src/agents/tool.py returns a _FunctionToolInvocationResult from three places. Two pass the marker:

# :2240, no timeout configured
# :2250, wait_for returned normally
return _FunctionToolInvocationResult(
    output,
    is_sdk_generated_error=_consume_function_tool_default_failure(context),
)

The third, at :2256, does not:

except asyncio.TimeoutError as exc:
    if tool_task.done() and not tool_task.cancelled():
        tool_exception = tool_task.exception()
        if tool_exception is None:
            return _FunctionToolInvocationResult(tool_task.result())

asyncio.wait_for can raise TimeoutError after the tool task has already finished. This branch then returns that finished result with is_sdk_generated_error at its default False, for a string the SDK itself generated.

run_internal/tool_execution.py:2085 reads that marker as bypass_output_schema. So for a tool that carries both a timeout and a declared output_type, the SDK's own "An error occurred while running the tool." message goes through the user's schema, and the run aborts.

A fourth return at :2279 also omits the marker. That one carries the user-supplied timeout_error_function result, which is correctly not SDK-generated, so it should stay as it is.

Debug information

  • Agents SDK version: repository main at 588826c
  • Related library versions: openai 3.0.0, pydantic as pinned
  • Python version: 3.12
  • Operating system: Ubuntu 24.04
  • Model and model provider: any. The defect is in tool invocation, before the model is involved.
  • Does the issue reproduce with the latest release? Yes, on main at 588826c.
  • Does it occur consistently or intermittently? It depends on a race between the wait_for deadline and the tool task finishing. So it is intermittent in general, and consistent at a very small timeout.

Repro steps

A failing tool, invoked at four timeout settings:

timeout output is is_sdk_generated_error correct
None SDK message True yes
1e-09 SDK message False no
1e-05 SDK message False no
0.001 SDK message True yes

A timeout whose deadline has passed by the time the task finishes lands in the branch. A larger one takes the normal wait_for return, and is correct today.

End to end through Runner.run, with output_type=Weather passed explicitly:

BEFORE:  RUN ABORTED   -> UserError: Function tool output does not match its declared output schema.
AFTER:   RUN COMPLETED -> tool output reaching the model:
                          'An error occurred while running the tool. Please try again.'

One thing to know before reproducing this. A return annotation alone does not declare an output schema for a non-direct tool. @function_tool(timeout=1e-09) on async def get_weather(...) -> Weather leaves output_json_schema and _output_type_adapter at None, so nothing aborts. output_type=Weather has to be passed explicitly, which matches what tool.py's own docstring says about inference applying to programmatic tools only.

Expected behavior

A late timeout that returns the tool task's own finished result carries the same is_sdk_generated_error value the other two returns carry, so an SDK-generated error message is not validated against the user's schema and the run continues.

Offer

CONTRIBUTING.md limits pull requests to collaborators and asks non-collaborators to open an issue, so this is an issue. The fix is written and verified locally. The marker on that one return, plus one test with the other timeout tests.

The test is deterministic rather than race-dependent, because tests/README.md asks for exactly that. It says "Tests should wait for observable state transitions rather than elapsed wall-clock time", and "Use events, deterministic fakes, immediate exceptions, and narrowly scoped mocks instead of real sleeps". It patches asyncio.wait_for to await the future and then raise TimeoutError. The tool task genuinely settles, and every production decision runs for real.

20 consecutive runs: pass=20 fail=0

uv run ruff format --check   → 994 files already formatted
uv run ruff check            → All checks passed!
uv run mypy src              → Success: no issues found in 317 source files
pyright                      → 0 errors, 0 warnings, 0 informations
pytest -n 8 -m "not serial"  → 11208 passed, 29 skipped
run_serial_tests.py          → 77 passed, 4 skipped, 77 deselected
.agents/skills/code-change-verification/scripts/run.sh → all commands passed

Falsification, production file reverted and the test kept:

assert result.is_sdk_generated_error is True
AssertionError: assert False is True
 +  where False = _FunctionToolInvocationResult(output='An error occurred while running
    the tool. Please try again.', is_sdk_generated_error=False).is_sdk_generated_error

Say the word and I will open the pull request, or take the change as described.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions