Skip to content

Responses streaming: the error event is not treated as terminal, so a teardown failure masks the provider's error #5181

Description

@chrikrah

Please read this first

  • Have you read the docs? Yes.
  • Have you searched for related issues? Yes. Ten open pull requests, none touching _TERMINAL_EVENT_TYPES; no open or closed issue naming the error stream event as non-terminal.

Describe the bug

_ResponseStreamWithRequestId._TERMINAL_EVENT_TYPES at src/agents/models/openai_responses.py:261 lists "response.error" but not "error". The SDK's error event is ResponseErrorEvent, whose type is Literal["error"]:

# openai/types/responses/response_error_event.py  (openai 3.0.0)
    type: Literal["error"]
    """The type of the event. Always `error`."""

So an error-terminated stream never sets _yielded_terminal_event. _cleanup_after_exhaustion then re-raises a transport teardown failure rather than debug-logging it. That teardown error escapes, and it masks the ModelBehaviorError built from the provider's own code and message.

The same file already holds the correct set. stream_response carries its own at :761-767, and that one lists both:

if chunk_type in {
    "response.completed",
    "response.failed",
    "response.incomplete",
    "error",
    "response.error",
}:

One line brings :261 into line with :761.

This is not a test-only path. _ResponseStreamWithRequestId is built at :916 with cleanup=lambda: api_response_cm.__aexit__(...), so the cleanup is a real httpx teardown.

Debug information

  • Agents SDK version: repository main at 588826c
  • Related library versions: openai 3.0.0 (the pin is openai>=3.0.0,<4)
  • Python version: 3.12
  • Operating system: Ubuntu 24.04
  • Model and model provider: OpenAI Responses API, any model. The trigger is the provider emitting an error event, not a particular model.
  • Does the issue reproduce with the latest release? Yes, on main at 588826c.
  • Does it occur consistently or intermittently? Consistently, whenever an error event arrives and the transport teardown then fails.

Repro steps

Driving _ResponseStreamWithRequestId with a cleanup that raises, one working event type and one that does not:

event type           terminal_seen   what __anext__ raised
------------------------------------------------------------------------------
response.failed      True            StopAsyncIteration (cleanup error suppressed)
error                False           RuntimeError: transport close failed

End to end on the real stream_response path, with the provider emitting an error event and the teardown failing:

on main:      caller sees -> RuntimeError: transport close failed
with the fix: caller sees -> ModelBehaviorError: Responses stream ended with terminal
              event `error`. code=server_error; message=the provider gave up.

The first hides the diagnosis the provider sent behind a transport artefact.

Expected behavior

An error event ends the stream, so the teardown failure is debug-logged and the ModelBehaviorError carrying the provider's code and message reaches the caller. That is what response.failed already does.

One related place that is not a second instance

A third terminal set at :1437-1442, in the websocket _stream_response, also omits "error". It is unreachable. event_type == "error" raises ResponsesWebSocketError at :1426, before that set is read. I left it alone.

Offer

CONTRIBUTING.md limits pull requests to collaborators and asks non-collaborators to open an issue, so this is an issue. The fix is written and verified locally all the same. One line in openai_responses.py, plus one parametrised test beside the existing _ResponseStreamWithRequestId cleanup tests.

uv run ruff format --check   → 994 files already formatted
uv run ruff check            → All checks passed!
uv run mypy src              → Success: no issues found in 317 source files
pyright                      → 0 errors, 0 warnings, 0 informations
pytest -n 8 -m "not serial"  → 11209 passed, 29 skipped
run_serial_tests.py          → 77 passed, 4 skipped, 77 deselected
.agents/skills/code-change-verification/scripts/run.sh → all commands passed

Baseline on unmodified main is 11207 passed, 29 skipped, so the delta is exactly the two new cases. Falsification, with the production file reverted and the test kept: the [error] case fails with RuntimeError: transport close failed, and the [response.failed] control still passes.

Say the word and I will open the pull request, or take the one line as it stands.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions