Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 8 additions & 6 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -244,10 +244,11 @@ up front. The demotion landed in #197; deepening LangGraph is tracked in #198.
(`python/tests/conformance/suite.py`) runs identically against every adapter. A
framework is "supported" only when it passes every scenario **or** records an
*asserted exemption* in `KNOWN_LIMITATIONS` — never a silent skip. Those
exemptions are published in the README as credibility. ADK and CrewAI cannot
propagate a per-run correlation ID or populate `donkey.last_call`, because the
framework owns the transport: LiteLLM for ADK, CrewAI's native OpenAI provider
for CrewAI. LlamaIndex and Microsoft Agent Framework have the
exemptions are published in the README as credibility. ADK's `model()` and
CrewAI cannot propagate a per-run correlation ID or populate `donkey.last_call`,
because the framework owns the transport: LiteLLM for ADK's `model()`, CrewAI's
native OpenAI provider for CrewAI. ADK's `gemini()` (a `Format=Gemini` proxy,
#691) injects the shared client, so it records no exemption. LlamaIndex and Microsoft Agent Framework have the
same two exemptions because they receive only a static `default_headers`
snapshot, which deliberately excludes the per-run correlation ID, rather than
the SDK's shared HTTP client.
Expand All @@ -257,8 +258,9 @@ matrix shrinks to LangGraph, and the deliverable becomes the **customer-facing
pytest plugin** users run against their own agent (#191).

Four adapters carry documented conformance exemptions for per-run correlation
and gateway-identity observation: ADK and CrewAI because their framework owns
the transport (LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI), plus
and gateway-identity observation: ADK's `model()` and CrewAI because their
framework owns the transport (LiteLLM for ADK, CrewAI's native OpenAI provider
for CrewAI), plus
LlamaIndex and Microsoft Agent Framework because they receive
only static headers. The conformance suite pins those exemptions to each
adapter's actual transport behavior.
Expand Down
5 changes: 4 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -383,7 +383,10 @@ This is the executable verification-discipline step for adapters' native constru
signatures (`docs/verified-apis.md` §8): does the exact class we name exist and
accept the exact kwargs we pass, against the framework as actually installed.
`--live` adds one real completion round-trip (needs the three
`DONKEY_LLM_PROXY_*` env vars); `--only <fw>` restricts scope;
`DONKEY_LLM_PROXY_*` env vars). The `adk.gemini` row needs a `Format=Gemini`
proxy, so its live check runs only when `DONKEY_GEMINI_PROXY_URL` is also set,
with a credential pair contracted on that proxy; otherwise it is skipped.
`--only <fw>` restricts scope (`--only adk` covers `adk` and `adk.gemini`);
`--emit-verified` prints `docs/verified-apis.md §8` markdown rows after maintainer sign-off. A
`_verify.blocked(...)`-guarded adapter correctly shows as `BLOCKED (verification discipline)`, not a
failure — don't "fix" the script to make a genuinely-blocked adapter pass.
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,7 @@ legitimately cannot satisfy a scenario, the reason is asserted in code

| Framework | Scenario | Why it's exempt |
| --- | --- | --- |
| ADK, CrewAI | correlation ID propagated | The framework owns the transport — LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI — so the SDK's `httpx` client cannot be injected and the correlation ID ends up per-client, not per-run. For ADK, a LiteLLM logger callback may recover trace correlation later. |
| ADK `model()`, CrewAI | correlation ID propagated | The framework owns the transport — LiteLLM for ADK's `model()`, CrewAI's native OpenAI provider for CrewAI — so the SDK's `httpx` client cannot be injected and the correlation ID ends up per-client, not per-run. For ADK, a LiteLLM logger callback may recover trace correlation later. ADK's `gemini()` (`Format=Gemini` proxy) injects the shared client and records no exemption (#691). |
| LlamaIndex, Microsoft Agent Framework | correlation ID propagated | These adapters receive a static `default_headers` snapshot, which deliberately excludes the per-run correlation ID. Without the SDK's `httpx` client, `donkey.run(id=...)` cannot update their request headers. |
| ADK, CrewAI | gateway identity observed | The framework owns the transport (LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI), so no response reaches the SDK's `_on_response` hook. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. |
| ADK `model()`, CrewAI | gateway identity observed | The framework owns the transport (LiteLLM for ADK's `model()`, CrewAI's native OpenAI provider for CrewAI), so no response reaches the SDK's `_on_response` hook. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. |
| LlamaIndex, Microsoft Agent Framework | gateway identity observed | These adapters receive `default_headers`, not the SDK's `httpx` client, so no response reaches `_on_response`. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. |
3 changes: 2 additions & 1 deletion docs/verified-apis.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,7 @@ request against the deployed gateway.
|---|---|---|---|---|---|
| Ingress base URL the SDK targets | `llm/client.py` | VERIFIED (LIVE) | `https://<ingress-gw-host>/<instance-path>/` — e.g. `https://agent-network-ingress-gw-zovwbn.jeg62f.usa-e2.cloudhub.io/openai-sdk/`. **No `/v1` at the ingress**; OpenAI path segment appended directly (`/responses`, and OpenAI-native routes) | 2026-08-28 | live probe |
| Ingress **Format** — Anthropic-native route (#304) | `integrations/anthropic.py`, `llm/client.py` | VERIFIED (LIVE) | A proxy provisioned `Format=Anthropic` exposes a **native Anthropic Messages ingress** at `POST /<base-path>/v1/messages`. The Anthropic request shape (header `anthropic-version: 2023-06-01`, body `{"model":…,"max_tokens":N,"messages":[{"role":"user","content":…}]}`) → **200** with a native Anthropic body (`type:"message"`, `role:"assistant"`, `content[].text`, `stop_reason:"end_turn"`, `usage.input_tokens`/`output_tokens`/`service_tier`) plus native `anthropic-ratelimit-*` / `request-id` / `anthropic-organization-id` response headers. It is a **passthrough, not a transcode** — gateway headers `x-llm-proxy-llm-provider: anthropic`, `x-llm-proxy-llm-model: <model>`, `x-llm-proxy-request-success` (`routingType` stays `model-based`; enabled by the per-upstream `routing[].upstreams[].llmConfigs.format: anthropic` field). An **OpenAI-shaped** request to `/<base-path>/chat/completions` → **404** (empty body) — the OpenAI route is simply not served on a native-Anthropic ingress (differs from Gemini's 400; both prove the ingress rejects the OpenAI wire format). Same `client_id`/`client_secret` CIE auth (no `client_id` → 401 "Client ID is not present", `www-authenticate: Client-ID-Enforcement`). Native ingress is **single-route only** (no multi-routing/fallback — that stays OpenAI-only); the upstream must speak native Anthropic (`provider: anthropic`/`api.anthropic.com` — **Bedrock is not usable** here). The `donkey.anthropic` adapter targets this route, but the SDK's own DDK proxies are `Format=OpenAI` — point `base_url` at a `Format=Anthropic` proxy to use the native surface. The roadmap "OAuth Client ID/Secret for configuring LLMs" is an **upstream/control-plane credential** concern and does **not** change this verified consumer `client_id`/`client_secret` request-header pair. | 2026-09-24 | live probe against `ddk-anthropic-inbound` (instance `21194086`, DDK/Sandbox) — captured in `python/tests/fixtures/anypoint/anthropic_inbound/` |
| Ingress **Format** — Gemini-native route (#540) | `llm/client.py` (no SDK adapter — see note) | VERIFIED (LIVE) | A proxy provisioned `Format=Gemini` exposes a **native Gemini ingress** at `POST /<base-path>/models/<model>:generateContent`. The Gemini request shape (`{"contents":[{"role":"user","parts":[{"text":…}]}]}`) → **200** with a native Gemini body (`candidates[].content.parts[].text`, `role:"model"`, `usageMetadata` incl. `thoughtsTokenCount`, `modelVersion`, `responseId`). It is a **passthrough, not a transcode** — response header `x-llm-proxy-model-based-routing-success: "Request passed through without model-based routing."` (`routingType` stays `model-based`; enabled by the per-upstream `routing[].upstreams[].llmConfigs.format: gemini` field). An **OpenAI-shaped** request to `/<base-path>/chat/completions` → **400** with a native-Gemini error envelope (a JSON *array* `[{"error":{"code":400,"status":"INVALID_ARGUMENT",…}}]`), confirming the ingress does not accept the OpenAI wire format. Same `client_id`/`client_secret` CIE auth (no `client_id` → 401 "Client ID is not present"). Native ingress is **single-route only** (no multi-routing/fallback — that stays OpenAI-only). **No SDK adapter ships** for this: a `google-genai` adapter is demand-driven (#540/#244, "do not guess it"); Gemini also remains reachable as an *upstream provider* behind an OpenAI ingress (supported-providers row below). | 2026-09-23 | live probe against `ddk-gemini-inbound` (instance `21193369`, DDK/Sandbox) — captured in `python/tests/fixtures/anypoint/gemini_inbound/` |
| Ingress **Format** — Gemini-native route (#540, #691) | `integrations/adk.py` (`gemini()`, #691); `core/transport.py` (model from the URL path) + `core/lastcall.py` (`usageMetadata`) | VERIFIED (LIVE) | A proxy provisioned `Format=Gemini` exposes a **native Gemini ingress** at `POST /<base-path>/models/<model>:generateContent`. The Gemini request shape (`{"contents":[{"role":"user","parts":[{"text":…}]}]}`) → **200** with a native Gemini body (`candidates[].content.parts[].text`, `role:"model"`, `usageMetadata` incl. `thoughtsTokenCount`, `modelVersion`, `responseId`). It is a **passthrough, not a transcode** — response header `x-llm-proxy-model-based-routing-success: "Request passed through without model-based routing."` (`routingType` stays `model-based`; enabled by the per-upstream `routing[].upstreams[].llmConfigs.format: gemini` field). An **OpenAI-shaped** request to `/<base-path>/chat/completions` → **400** with a native-Gemini error envelope (a JSON *array* `[{"error":{"code":400,"status":"INVALID_ARGUMENT",…}}]`), confirming the ingress does not accept the OpenAI wire format. Same `client_id`/`client_secret` CIE auth (no `client_id` → 401 "Client ID is not present"). Native ingress is **single-route only** (no multi-routing/fallback — that stays OpenAI-only). **Re-probed 2026-09-29 for the ADK native adapter (#691):** (a) `google-genai` always sends the API key as `x-goog-api-key`; a non-empty placeholder alongside `client_id`/`client_secret` does not clash → 200; the gateway echoes `x-correlation-id`. (b) `POST /<base-path>/models/<model>:streamGenerateContent?alt=sse` is routed → `text/event-stream`, each event a Gemini chunk with **cumulative** `usageMetadata` (last event is the total). (c) A `functionCall`/`functionResponse` round-trip works. (d) The gateway **ignores a body `model` field** — absent, `gemini-2.5-flash` and `gemini/gemini-2.5-flash` all → 200 with the same passthrough header; the URL path picks the model, so the SDK reads it from there for `last_call`/spans. No `x-llm-proxy-llm-model`/`-provider` header is emitted, so `served_model`/`served_provider` stay `None`. (e) Refusals: bad secret → **401** (`classify` → `AuthError`); unknown model → Google's **404** `NOT_FOUND` passed through (`classify` → `UpstreamRequestError`). `usageMetadata.totalTokenCount` **includes** `thoughtsTokenCount`. Gemini also remains reachable as an *upstream provider* behind an OpenAI ingress (supported-providers row below). | 2026-09-29 | live probes against `ddk-gemini-inbound` (instance `21193369`, DDK/Sandbox) on 2026-09-23 and 2026-09-29 — captured in `python/tests/fixtures/anypoint/gemini_inbound/` |
| Endpoint/API surface | `llm/client.py` | VERIFIED (LIVE) | OpenAI **Responses API** works (`POST /openai-sdk/responses`, body `{model, input}`). Upstream registered as `https://api.openai.com/v1/`; `proxyUri http://0.0.0.0:8081/openai-sdk` | 2026-08-28 | `api:describe`, live probe |
| Auth: header name / model | `core/transport.py`, `core/config.py` | VERIFIED (LIVE) | **`client_id` + `client_secret` request headers** (NOT bearer) — enforced by `client-id-enforcement` 1.3.3. A consumer **credential pair**, mapped to an Anypoint client application | 2026-08-28 | live probe + `policy:list` |
| Auth: model-wallet ingress (alt to `client_id`/`client_secret`, #372) | `core/transport.py`, `core/config.py` — contract verified; SDK wiring landed in #509 (`llm_proxy_auth="jwt"` + `llm_proxy_wallet_client_id` + `Donkey(llm_auth=…)`) | VERIFIED (LIVE) | Wallet-backed model proxies auth via an **IdP-issued JWT + a client ID**, with **no `client_secret`**: the client ID travels as the **`X-Client-Id`** request header (exact casing confirmed; value = the wallet's system-generated `clientId`, `ddk-model-wallet` here — read it from the omni API response, don't assume) which selects the wallet (echoed back as response header **`x-model-wallet-selected`**), *and* as a `client_id` **JWT claim** read by the LLM Proxy Core Policy at `#[authentication.properties.claims.client_id]` after the **JWT Validation** policy validates + publishes claims. **The JWT rides as `Authorization: Bearer <JWT>`** — this was an *assumption* from the doc read (Bearer is nowhere quoted on the source page); now **confirmed** (policy `jwtOrigin: httpBearerAuthenticationHeader`; missing header → `400 {"error":"JWT Token is required."}`, `www-authenticate: Bearer`; invalid/expired token → `401 {"error":"Invalid token."}`). No `client_id`/`client_secret` request headers are sent and the call still succeeds → the default **DataWeave Headers Transformation** + **Client ID Enforcement** policies are confirmed **disabled**. Requires an org **IdP** (JWKS URL / signing key — "orgs without an IdP can't use model wallets"). A **parallel** ingress model, NOT a replacement of the LIVE `client_id`/`client_secret` pair above. | 2026-09-21 | live probe — `tests/fixtures/anypoint/model_wallet/` (`ddk-model-wallet`, instance `21186246`, Sandbox); shape corroborates `docs.mulesoft.com/general/exp-model-wallets-manage` |
Expand Down Expand Up @@ -495,6 +495,7 @@ regardless of which issue/PR produced the offline pass.
|---|---|---|---|---|---|
| LangGraph | `langchain_openai.ChatOpenAI(model, base_url, api_key, default_headers, http_async_client, max_retries=0, use_responses_api=True)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `langgraph==1.2.12`, `langchain-openai==1.6.6`, `langchain-core==1.6.5`. The deep adapter (#198) retains `use_responses_api=True` for the live-verified `/responses` data plane (§2); overridable via `chat_model(..., use_responses_api=False)`. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only langgraph` (also `--emit-verified`) |
| Google ADK | `google.adk.models.lite_llm.LiteLlm(model="openai/…", api_base, api_key, extra_headers)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `google-adk==2.10.0`, `litellm==1.102.1`. LiteLLM owns the transport; this does not verify request/header forwarding. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only adk` (also `--emit-verified`) |
| Google ADK (native Gemini) | `google.adk.models.Gemini(model, base_url, client_kwargs={api_key, http_options={base_url, api_version="", headers, timeout, httpx_async_client}})` | VERIFIED (LIVE) | Round-trip through a `Format=Gemini` proxy with `google-adk==2.10.0`, `google-genai==2.25.0`: `generate_content_async` → 200, SSE streaming, and a tool-calling agent run. `client_kwargs` **replaces** ADK's default `http_options` wholesale, so the adapter passes the full set; an injected `httpx_async_client` disables genai's aiohttp path, so buffered and streamed calls both go through the shared `DonkeyAsyncClient`. `api_key` is required by genai and sent as `x-goog-api-key` (gateway-ignored placeholder). `timeout` is milliseconds (`None` would disable it). genai raises its own `google.genai.errors.APIError` whose `.response` is the httpx response, so `classify(exc.response)` works. ADK caches one genai `Client` per event loop, but the shared httpx pool is bound to the first loop — one Donkey per event loop. | 2026-09-29 | #691; `python/scripts/verify_frameworks.py --only adk.gemini --live` against `ddk-gemini-inbound` |
| MS Agent Framework | `agent_framework.openai.OpenAIChatClient(model, base_url, api_key, default_headers)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `agent-framework==1.19.0`. The constructor kwarg is `model=` — `model_id` is **not** accepted. `base_url`/`api_key`/`default_headers` all accepted, so `connection_kwargs()` is unchanged. No live round-trip. (Reclassified from a bare `VERIFIED` per #681 — same offline-only evidence class as its peer rows; §0.3 reserves `VERIFIED` for a real-sandbox round-trip.) | 2026-09-22 | #520 (`verify_frameworks.py --only agent_framework`) |
| OpenAI Agents SDK | `agents.OpenAIChatCompletionsModel(model, openai_client=openai.AsyncOpenAI(base_url, api_key, default_headers, http_client, max_retries=0))` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `openai-agents==0.20.0`. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only openai_agents` (also `--emit-verified`) |
| Anthropic SDK | `anthropic.AsyncAnthropic(base_url, api_key, default_headers, http_client, max_retries=0)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `anthropic==0.116.0`. Model id remains per-call; the separately live-verified Anthropic Messages ingress (§2, #304) requires a `Format=Anthropic` proxy. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only anthropic` (also `--emit-verified`) |
Expand Down
Loading
Loading