diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 2df77c00..08cbf392 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -244,10 +244,11 @@ up front. The demotion landed in #197; deepening LangGraph is tracked in #198. (`python/tests/conformance/suite.py`) runs identically against every adapter. A framework is "supported" only when it passes every scenario **or** records an *asserted exemption* in `KNOWN_LIMITATIONS` — never a silent skip. Those -exemptions are published in the README as credibility. ADK and CrewAI cannot -propagate a per-run correlation ID or populate `donkey.last_call`, because the -framework owns the transport: LiteLLM for ADK, CrewAI's native OpenAI provider -for CrewAI. LlamaIndex and Microsoft Agent Framework have the +exemptions are published in the README as credibility. ADK's `model()` and +CrewAI cannot propagate a per-run correlation ID or populate `donkey.last_call`, +because the framework owns the transport: LiteLLM for ADK's `model()`, CrewAI's +native OpenAI provider for CrewAI. ADK's `gemini()` (a `Format=Gemini` proxy, +#691) injects the shared client, so it records no exemption. LlamaIndex and Microsoft Agent Framework have the same two exemptions because they receive only a static `default_headers` snapshot, which deliberately excludes the per-run correlation ID, rather than the SDK's shared HTTP client. @@ -257,8 +258,9 @@ matrix shrinks to LangGraph, and the deliverable becomes the **customer-facing pytest plugin** users run against their own agent (#191). Four adapters carry documented conformance exemptions for per-run correlation -and gateway-identity observation: ADK and CrewAI because their framework owns -the transport (LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI), plus +and gateway-identity observation: ADK's `model()` and CrewAI because their +framework owns the transport (LiteLLM for ADK, CrewAI's native OpenAI provider +for CrewAI), plus LlamaIndex and Microsoft Agent Framework because they receive only static headers. The conformance suite pins those exemptions to each adapter's actual transport behavior. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index bc90212d..0b0164e7 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -383,7 +383,10 @@ This is the executable verification-discipline step for adapters' native constru signatures (`docs/verified-apis.md` §8): does the exact class we name exist and accept the exact kwargs we pass, against the framework as actually installed. `--live` adds one real completion round-trip (needs the three -`DONKEY_LLM_PROXY_*` env vars); `--only ` restricts scope; +`DONKEY_LLM_PROXY_*` env vars). The `adk.gemini` row needs a `Format=Gemini` +proxy, so its live check runs only when `DONKEY_GEMINI_PROXY_URL` is also set, +with a credential pair contracted on that proxy; otherwise it is skipped. +`--only ` restricts scope (`--only adk` covers `adk` and `adk.gemini`); `--emit-verified` prints `docs/verified-apis.md §8` markdown rows after maintainer sign-off. A `_verify.blocked(...)`-guarded adapter correctly shows as `BLOCKED (verification discipline)`, not a failure — don't "fix" the script to make a genuinely-blocked adapter pass. diff --git a/README.md b/README.md index a18f0147..c1765114 100644 --- a/README.md +++ b/README.md @@ -131,7 +131,7 @@ legitimately cannot satisfy a scenario, the reason is asserted in code | Framework | Scenario | Why it's exempt | | --- | --- | --- | -| ADK, CrewAI | correlation ID propagated | The framework owns the transport — LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI — so the SDK's `httpx` client cannot be injected and the correlation ID ends up per-client, not per-run. For ADK, a LiteLLM logger callback may recover trace correlation later. | +| ADK `model()`, CrewAI | correlation ID propagated | The framework owns the transport — LiteLLM for ADK's `model()`, CrewAI's native OpenAI provider for CrewAI — so the SDK's `httpx` client cannot be injected and the correlation ID ends up per-client, not per-run. For ADK, a LiteLLM logger callback may recover trace correlation later. ADK's `gemini()` (`Format=Gemini` proxy) injects the shared client and records no exemption (#691). | | LlamaIndex, Microsoft Agent Framework | correlation ID propagated | These adapters receive a static `default_headers` snapshot, which deliberately excludes the per-run correlation ID. Without the SDK's `httpx` client, `donkey.run(id=...)` cannot update their request headers. | -| ADK, CrewAI | gateway identity observed | The framework owns the transport (LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI), so no response reaches the SDK's `_on_response` hook. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. | +| ADK `model()`, CrewAI | gateway identity observed | The framework owns the transport (LiteLLM for ADK's `model()`, CrewAI's native OpenAI provider for CrewAI), so no response reaches the SDK's `_on_response` hook. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. | | LlamaIndex, Microsoft Agent Framework | gateway identity observed | These adapters receive `default_headers`, not the SDK's `httpx` client, so no response reaches `_on_response`. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. | diff --git a/docs/verified-apis.md b/docs/verified-apis.md index d1f1bf63..3a7c35ea 100644 --- a/docs/verified-apis.md +++ b/docs/verified-apis.md @@ -96,7 +96,7 @@ request against the deployed gateway. |---|---|---|---|---|---| | Ingress base URL the SDK targets | `llm/client.py` | VERIFIED (LIVE) | `https:////` — e.g. `https://agent-network-ingress-gw-zovwbn.jeg62f.usa-e2.cloudhub.io/openai-sdk/`. **No `/v1` at the ingress**; OpenAI path segment appended directly (`/responses`, and OpenAI-native routes) | 2026-08-28 | live probe | | Ingress **Format** — Anthropic-native route (#304) | `integrations/anthropic.py`, `llm/client.py` | VERIFIED (LIVE) | A proxy provisioned `Format=Anthropic` exposes a **native Anthropic Messages ingress** at `POST //v1/messages`. The Anthropic request shape (header `anthropic-version: 2023-06-01`, body `{"model":…,"max_tokens":N,"messages":[{"role":"user","content":…}]}`) → **200** with a native Anthropic body (`type:"message"`, `role:"assistant"`, `content[].text`, `stop_reason:"end_turn"`, `usage.input_tokens`/`output_tokens`/`service_tier`) plus native `anthropic-ratelimit-*` / `request-id` / `anthropic-organization-id` response headers. It is a **passthrough, not a transcode** — gateway headers `x-llm-proxy-llm-provider: anthropic`, `x-llm-proxy-llm-model: `, `x-llm-proxy-request-success` (`routingType` stays `model-based`; enabled by the per-upstream `routing[].upstreams[].llmConfigs.format: anthropic` field). An **OpenAI-shaped** request to `//chat/completions` → **404** (empty body) — the OpenAI route is simply not served on a native-Anthropic ingress (differs from Gemini's 400; both prove the ingress rejects the OpenAI wire format). Same `client_id`/`client_secret` CIE auth (no `client_id` → 401 "Client ID is not present", `www-authenticate: Client-ID-Enforcement`). Native ingress is **single-route only** (no multi-routing/fallback — that stays OpenAI-only); the upstream must speak native Anthropic (`provider: anthropic`/`api.anthropic.com` — **Bedrock is not usable** here). The `donkey.anthropic` adapter targets this route, but the SDK's own DDK proxies are `Format=OpenAI` — point `base_url` at a `Format=Anthropic` proxy to use the native surface. The roadmap "OAuth Client ID/Secret for configuring LLMs" is an **upstream/control-plane credential** concern and does **not** change this verified consumer `client_id`/`client_secret` request-header pair. | 2026-09-24 | live probe against `ddk-anthropic-inbound` (instance `21194086`, DDK/Sandbox) — captured in `python/tests/fixtures/anypoint/anthropic_inbound/` | -| Ingress **Format** — Gemini-native route (#540) | `llm/client.py` (no SDK adapter — see note) | VERIFIED (LIVE) | A proxy provisioned `Format=Gemini` exposes a **native Gemini ingress** at `POST //models/:generateContent`. The Gemini request shape (`{"contents":[{"role":"user","parts":[{"text":…}]}]}`) → **200** with a native Gemini body (`candidates[].content.parts[].text`, `role:"model"`, `usageMetadata` incl. `thoughtsTokenCount`, `modelVersion`, `responseId`). It is a **passthrough, not a transcode** — response header `x-llm-proxy-model-based-routing-success: "Request passed through without model-based routing."` (`routingType` stays `model-based`; enabled by the per-upstream `routing[].upstreams[].llmConfigs.format: gemini` field). An **OpenAI-shaped** request to `//chat/completions` → **400** with a native-Gemini error envelope (a JSON *array* `[{"error":{"code":400,"status":"INVALID_ARGUMENT",…}}]`), confirming the ingress does not accept the OpenAI wire format. Same `client_id`/`client_secret` CIE auth (no `client_id` → 401 "Client ID is not present"). Native ingress is **single-route only** (no multi-routing/fallback — that stays OpenAI-only). **No SDK adapter ships** for this: a `google-genai` adapter is demand-driven (#540/#244, "do not guess it"); Gemini also remains reachable as an *upstream provider* behind an OpenAI ingress (supported-providers row below). | 2026-09-23 | live probe against `ddk-gemini-inbound` (instance `21193369`, DDK/Sandbox) — captured in `python/tests/fixtures/anypoint/gemini_inbound/` | +| Ingress **Format** — Gemini-native route (#540, #691) | `integrations/adk.py` (`gemini()`, #691); `core/transport.py` (model from the URL path) + `core/lastcall.py` (`usageMetadata`) | VERIFIED (LIVE) | A proxy provisioned `Format=Gemini` exposes a **native Gemini ingress** at `POST //models/:generateContent`. The Gemini request shape (`{"contents":[{"role":"user","parts":[{"text":…}]}]}`) → **200** with a native Gemini body (`candidates[].content.parts[].text`, `role:"model"`, `usageMetadata` incl. `thoughtsTokenCount`, `modelVersion`, `responseId`). It is a **passthrough, not a transcode** — response header `x-llm-proxy-model-based-routing-success: "Request passed through without model-based routing."` (`routingType` stays `model-based`; enabled by the per-upstream `routing[].upstreams[].llmConfigs.format: gemini` field). An **OpenAI-shaped** request to `//chat/completions` → **400** with a native-Gemini error envelope (a JSON *array* `[{"error":{"code":400,"status":"INVALID_ARGUMENT",…}}]`), confirming the ingress does not accept the OpenAI wire format. Same `client_id`/`client_secret` CIE auth (no `client_id` → 401 "Client ID is not present"). Native ingress is **single-route only** (no multi-routing/fallback — that stays OpenAI-only). **Re-probed 2026-09-29 for the ADK native adapter (#691):** (a) `google-genai` always sends the API key as `x-goog-api-key`; a non-empty placeholder alongside `client_id`/`client_secret` does not clash → 200; the gateway echoes `x-correlation-id`. (b) `POST //models/:streamGenerateContent?alt=sse` is routed → `text/event-stream`, each event a Gemini chunk with **cumulative** `usageMetadata` (last event is the total). (c) A `functionCall`/`functionResponse` round-trip works. (d) The gateway **ignores a body `model` field** — absent, `gemini-2.5-flash` and `gemini/gemini-2.5-flash` all → 200 with the same passthrough header; the URL path picks the model, so the SDK reads it from there for `last_call`/spans. No `x-llm-proxy-llm-model`/`-provider` header is emitted, so `served_model`/`served_provider` stay `None`. (e) Refusals: bad secret → **401** (`classify` → `AuthError`); unknown model → Google's **404** `NOT_FOUND` passed through (`classify` → `UpstreamRequestError`). `usageMetadata.totalTokenCount` **includes** `thoughtsTokenCount`. Gemini also remains reachable as an *upstream provider* behind an OpenAI ingress (supported-providers row below). | 2026-09-29 | live probes against `ddk-gemini-inbound` (instance `21193369`, DDK/Sandbox) on 2026-09-23 and 2026-09-29 — captured in `python/tests/fixtures/anypoint/gemini_inbound/` | | Endpoint/API surface | `llm/client.py` | VERIFIED (LIVE) | OpenAI **Responses API** works (`POST /openai-sdk/responses`, body `{model, input}`). Upstream registered as `https://api.openai.com/v1/`; `proxyUri http://0.0.0.0:8081/openai-sdk` | 2026-08-28 | `api:describe`, live probe | | Auth: header name / model | `core/transport.py`, `core/config.py` | VERIFIED (LIVE) | **`client_id` + `client_secret` request headers** (NOT bearer) — enforced by `client-id-enforcement` 1.3.3. A consumer **credential pair**, mapped to an Anypoint client application | 2026-08-28 | live probe + `policy:list` | | Auth: model-wallet ingress (alt to `client_id`/`client_secret`, #372) | `core/transport.py`, `core/config.py` — contract verified; SDK wiring landed in #509 (`llm_proxy_auth="jwt"` + `llm_proxy_wallet_client_id` + `Donkey(llm_auth=…)`) | VERIFIED (LIVE) | Wallet-backed model proxies auth via an **IdP-issued JWT + a client ID**, with **no `client_secret`**: the client ID travels as the **`X-Client-Id`** request header (exact casing confirmed; value = the wallet's system-generated `clientId`, `ddk-model-wallet` here — read it from the omni API response, don't assume) which selects the wallet (echoed back as response header **`x-model-wallet-selected`**), *and* as a `client_id` **JWT claim** read by the LLM Proxy Core Policy at `#[authentication.properties.claims.client_id]` after the **JWT Validation** policy validates + publishes claims. **The JWT rides as `Authorization: Bearer `** — this was an *assumption* from the doc read (Bearer is nowhere quoted on the source page); now **confirmed** (policy `jwtOrigin: httpBearerAuthenticationHeader`; missing header → `400 {"error":"JWT Token is required."}`, `www-authenticate: Bearer`; invalid/expired token → `401 {"error":"Invalid token."}`). No `client_id`/`client_secret` request headers are sent and the call still succeeds → the default **DataWeave Headers Transformation** + **Client ID Enforcement** policies are confirmed **disabled**. Requires an org **IdP** (JWKS URL / signing key — "orgs without an IdP can't use model wallets"). A **parallel** ingress model, NOT a replacement of the LIVE `client_id`/`client_secret` pair above. | 2026-09-21 | live probe — `tests/fixtures/anypoint/model_wallet/` (`ddk-model-wallet`, instance `21186246`, Sandbox); shape corroborates `docs.mulesoft.com/general/exp-model-wallets-manage` | @@ -495,6 +495,7 @@ regardless of which issue/PR produced the offline pass. |---|---|---|---|---|---| | LangGraph | `langchain_openai.ChatOpenAI(model, base_url, api_key, default_headers, http_async_client, max_retries=0, use_responses_api=True)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `langgraph==1.2.12`, `langchain-openai==1.6.6`, `langchain-core==1.6.5`. The deep adapter (#198) retains `use_responses_api=True` for the live-verified `/responses` data plane (§2); overridable via `chat_model(..., use_responses_api=False)`. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only langgraph` (also `--emit-verified`) | | Google ADK | `google.adk.models.lite_llm.LiteLlm(model="openai/…", api_base, api_key, extra_headers)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `google-adk==2.10.0`, `litellm==1.102.1`. LiteLLM owns the transport; this does not verify request/header forwarding. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only adk` (also `--emit-verified`) | +| Google ADK (native Gemini) | `google.adk.models.Gemini(model, base_url, client_kwargs={api_key, http_options={base_url, api_version="", headers, timeout, httpx_async_client}})` | VERIFIED (LIVE) | Round-trip through a `Format=Gemini` proxy with `google-adk==2.10.0`, `google-genai==2.25.0`: `generate_content_async` → 200, SSE streaming, and a tool-calling agent run. `client_kwargs` **replaces** ADK's default `http_options` wholesale, so the adapter passes the full set; an injected `httpx_async_client` disables genai's aiohttp path, so buffered and streamed calls both go through the shared `DonkeyAsyncClient`. `api_key` is required by genai and sent as `x-goog-api-key` (gateway-ignored placeholder). `timeout` is milliseconds (`None` would disable it). genai raises its own `google.genai.errors.APIError` whose `.response` is the httpx response, so `classify(exc.response)` works. ADK caches one genai `Client` per event loop, but the shared httpx pool is bound to the first loop — one Donkey per event loop. | 2026-09-29 | #691; `python/scripts/verify_frameworks.py --only adk.gemini --live` against `ddk-gemini-inbound` | | MS Agent Framework | `agent_framework.openai.OpenAIChatClient(model, base_url, api_key, default_headers)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `agent-framework==1.19.0`. The constructor kwarg is `model=` — `model_id` is **not** accepted. `base_url`/`api_key`/`default_headers` all accepted, so `connection_kwargs()` is unchanged. No live round-trip. (Reclassified from a bare `VERIFIED` per #681 — same offline-only evidence class as its peer rows; §0.3 reserves `VERIFIED` for a real-sandbox round-trip.) | 2026-09-22 | #520 (`verify_frameworks.py --only agent_framework`) | | OpenAI Agents SDK | `agents.OpenAIChatCompletionsModel(model, openai_client=openai.AsyncOpenAI(base_url, api_key, default_headers, http_client, max_retries=0))` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `openai-agents==0.20.0`. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only openai_agents` (also `--emit-verified`) | | Anthropic SDK | `anthropic.AsyncAnthropic(base_url, api_key, default_headers, http_client, max_retries=0)` | UNVERIFIED (signature confirmed; pending maintainer `--live` + sign-off) | Construction and `isinstance` against the recorded path confirmed offline with `anthropic==0.116.0`. Model id remains per-call; the separately live-verified Anthropic Messages ingress (§2, #304) requires a `Format=Anthropic` proxy. Tested with `openai==2.54.0`; no live round-trip. | 2026-09-27 | #34; `python/scripts/verify_frameworks.py --only anthropic` (also `--emit-verified`) | diff --git a/python/README.md b/python/README.md index 50e68237..572ff3d9 100644 --- a/python/README.md +++ b/python/README.md @@ -131,7 +131,7 @@ legitimately cannot satisfy a scenario, the reason is asserted in code | Framework | Scenario | Why it's exempt | | --- | --- | --- | -| ADK, CrewAI | correlation ID propagated | The framework owns the transport — LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI — so the SDK's `httpx` client cannot be injected and the correlation ID ends up per-client, not per-run. For ADK, a LiteLLM logger callback may recover trace correlation later. | +| ADK `model()`, CrewAI | correlation ID propagated | The framework owns the transport — LiteLLM for ADK's `model()`, CrewAI's native OpenAI provider for CrewAI — so the SDK's `httpx` client cannot be injected and the correlation ID ends up per-client, not per-run. For ADK, a LiteLLM logger callback may recover trace correlation later. ADK's `gemini()` (`Format=Gemini` proxy) injects the shared client and records no exemption (#691). | | LlamaIndex, Microsoft Agent Framework | correlation ID propagated | These adapters receive a static `default_headers` snapshot, which deliberately excludes the per-run correlation ID. Without the SDK's `httpx` client, `donkey.run(id=...)` cannot update their request headers. | -| ADK, CrewAI | gateway identity observed | The framework owns the transport (LiteLLM for ADK, CrewAI's native OpenAI provider for CrewAI), so no response reaches the SDK's `_on_response` hook. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. | +| ADK `model()`, CrewAI | gateway identity observed | The framework owns the transport (LiteLLM for ADK's `model()`, CrewAI's native OpenAI provider for CrewAI), so no response reaches the SDK's `_on_response` hook. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. | | LlamaIndex, Microsoft Agent Framework | gateway identity observed | These adapters receive `default_headers`, not the SDK's `httpx` client, so no response reaches `_on_response`. When every resolved adapter is non-observing, `donkey.last_call` reports `UNAVAILABLE` and names them in `surface`. | diff --git a/python/examples/adk/README.md b/python/examples/adk/README.md index b508668c..83b9aa0c 100644 --- a/python/examples/adk/README.md +++ b/python/examples/adk/README.md @@ -55,6 +55,23 @@ The factory (`donkey_kit.integrations.adk.model`) fills in `api_base`, `api_key`, `extra_headers`, and the `openai/` prefix from one governed config source. +## Native Gemini (`Format=Gemini` proxy) + +On a proxy provisioned **Format = Gemini**, `gemini()` returns ADK's native +`google.adk.models.Gemini` with the SDK's shared http client injected, so +per-run correlation, spans, usage and `donkey.last_call` all work. Pass +`base_url` when the Gemini proxy is not `DONKEY_LLM_PROXY_URL`: + +```python +from donkey_kit.integrations.adk import gemini + +m = gemini("gemini-2.5-flash", base_url="https:////") +``` + +The manual equivalent is +`Gemini(model="gemini-2.5-flash", **donkey.adk.gemini_connection_kwargs())`; +the docs page lists every kwarg it fills in. + ## Links - Google Agent Development Kit (ADK) docs: see the framework's official diff --git a/python/examples/adk/main.py b/python/examples/adk/main.py index f4fb0a76..c1f3f47f 100644 --- a/python/examples/adk/main.py +++ b/python/examples/adk/main.py @@ -8,6 +8,9 @@ from donkey_kit.integrations.adk import model m = model("gpt-4o") # sent to LiteLLM as "openai/gpt-4o" +On a ``Format=Gemini`` proxy, ``gemini("gemini-2.5-flash")`` from the same +module returns ADK's native ``google.adk.models.Gemini`` instead (see README). + Honest status (verification discipline / docs/verified-apis.md §8): the proxy *contract* (base URL, client_id/secret auth, attribution headers) is live-verified, and ``LiteLlm``/its kwargs are verified per the FACTS table. What is NOT attempted here is a live diff --git a/python/scripts/verify_frameworks.py b/python/scripts/verify_frameworks.py index 80d6005b..511398fd 100644 --- a/python/scripts/verify_frameworks.py +++ b/python/scripts/verify_frameworks.py @@ -62,6 +62,11 @@ class path is wrong -> ImportError/AttributeError. If a kwarg name is wrong "langchain_openai.ChatOpenAI", "langchain-openai"), ("adk", "donkey_kit.integrations.adk", "model", "google.adk.models.lite_llm.LiteLlm", "google-adk"), + # ADK's native Gemini on a Format=Gemini proxy (#691). A dotted key is a second + # factory of the same framework: `--only adk` selects both, and its extra is + # the part before the dot. + ("adk.gemini", "donkey_kit.integrations.adk", "gemini", + "google.adk.models.Gemini", "google-adk"), ("strands", "donkey_kit.integrations.strands", "model", "strands.models.openai.OpenAIModel", "strands-agents"), ("agent_framework", "donkey_kit.integrations.agent_framework", "chat_client", @@ -84,6 +89,10 @@ class path is wrong -> ImportError/AttributeError. If a kwarg name is wrong "DONKEY_LLM_PROXY_CLIENT_SECRET", ) MODEL = os.environ.get("DEMO_MODEL", "gpt-4o") +# adk.gemini needs a Format=Gemini proxy; the default DDK proxies are Format=OpenAI. +# Its live round-trip runs only when this points at one (same consumer creds). +GEMINI_PROXY_ENV = "DONKEY_GEMINI_PROXY_URL" +GEMINI_MODEL = os.environ.get("DEMO_GEMINI_MODEL", "gemini-2.5-flash") @dataclass @@ -184,7 +193,13 @@ def check_signature( try: # Anthropic's native surface is a client; the model id is a per-call # argument, so its factory takes no positional model (BG §1.8 divergence). - obj: object = fn() if res.framework == "anthropic" else fn(MODEL) + obj: object + if res.framework == "anthropic": + obj = fn() + elif res.framework == "adk.gemini": + obj = fn(GEMINI_MODEL, base_url=os.environ.get(GEMINI_PROXY_ENV) or None) + else: + obj = fn(MODEL) except NotImplementedError as exc: # blocked on verification res.installed = True res.blocked = True @@ -242,6 +257,29 @@ async def check_live(res: Result, obj: object) -> None: except Exception as exc: # noqa: BLE001 res.live = f"fail: {type(exc).__name__}: {exc}" return + if res.framework == "adk.gemini": + # ADK's model-level call (BaseLlm.generate_content_async), live-verified + # against ddk-gemini-inbound on 2026-09-29 (#691). + if not os.environ.get(GEMINI_PROXY_ENV): + res.live = f"skipped: set {GEMINI_PROXY_ENV} to a Format=Gemini proxy" + return + try: + from google.adk.models.llm_request import LlmRequest + from google.genai import types + + request = LlmRequest( + model=GEMINI_MODEL, + contents=[types.Content(role="user", parts=[types.Part(text="Say hi.")])], + ) + text = "" + async for resp in obj.generate_content_async(request): # type: ignore[attr-defined] + if resp.content and resp.content.parts: + text += "".join(p.text or "" for p in resp.content.parts) + res.live = "ok" + res.detail = f"completion: {text!r}" + except Exception as exc: # noqa: BLE001 + res.live = f"fail: {type(exc).__name__}: {exc}" + return res.live = ( "skipped: framework runtime call API not verified (docs/verified-apis.md §8/§9); " "shared proxy path verified via raw client (docs/verified-apis.md §2)" @@ -252,7 +290,7 @@ async def run(only: list[str] | None, live: bool) -> list[Result]: have_real = _ensure_proxy_env_for_offline() results: list[Result] = [] for key, import_path, factory, expected, distribution in FRAMEWORKS: - if only and key not in only: + if only and key not in only and key.partition(".")[0] not in only: continue res = Result(framework=key, expected_class=expected) obj = check_signature(res, import_path, factory, distribution) diff --git a/python/src/donkey_kit/core/lastcall.py b/python/src/donkey_kit/core/lastcall.py index 18647d7c..b3b82f3b 100644 --- a/python/src/donkey_kit/core/lastcall.py +++ b/python/src/donkey_kit/core/lastcall.py @@ -32,7 +32,7 @@ correlation id. **Three honest states (hazard #3).** A bare ``request_id is None`` is a lie of -omission on the adapters where the SDK does not own the transport (ADK, CrewAI, +omission on the adapters where the SDK does not own the transport (ADK ``model()``, CrewAI, LlamaIndex, MS Agent Framework — see :attr:`donkey_kit.integrations._base.Adapter.observes_last_call`). A developer there would read ``None`` as "the gateway sent no id" when the truth @@ -355,10 +355,13 @@ def is_fallback(response: httpx.Response) -> bool: # * ``reasoning_tokens`` — output tokens spent on reasoning the developer never # sees. A reasoning model can spend most of its output here, so reading only # ``total_tokens`` draws the wrong conclusion about both cost and latency. -# Both wire shapes are read so the raw client, the deep LangGraph adapter, and any -# OpenAI-compatible call populate identically: the Responses API -# (``input_tokens`` + ``input_tokens_details``) and Chat Completions -# (``prompt_tokens`` + ``prompt_tokens_details``). +# All three wire shapes are read so the raw client, the deep LangGraph adapter, +# any OpenAI-compatible call and a native Gemini call populate identically: the +# Responses API (``input_tokens`` + ``input_tokens_details``), Chat Completions +# (``prompt_tokens`` + ``prompt_tokens_details``), and Gemini's ``usageMetadata`` +# (``promptTokenCount`` / ``candidatesTokenCount`` / ``totalTokenCount`` / +# ``cachedContentTokenCount`` / ``thoughtsTokenCount``; LIVE, #540/#691). Gemini's +# ``totalTokenCount`` already includes the thoughts, so it is taken as reported. _USAGE_FIELDS = ( "input_tokens", "output_tokens", @@ -395,8 +398,9 @@ def parse_usage(usage: object) -> dict[str, int | None]: dict keyed by :data:`_USAGE_FIELDS`. Handles the Responses API (``input_tokens`` / ``input_tokens_details`` / - ``output_tokens_details``) and Chat Completions (``prompt_tokens`` / - ``prompt_tokens_details`` / ``completion_tokens_details``) shapes. A non-dict + ``output_tokens_details``), Chat Completions (``prompt_tokens`` / + ``prompt_tokens_details`` / ``completion_tokens_details``) and Gemini + ``usageMetadata`` (flat camelCase counts) shapes. A non-dict ``usage`` (absent, ``None``, wrong type) yields all-``None`` — an absent count is ``None``, never ``0`` (the same honesty rule ``Budget`` applies to an unobserved window). Never raises (verification discipline).""" @@ -411,25 +415,41 @@ def parse_usage(usage: object) -> dict[str, int | None]: } input_details = _usage_details(usage, "input_tokens_details", "prompt_tokens_details") output_details = _usage_details(usage, "output_tokens_details", "completion_tokens_details") + cached = _first_int(input_details, "cached_tokens") + reasoning = _first_int(output_details, "reasoning_tokens") return { - "input_tokens": _first_int(usage, "input_tokens", "prompt_tokens"), - "output_tokens": _first_int(usage, "output_tokens", "completion_tokens"), - "total_tokens": _first_int(usage, "total_tokens"), - "cached_tokens": _first_int(input_details, "cached_tokens"), + "input_tokens": _first_int(usage, "input_tokens", "prompt_tokens", "promptTokenCount"), + "output_tokens": _first_int( + usage, "output_tokens", "completion_tokens", "candidatesTokenCount" + ), + "total_tokens": _first_int(usage, "total_tokens", "totalTokenCount"), + "cached_tokens": cached + if cached is not None + else _first_int(usage, "cachedContentTokenCount"), "cache_write_tokens": _first_int(input_details, "cache_write_tokens"), - "reasoning_tokens": _first_int(output_details, "reasoning_tokens"), + "reasoning_tokens": reasoning + if reasoning is not None + else _first_int(usage, "thoughtsTokenCount"), } +def _body_usage(body: dict[str, object]) -> object: + """A response body's usage object: OpenAI's ``usage``, else Gemini's + ``usageMetadata``.""" + usage = body.get("usage") + return usage if usage is not None else body.get("usageMetadata") + + def usage_mapping(obj: object) -> dict[str, object] | None: """The ``usage`` mapping from a parsed SSE ``data:`` object, or ``None``. - Handles Chat Completions (top-level ``usage``) and the Responses API (``usage`` + Handles Chat Completions (top-level ``usage``), the Responses API (``usage`` nested under ``response``, as the terminal ``response.completed`` event - carries it). Feeds the streaming scanner, which then :func:`parse_usage` it.""" + carries it) and Gemini (top-level ``usageMetadata`` on each chunk). Feeds the + streaming scanner, which then :func:`parse_usage` it.""" if not isinstance(obj, dict): return None - usage = obj.get("usage") + usage = _body_usage(obj) if not isinstance(usage, dict): nested = obj.get("response") usage = nested.get("usage") if isinstance(nested, dict) else None @@ -450,7 +470,7 @@ def usage_from_response(response: httpx.Response) -> dict[str, int | None]: return parse_usage(None) if not isinstance(body, dict): return parse_usage(None) - return parse_usage(body.get("usage")) + return parse_usage(_body_usage(body)) @dataclass(frozen=True) diff --git a/python/src/donkey_kit/core/transport.py b/python/src/donkey_kit/core/transport.py index 796ef447..8a3ec8b4 100644 --- a/python/src/donkey_kit/core/transport.py +++ b/python/src/donkey_kit/core/transport.py @@ -35,6 +35,7 @@ import asyncio import json import random +import re import time from collections.abc import AsyncIterator, Iterator @@ -337,11 +338,9 @@ def _gateway_unavailable( _STREAM_CONTENT_TYPE = "text/event-stream" -def _request_model(request: httpx.Request) -> str | None: - """The requested model from the request's JSON body (``gen_ai.request.model``), - or ``None`` when the body is absent, unreadable, not JSON, or carries no - ``model``. A ``None`` marks "not a GenAI call": no span is opened, so GETs, - token fetches and bodyless POSTs stay byte-identical.""" +def _body_model(request: httpx.Request) -> str | None: + """The ``model`` from the request's JSON body, or ``None`` when the body is + absent, unreadable, not JSON, or carries no ``model``.""" try: raw = request.content except Exception: # noqa: BLE001 — streaming/unread body is not a model call @@ -358,6 +357,28 @@ def _request_model(request: httpx.Request) -> str | None: return None +# A Format=Gemini proxy carries the model in the URL path, never the body +# (docs/verified-apis.md §2, #540/#691): ``/models/:generateContent`` +# (LIVE-verified) and its SSE twin ``:streamGenerateContent``. The ingress ignores +# a body ``model``, so this is read for the SDK's own bookkeeping only — the +# request on the wire is never changed. +_GEMINI_MODEL_PATH = re.compile(r"/models/([^/:]+):(?:generateContent|streamGenerateContent)$") + + +def _request_model(request: httpx.Request) -> str | None: + """The requested model (``gen_ai.request.model``): the JSON body's ``model``, + else the model segment of a native Gemini ``POST`` path. ``None`` marks "not a + GenAI call": no span is opened, so GETs, token fetches and bodyless POSTs + stay byte-identical.""" + model = _body_model(request) + if model is not None: + return model + if request.method != "POST": + return None + match = _GEMINI_MODEL_PATH.search(request.url.path) + return match.group(1) if match else None + + def _span_decision(response: httpx.Response) -> tuple[str | None, str | None]: """``(donkey.policy.decision, donkey.policy.type)`` for the final response: ``allow`` on 2xx; ``refuse`` + a policy-type slug when :func:`classify` maps @@ -526,7 +547,8 @@ def close(self) -> None: self._buf = b"" def _scan(self, line: bytes) -> None: - if b'"usage"' not in line: + # ``"usage`` matches both OpenAI's ``"usage"`` and Gemini's ``"usageMetadata"``. + if b'"usage' not in line: return stripped = line.strip() if not stripped.startswith(b"data:"): diff --git a/python/src/donkey_kit/donkey.py b/python/src/donkey_kit/donkey.py index 77ccbc0f..f790f157 100644 --- a/python/src/donkey_kit/donkey.py +++ b/python/src/donkey_kit/donkey.py @@ -233,7 +233,8 @@ def last_call(self) -> LastCall: :class:`~donkey_kit.core.errors.DonkeyError` hands you on a refusal. Usage counts are read from the response body, so they are ``None`` (never - ``0``) when the gateway sent no ``usage`` object; on a streamed response + ``0``) when the gateway sent no ``usage`` (or Gemini ``usageMetadata``) + object; on a streamed response they land once the terminal SSE event has been consumed, not at first read. Contextvar-scoped, not instance-scoped (hazard #2): under the parallel @@ -249,7 +250,7 @@ def last_call(self) -> LastCall: it said nothing"). * **UNOBSERVED** — no governed model call has returned in this context yet. * **UNAVAILABLE** — every adapter used on this Donkey routes outside our - transport (ADK via LiteLLM, CrewAI via its native OpenAI provider, or + transport (ADK ``model()`` via LiteLLM, CrewAI via its native OpenAI provider, or ``default_headers``-only LlamaIndex / MS Agent Framework), so a response can never reach the record. :attr:`LastCall.surface` names which. This is derived from the adapters actually resolved, and the diff --git a/python/src/donkey_kit/integrations/_base.py b/python/src/donkey_kit/integrations/_base.py index 9ce6fda7..b907512e 100644 --- a/python/src/donkey_kit/integrations/_base.py +++ b/python/src/donkey_kit/integrations/_base.py @@ -30,7 +30,7 @@ class Adapter: #: (#362). True when the adapter hands the framework our shared #: :class:`DonkeyAsyncClient` (its ``_on_response`` observes the response); #: False when the SDK does not own the transport — the framework builds its own - #: client (ADK via LiteLLM, CrewAI via its native OpenAI provider) or the + #: client (ADK ``model()`` via LiteLLM, CrewAI via its native OpenAI provider) or the #: adapter is given only ``default_headers`` (LlamaIndex, MS Agent Framework). #: A ``False`` here is why ``donkey.last_call`` reports "not available on this #: surface" rather than a bare ``None`` (hazard #3), diff --git a/python/src/donkey_kit/integrations/adk.py b/python/src/donkey_kit/integrations/adk.py index d9ee20a6..88f3f386 100644 --- a/python/src/donkey_kit/integrations/adk.py +++ b/python/src/donkey_kit/integrations/adk.py @@ -2,17 +2,28 @@ Supported at connection_kwargs() — not conformance-tested (BG §1.8). -ADK is Gemini-first and reaches other providers through the ``LiteLlm`` wrapper, -which takes LiteLLM-format model strings. +Two factories, one per proxy ingress Format (docs/verified-apis.md §2): -Header injection: via LiteLLM's ``extra_headers``. We CANNOT inject our httpx -client — LiteLLM owns the transport. Consequence: transport retries and -correlation-ID-per-run degrade to per-client. This is a documented, asserted -conformance exemption (the conformance kit's ``correlation_id_propagated``). A LiteLLM custom -logger callback may later recover trace correlation. +* ``model()`` — ADK's ``LiteLlm`` wrapper for a ``Format=OpenAI`` proxy (the + SDK's default DDK proxies). LiteLLM takes ``openai/`` model strings and + speaks ``/chat/completions``. Header injection is via LiteLLM's + ``extra_headers``; we CANNOT inject our httpx client — LiteLLM owns the + transport. Consequence: transport retries and correlation-ID-per-run degrade + to per-client, and ``donkey.last_call`` is not populated. These are + documented, asserted conformance exemptions (the conformance kit's + ``correlation_id_propagated`` / ``gateway_identity_observed``). ADK requires + ``litellm>=1.84`` (floor, not ceiling). +* ``gemini()`` — ADK's native ``google.adk.models.Gemini`` for a + ``Format=Gemini`` proxy (#691). The LIVE-verified native route is + ``POST /models/:generateContent`` (#540); the model travels in + the URL only, and the ingress ignores a body ``model`` (probed 2026-09-29). + ``google-genai`` accepts ``HttpOptions.httpx_async_client``, so we hand it the + shared :class:`~donkey_kit.core.transport.DonkeyAsyncClient`: full injection — + per-run correlation, SDK retries, rotating JWTs and ``donkey.last_call`` all + work, and none of ``model()``'s exemptions apply. The default DDK proxies are + ``Format=OpenAI``, so point ``base_url`` at a ``Format=Gemini`` proxy. -Class names / kwargs UNVERIFIED — docs/verified-apis.md §8. ADK requires -``litellm>=1.84`` (floor, not ceiling). +Class names / kwargs UNVERIFIED — docs/verified-apis.md §8. """ from __future__ import annotations @@ -22,13 +33,16 @@ from ._base import Adapter, default_adapter if TYPE_CHECKING: + from google.adk.models import Gemini from google.adk.models.lite_llm import LiteLlm class ADKAdapter(Adapter): extra = "adk" - # LiteLLM owns the transport, so no response reaches donkey.last_call (#362, - # the same reason as the conformance kit's correlation_id_propagated exemption). + # LiteLLM owns the transport, so no response from ``model()`` reaches + # donkey.last_call (#362, the same reason as the conformance kit's + # correlation_id_propagated exemption). ``gemini()`` routes through our + # transport and flips this on the instance that built it (#691). observes_last_call = False def connection_kwargs(self) -> dict[str, Any]: @@ -43,15 +57,60 @@ def connection_kwargs(self) -> dict[str, Any]: "extra_headers": conn["default_headers"], } + def gemini_connection_kwargs(self, *, base_url: str | None = None) -> dict[str, Any]: + """Governed kwargs for a ``Gemini(model=, **kwargs)`` you build + yourself, bound to a ``Format=Gemini`` proxy (``base_url``, default the + configured proxy URL). + + ADK replaces its own ``http_options`` with ``client_kwargs``, so every + governed value rides there: the shared http client (header injection, + retries, last_call), the consumer-auth headers, the ``api_key`` slot + google-genai requires for the Gemini API backend, and the SDK timeout — + google-genai otherwise sends ``timeout=None``, which disables the + client's. ``api_version=""`` because the proxy route has no + ``/v1beta`` segment.""" + conn = self._openai_connection() + url = base_url or conn["base_url"] + return { + "base_url": url, + "client_kwargs": { + "api_key": conn["api_key"], + "http_options": { + "base_url": url, + "api_version": "", + "headers": conn["default_headers"], + "timeout": int(self._cfg.timeout_s * 1000), + "httpx_async_client": self._http_client(), + }, + }, + } + def model(self, model: str, **kw: Any) -> LiteLlm: from google.adk.models.lite_llm import LiteLlm # VERIFY name/path: docs/verified-apis.md §8 # LiteLLM's OpenAI-compatible route needs the ``openai/`` prefix. return LiteLlm(model=f"openai/{model}", **{**self.connection_kwargs(), **kw}) + def gemini(self, model: str, *, base_url: str | None = None, **kw: Any) -> Gemini: + """Return ADK's native ``google.adk.models.Gemini`` bound to a + ``Format=Gemini`` proxy, with the shared http client injected (#691). + Pass the bare model id (``"gemini-2.5-flash"``) — no provider prefix.""" + from google.adk.models import Gemini # VERIFY name/path: docs/verified-apis.md §8 + + native = Gemini(model=model, **{**self.gemini_connection_kwargs(base_url=base_url), **kw}) + self.observes_last_call = True + return native + def model(model: str, **kw: Any) -> LiteLlm: """Module-level convenience: a native ``LiteLlm`` at the proxy using a cached default env-configured Donkey. Equivalent to ``Donkey.from_env().adk.model(model, **kw)``.""" return default_adapter(ADKAdapter).model(model, **kw) + + +def gemini(model: str, *, base_url: str | None = None, **kw: Any) -> Gemini: + """Module-level convenience: a native ``Gemini`` at a ``Format=Gemini`` proxy + using a cached default env-configured Donkey. Equivalent to + ``Donkey.from_env().adk.gemini(model, base_url=..., **kw)``.""" + return default_adapter(ADKAdapter).gemini(model, base_url=base_url, **kw) diff --git a/python/tests/conformance/suite.py b/python/tests/conformance/suite.py index 8697eaee..f9e3caa5 100644 --- a/python/tests/conformance/suite.py +++ b/python/tests/conformance/suite.py @@ -54,7 +54,7 @@ # Documented, ASSERTED exemptions — published in the README (the conformance kit). A framework # that cannot satisfy a scenario records WHY here rather than skipping silently. _LITELLM_TRANSPORT_EXEMPTION = ( - "LiteLLM owns the transport; we cannot inject our httpx client, so the " + "adk.model() only: LiteLLM owns the transport; we cannot inject our httpx client, so the " "correlation ID is per-client, not per-run (BG §1.8). A LiteLLM custom " "logger callback may later recover trace correlation." ) @@ -76,7 +76,7 @@ # surface) rather than a bare None — the honest-state contract (hazard #3) — # and mirrors Adapter.observes_last_call = False on each of these adapters. _LITELLM_LAST_CALL_EXEMPTION = ( - "LiteLLM owns the transport; no response reaches our _on_response, so " + "adk.model() only: LiteLLM owns the transport; no response reaches our _on_response, so " "donkey.last_call cannot observe the gateway identity of the call and " "reports UNAVAILABLE (#362, same cause as correlation_id_propagated BG §1.8)." ) @@ -99,8 +99,8 @@ # starts 401-ing after it expires — so jwt auth mode is unsupported on those # adapters, asserted here rather than silently skipped. _LITELLM_JWT_EXEMPTION = ( - "LiteLLM owns the transport; we cannot inject our httpx client, so a rotating " - "model-wallet JWT cannot be refreshed per-send and would expire. Use client-id " + "adk.model() only: LiteLLM owns the transport; we cannot inject our httpx client, so a " + "rotating model-wallet JWT cannot be refreshed per-send and would expire. Use client-id " "auth with this adapter, or route the raw/LangGraph client for jwt mode (#509)." ) _CREWAI_JWT_EXEMPTION = ( @@ -116,8 +116,9 @@ ) KNOWN_LIMITATIONS: dict[str, dict[str, str]] = { - # ADK reaches models through LiteLLM and CrewAI through its native OpenAI - # provider; either way the framework owns the transport (BG §1.8). + # ADK's model() reaches models through LiteLLM and CrewAI through its native + # OpenAI provider; either way the framework owns the transport (BG §1.8). + # adk.gemini() is handed our httpx client and records none of these (#691). "adk": { "correlation_id_propagated": _LITELLM_TRANSPORT_EXEMPTION, "gateway_identity_observed": _LITELLM_LAST_CALL_EXEMPTION, diff --git a/python/tests/conformance/test_transport_exemptions.py b/python/tests/conformance/test_transport_exemptions.py index 5b6fdc05..10df1303 100644 --- a/python/tests/conformance/test_transport_exemptions.py +++ b/python/tests/conformance/test_transport_exemptions.py @@ -104,6 +104,26 @@ def test_exemption_matches_observes_last_call_flag() -> None: assert non_observing == {"adk", "crewai", "llamaindex", "agent_framework"} +async def test_adk_gemini_is_not_exempt() -> None: + # The adk exemptions are scoped to model() (LiteLLM). adk.gemini() is handed + # the shared DonkeyAsyncClient — the fact that makes correlation, last_call and + # per-send JWT work — and so flips observes_last_call on its instance (#691). + cfg = DonkeyConfig( + llm_proxy_url="https://proxy", + llm_proxy_client_id="cid", + llm_proxy_client_secret="secret", + ) + client = build_http_client(cfg, None) + try: + adapter = _adapter_class("adk")(cfg, client) + kw = adapter.gemini_connection_kwargs() # type: ignore[attr-defined] + assert kw["client_kwargs"]["http_options"]["httpx_async_client"] is client + finally: + await client.aclose() + for reason in KNOWN_LIMITATIONS["adk"].values(): + assert reason.startswith("adk.model() only:") + + async def test_header_only_correlation_exemptions_match_connection_kwargs() -> None: cfg = DonkeyConfig( llm_proxy_url="https://proxy", diff --git a/python/tests/fixtures/anypoint/gemini_inbound/README.md b/python/tests/fixtures/anypoint/gemini_inbound/README.md index bf9986db..77129194 100644 --- a/python/tests/fixtures/anypoint/gemini_inbound/README.md +++ b/python/tests/fixtures/anypoint/gemini_inbound/README.md @@ -56,12 +56,34 @@ directly. Consumer `client_id`/`client_secret` are **redacted** from `request.success.http` and were never persisted; the probe applications were removed after capture. -## There is still no Gemini adapter — by design +## Streaming and refusal captures (2026-09-29, #691) -Verifying the route does **not** mean the SDK ships a `donkey.gemini` adapter. -Per #540 / #244 a native adapter (a `google-genai` client bound at -`connection_kwargs()` level, `BG §1.8`) is **demand-driven** — "do not guess it". +Captured for the ADK native Gemini adapter (`donkey.adk.gemini(...)`), same +instance `21193369`, with a consumer pair contracted on the proxy. The pair is +**redacted** from both `request.*.http` files. + +- **Streaming is routed.** `POST //models/:streamGenerateContent?alt=sse` + → **HTTP 200**, `content-type: text/event-stream`. Each `data:` event is a + native Gemini chunk; `usageMetadata` is **cumulative**, so the last event + carries the call's total. See `request.stream.http` + + `responses.stream.{body.txt,headers.txt}` (sent with `thinkingBudget: 0`). +- **Unknown model is Google's 404, passed through.** `models/gemini-no-such-model:generateContent` + → **HTTP 404** with Google's `NOT_FOUND` envelope; `classify()` maps it to + `UpstreamRequestError`. See `request.unknown-model.http` + + `reject.unknown-model.*`. + +`usageMetadata.totalTokenCount` includes `thoughtsTokenCount`, and no +`x-llm-proxy-llm-model` / `-llm-provider` header is emitted, so +`donkey.last_call.served_model` stays `None` on this route. The body `model` +field is ignored by the gateway (the URL path picks the model), so the SDK reads +the requested model from the path. `tests/unit/test_gemini_inbound_contract.py` +pins these shapes. The simulator does not serve these files. + +## The SDK adapter + +Google ADK's native `google.adk.models.Gemini` is bound to this ingress by +`donkey.adk.gemini(...)` (#691, `BG §1.8`), with the shared http client injected +through `HttpOptions.httpx_async_client`. There is no standalone `google-genai` +adapter; the manual equivalent is `donkey.adk.gemini_connection_kwargs()`. Gemini also remains reachable as an *upstream provider* behind an OpenAI-format -ingress via model-based routing (the §2 supported-providers row), which needs no -new code. This capture records that the native ingress **exists and works**; a -future adapter issue is filed only if demand warrants. +ingress via model-based routing (the §2 supported-providers row). diff --git a/python/tests/fixtures/anypoint/gemini_inbound/reject.unknown-model.body.json b/python/tests/fixtures/anypoint/gemini_inbound/reject.unknown-model.body.json new file mode 100644 index 00000000..87c89099 --- /dev/null +++ b/python/tests/fixtures/anypoint/gemini_inbound/reject.unknown-model.body.json @@ -0,0 +1,7 @@ +{ + "error": { + "code": 404, + "message": "models/gemini-no-such-model is not found for API version v1beta, or is not supported for generateContent. Call ModelService.ListModels to see the list of available models and their supported methods.", + "status": "NOT_FOUND" + } +} diff --git a/python/tests/fixtures/anypoint/gemini_inbound/reject.unknown-model.headers.txt b/python/tests/fixtures/anypoint/gemini_inbound/reject.unknown-model.headers.txt new file mode 100644 index 00000000..878a00af --- /dev/null +++ b/python/tests/fixtures/anypoint/gemini_inbound/reject.unknown-model.headers.txt @@ -0,0 +1,18 @@ +HTTP/1.1 404 Not Found +vary: X-Origin,Referer,Origin,Accept-Encoding +content-type: application/json; charset=UTF-8 +date: Tue, 29 Sep 2026 19:32:04 GMT +server: Anypoint Flex Gateway +x-xss-protection: 0 +x-frame-options: SAMEORIGIN +x-content-type-options: nosniff +server-timing: gfet4t7; dur=93 +alt-svc: h3=":443"; ma=2592000,h3-29=":443"; ma=2592000 +accept-ranges: none +x-llm-proxy-model-based-routing-success: Request passed through without model-based routing. +x-envoy-decorator-operation: api-instance-21193369.14d3b31e-4e3b-4d90-b77a-63c9d6b7ea6a.svc +transfer-encoding: chunked +x-correlation-id: 077c5100-90b3-489e-b7b8-a0edf5264f5b +Strict-Transport-Security: max-age=31536000; includeSubdomains +Connection: Keep-Alive + diff --git a/python/tests/fixtures/anypoint/gemini_inbound/request.stream.http b/python/tests/fixtures/anypoint/gemini_inbound/request.stream.http new file mode 100644 index 00000000..bf33b917 --- /dev/null +++ b/python/tests/fixtures/anypoint/gemini_inbound/request.stream.http @@ -0,0 +1,7 @@ +POST /ddk-gemini-inbound/models/gemini-2.5-flash:streamGenerateContent?alt=sse HTTP/1.1 +Host: shared-omni-gateway-qrpud8.5sc6y6-1.usa-e2.cloudhub.io +Content-Type: application/json +client_id: <32-char consumer client id — redacted> +client_secret: <32-char consumer client secret — redacted> + +{"contents":[{"role":"user","parts":[{"text":"Write three short sentences about donkeys."}]}],"generationConfig":{"thinkingConfig":{"thinkingBudget":0}}} diff --git a/python/tests/fixtures/anypoint/gemini_inbound/request.unknown-model.http b/python/tests/fixtures/anypoint/gemini_inbound/request.unknown-model.http new file mode 100644 index 00000000..7e1916d5 --- /dev/null +++ b/python/tests/fixtures/anypoint/gemini_inbound/request.unknown-model.http @@ -0,0 +1,7 @@ +POST /ddk-gemini-inbound/models/gemini-no-such-model:generateContent HTTP/1.1 +Host: shared-omni-gateway-qrpud8.5sc6y6-1.usa-e2.cloudhub.io +Content-Type: application/json +client_id: <32-char consumer client id — redacted> +client_secret: <32-char consumer client secret — redacted> + +{"contents":[{"role":"user","parts":[{"text":"hi"}]}]} diff --git a/python/tests/fixtures/anypoint/gemini_inbound/responses.stream.body.txt b/python/tests/fixtures/anypoint/gemini_inbound/responses.stream.body.txt new file mode 100644 index 00000000..38596ce2 --- /dev/null +++ b/python/tests/fixtures/anypoint/gemini_inbound/responses.stream.body.txt @@ -0,0 +1,6 @@ +data: {"candidates": [{"content": {"parts": [{"text": "Donkeys are known"}],"role": "model"},"index": 0}],"usageMetadata": {"promptTokenCount": 9,"candidatesTokenCount": 4,"totalTokenCount": 13,"promptTokensDetails": [{"modality": "TEXT","tokenCount": 9}],"serviceTier": "standard"},"modelVersion": "gemini-2.5-flash","responseId": "PRK8aqqHN6iU39IPvtWW6A4"} + +data: {"candidates": [{"content": {"parts": [{"text": " for their long ears. They are hardy animals often used for work. Many people"}],"role": "model"},"index": 0}],"usageMetadata": {"promptTokenCount": 9,"candidatesTokenCount": 20,"totalTokenCount": 29,"promptTokensDetails": [{"modality": "TEXT","tokenCount": 9}],"serviceTier": "standard"},"modelVersion": "gemini-2.5-flash","responseId": "PRK8aqqHN6iU39IPvtWW6A4"} + +data: {"candidates": [{"content": {"parts": [{"text": " find their bray to be quite distinctive."}],"role": "model"},"finishReason": "STOP","index": 0}],"usageMetadata": {"promptTokenCount": 9,"candidatesTokenCount": 29,"totalTokenCount": 38,"promptTokensDetails": [{"modality": "TEXT","tokenCount": 9}],"serviceTier": "standard"},"modelVersion": "gemini-2.5-flash","responseId": "PRK8aqqHN6iU39IPvtWW6A4"} + diff --git a/python/tests/fixtures/anypoint/gemini_inbound/responses.stream.headers.txt b/python/tests/fixtures/anypoint/gemini_inbound/responses.stream.headers.txt new file mode 100644 index 00000000..34fcba8e --- /dev/null +++ b/python/tests/fixtures/anypoint/gemini_inbound/responses.stream.headers.txt @@ -0,0 +1,18 @@ +HTTP/1.1 200 OK +content-type: text/event-stream +content-disposition: attachment +vary: Origin,X-Origin,Referer +date: Tue, 29 Sep 2026 19:32:14 GMT +server: Anypoint Flex Gateway +x-xss-protection: 0 +x-frame-options: SAMEORIGIN +x-content-type-options: nosniff +server-timing: gfet4t7; dur=394 +alt-svc: h3=":443"; ma=2592000,h3-29=":443"; ma=2592000 +x-llm-proxy-model-based-routing-success: Request passed through without model-based routing. +x-envoy-decorator-operation: api-instance-21193369.14d3b31e-4e3b-4d90-b77a-63c9d6b7ea6a.svc +transfer-encoding: chunked +x-correlation-id: 2b7a3b6e-7354-4ac1-b71c-ddc265213298 +Strict-Transport-Security: max-age=31536000; includeSubdomains +Connection: Keep-Alive + diff --git a/python/tests/unit/test_adapter_ergonomics.py b/python/tests/unit/test_adapter_ergonomics.py index d4881a83..8a07f29d 100644 --- a/python/tests/unit/test_adapter_ergonomics.py +++ b/python/tests/unit/test_adapter_ergonomics.py @@ -543,3 +543,148 @@ def test_openai_agents_factory_caller_openai_client_overrides_default( model("gpt-4o", openai_client=caller_client) assert captured["openai_client"] is caller_client + + +# --- ADK native Gemini on a Format=Gemini proxy (#691) ----------------------- + + +def test_adk_gemini_connection_kwargs_inject_the_shared_client() -> None: + # ADK replaces its own http_options with client_kwargs, so every governed + # value rides there — including OUR http client (full injection). + from donkey_kit.integrations.adk import ADKAdapter + + cfg = _cfg() + http = build_http_client(cfg, None) + kw = ADKAdapter(cfg, http).gemini_connection_kwargs() + opts = kw["client_kwargs"]["http_options"] + assert kw["base_url"] == opts["base_url"] == "https://proxy" + assert opts["api_version"] == "" # the proxy route has no /v1beta segment + assert opts["httpx_async_client"] is http + assert opts["headers"]["client_id"] == "cid" + assert opts["timeout"] == int(cfg.timeout_s * 1000) # genai sends None otherwise + assert kw["client_kwargs"]["api_key"] # google-genai requires the slot + + +def test_adk_gemini_base_url_override_reaches_both_slots() -> None: + from donkey_kit.integrations.adk import ADKAdapter + + kw = ADKAdapter(_cfg(), _http()).gemini_connection_kwargs(base_url="https://gw/gem/") + assert kw["base_url"] == kw["client_kwargs"]["http_options"]["base_url"] == "https://gw/gem/" + + +def test_adk_gemini_factory_and_connection_kwargs_do_not_drift( + monkeypatch: pytest.MonkeyPatch, +) -> None: + _set_proxy_env(monkeypatch) + captured = _install_native_stub(monkeypatch, "google.adk.models", "Gemini") + from donkey_kit.integrations.adk import ADKAdapter, gemini + + gemini("gemini-2.5-flash", base_url="https://gw/gem/") + + adapter = default_adapter(ADKAdapter) + expected = adapter.gemini_connection_kwargs(base_url="https://gw/gem/") + assert captured["model"] == "gemini-2.5-flash" # bare id, no provider prefix + assert {k: captured[k] for k in expected} == expected + + +def test_adk_gemini_caller_kwargs_override_connection_defaults( + monkeypatch: pytest.MonkeyPatch, +) -> None: + _set_proxy_env(monkeypatch) + captured = _install_native_stub(monkeypatch, "google.adk.models", "Gemini") + from donkey_kit.integrations.adk import gemini + + caller = {"api_key": "k"} + gemini("gemini-2.5-flash", client_kwargs=caller) + assert captured["client_kwargs"] is caller + + +def test_adk_gemini_flips_observes_last_call_on_the_instance_only( + monkeypatch: pytest.MonkeyPatch, +) -> None: + # model() (LiteLLM) cannot observe; gemini() routes through our transport. The + # class default stays False so the exemption table keeps matching model(). + _install_native_stub(monkeypatch, "google.adk.models", "Gemini") + from donkey_kit.integrations.adk import ADKAdapter + + adapter = ADKAdapter(_cfg(), _http()) + assert adapter.observes_last_call is False + adapter.gemini("gemini-2.5-flash") + assert adapter.observes_last_call is True + assert ADKAdapter.observes_last_call is False + + +async def test_adk_gemini_real_round_trip_is_governed_by_our_transport() -> None: + """With google-adk installed: the native Gemini sends through the shared + DonkeyAsyncClient — consumer auth, the per-run correlation id, the native + route — and the reply populates donkey.last_call; a refusal surfaces as + google-genai's APIError whose ``.response`` classify() types (#691).""" + pytest.importorskip("google.adk") + import httpx + from google.adk.models.llm_request import LlmRequest + from google.genai import errors as genai_errors + from google.genai import types as genai_types + + from donkey_kit.core.errors import UpstreamRequestError, classify + from donkey_kit.core.lastcall import LastCallStatus, current_last_call + from donkey_kit.core.telemetry import run_context + from donkey_kit.integrations.adk import ADKAdapter + + seen: list[httpx.Request] = [] + + def handler(request: httpx.Request) -> httpx.Response: + seen.append(request) + if "no-such-model" in request.url.path: + return httpx.Response(404, json={"error": {"code": 404, "status": "NOT_FOUND"}}) + return httpx.Response( + 200, + json={ + "candidates": [{"content": {"role": "model", "parts": [{"text": "PONG"}]}}], + "usageMetadata": { + "promptTokenCount": 9, + "candidatesTokenCount": 2, + "totalTokenCount": 32, + "thoughtsTokenCount": 21, + }, + }, + headers={"x-envoy-decorator-operation": "api-instance-21193369.env.svc"}, + ) + + cfg = _cfg() + http = DonkeyAsyncClient(cfg, None, transport=httpx.MockTransport(handler)) + adapter = ADKAdapter(cfg, http) + + def _request(model: str) -> LlmRequest: + return LlmRequest( + model=model, + contents=[genai_types.Content(role="user", parts=[genai_types.Part(text="hi")])], + ) + + async with http: + with run_context("run-691"): + m = adapter.gemini("gemini-2.5-flash", base_url="https://gw/ddk-gemini-inbound/") + async for _ in m.generate_content_async(_request("gemini-2.5-flash")): + pass + record = current_last_call() + + bad = adapter.gemini("no-such-model", base_url="https://gw/ddk-gemini-inbound/") + with pytest.raises(genai_errors.APIError) as exc_info: + async for _ in bad.generate_content_async(_request("no-such-model")): + pass + + sent = seen[0] + assert sent.url.path == "/ddk-gemini-inbound/models/gemini-2.5-flash:generateContent" + assert sent.headers["client_id"] == "cid" + assert sent.headers[http._correlation_header] == "run-691" + assert "model" not in json_body(sent) # the wire body stays pure Gemini + assert record is not None and record.status is LastCallStatus.OBSERVED + assert record.requested_model == "gemini-2.5-flash" + assert (record.input_tokens, record.total_tokens, record.reasoning_tokens) == (9, 32, 21) + assert isinstance(classify(exc_info.value.response), UpstreamRequestError) + + +def json_body(request: Any) -> dict[str, Any]: + import json + + body: dict[str, Any] = json.loads(request.content) + return body diff --git a/python/tests/unit/test_gemini_inbound_contract.py b/python/tests/unit/test_gemini_inbound_contract.py new file mode 100644 index 00000000..afc4ca04 --- /dev/null +++ b/python/tests/unit/test_gemini_inbound_contract.py @@ -0,0 +1,153 @@ +"""Pins the native Gemini ingress contract to LIVE captures from a real +``Format=Gemini`` proxy (``ddk-gemini-inbound``, instance 21193369) — see +tests/fixtures/anypoint/gemini_inbound/README.md and docs/verified-apis.md §2 +(#540, #691). + +Two halves: the captured shapes themselves, and the framework-free transport +consuming them. A native Gemini request carries the model in the URL path and +reports usage as ``usageMetadata``, so the transport must read both for +``donkey.last_call``, the GenAI span and usage to work on ``adk.gemini()``. +""" + +from __future__ import annotations + +import json +from pathlib import Path + +import httpx + +from donkey_kit.core.config import DonkeyConfig +from donkey_kit.core.errors import UpstreamRequestError, classify +from donkey_kit.core.lastcall import LastCallStatus, current_last_call +from donkey_kit.core.transport import DonkeyAsyncClient, _request_model +from donkey_kit.simulator.fixtures import parse_headers + +FIXTURES = Path(__file__).resolve().parents[1] / "fixtures" / "anypoint" / "gemini_inbound" +_CFG = DonkeyConfig(llm_proxy_url="https://proxy") +_GENERATE = "https://proxy/ddk-gemini-inbound/models/gemini-2.5-flash:generateContent" +_STREAM = "https://proxy/ddk-gemini-inbound/models/gemini-2.5-flash:streamGenerateContent?alt=sse" +_CONTENTS = {"contents": [{"role": "user", "parts": [{"text": "hi"}]}]} + + +def _headers(name: str) -> dict[str, str]: + return parse_headers((FIXTURES / name).read_text()) + + +def _fixture_response(body: str, headers: str, status: int = 200) -> httpx.Response: + return httpx.Response( + status, content=(FIXTURES / body).read_bytes(), headers=_headers(headers) + ) + + +# --- the captured shapes ------------------------------------------------------ + + +def test_success_body_is_native_gemini_with_usage_metadata() -> None: + body = json.loads((FIXTURES / "responses.success.body.json").read_text()) + usage = body["usageMetadata"] + assert body["candidates"][0]["content"]["role"] == "model" + # totalTokenCount already INCLUDES the thoughts — it is not prompt + candidates. + assert usage["totalTokenCount"] == ( + usage["promptTokenCount"] + usage["candidatesTokenCount"] + usage["thoughtsTokenCount"] + ) + + +def test_passthrough_carries_no_served_model_routing_headers() -> None: + # A native ingress passes through: no provider/model/routing-type headers, so + # last_call's served_* fields stay None on this path (correct, not a gap). + for name in ("responses.success.headers.txt", "responses.stream.headers.txt"): + h = _headers(name) + assert h["x-llm-proxy-model-based-routing-success"].startswith("Request passed through") + assert "x-llm-proxy-llm-model" not in h + assert "x-llm-proxy-routing-type" not in h + assert h["x-envoy-decorator-operation"].startswith("api-instance-21193369.") + + +def test_stream_is_sse_with_cumulative_usage_per_event() -> None: + assert _headers("responses.stream.headers.txt")["content-type"] == "text/event-stream" + events = [ + json.loads(line[len("data:") :]) + for line in (FIXTURES / "responses.stream.body.txt").read_text().splitlines() + if line.startswith("data:") + ] + totals = [e["usageMetadata"]["totalTokenCount"] for e in events] + assert len(events) > 1 + assert totals == sorted(totals) # each event reports the running total + + +def test_unknown_model_is_an_upstream_404_passed_through() -> None: + resp = _fixture_response( + "reject.unknown-model.body.json", "reject.unknown-model.headers.txt", status=404 + ) + assert resp.json()["error"]["status"] == "NOT_FOUND" + assert isinstance(classify(resp), UpstreamRequestError) + + +# --- the transport consuming them --------------------------------------------- + + +def test_request_model_reads_the_gemini_url_path() -> None: + assert _request_model(httpx.Request("POST", _GENERATE, json=_CONTENTS)) == "gemini-2.5-flash" + assert _request_model(httpx.Request("POST", _STREAM, json=_CONTENTS)) == "gemini-2.5-flash" + + +def test_request_model_prefers_the_body_and_ignores_non_model_calls() -> None: + body_model = {"model": "gpt-4o", **_CONTENTS} + assert _request_model(httpx.Request("POST", _GENERATE, json=body_model)) == "gpt-4o" + # A GET on the same path, an unrelated method suffix, or a plain POST is not a + # model call, so it opens no span and records no last_call. + assert _request_model(httpx.Request("GET", _GENERATE)) is None + assert ( + _request_model(httpx.Request("POST", _GENERATE.replace("generateContent", "countTokens"))) + is None + ) + assert _request_model(httpx.Request("POST", "https://proxy/token", json={"a": 1})) is None + + +async def test_buffered_gemini_call_populates_last_call() -> None: + def handler(request: httpx.Request) -> httpx.Response: + return _fixture_response("responses.success.body.json", "responses.success.headers.txt") + + async with DonkeyAsyncClient(_CFG, None, transport=httpx.MockTransport(handler)) as client: + await client.post(_GENERATE, json=_CONTENTS) + + record = current_last_call() + assert record is not None and record.status is LastCallStatus.OBSERVED + assert record.requested_model == "gemini-2.5-flash" + assert record.api_instance_id == "21193369" + assert (record.input_tokens, record.output_tokens, record.total_tokens) == (9, 2, 32) + assert record.reasoning_tokens == 21 + assert record.served_model is None and record.routing_type is None + assert record.substituted is False + + +class _AsyncChunks(httpx.AsyncByteStream): + """A genuinely unread stream — ``content=`` would buffer the SSE body and + bypass the transport's stream wrapper.""" + + def __init__(self, data: bytes, size: int = 64) -> None: + self._chunks = [data[i : i + size] for i in range(0, len(data), size)] + + async def __aiter__(self): # type: ignore[override] + for chunk in self._chunks: + yield chunk + + +async def test_streamed_gemini_call_fills_usage_after_drain() -> None: + def handler(request: httpx.Request) -> httpx.Response: + body = (FIXTURES / "responses.stream.body.txt").read_bytes() + return httpx.Response( + 200, headers=_headers("responses.stream.headers.txt"), stream=_AsyncChunks(body) + ) + + async with DonkeyAsyncClient(_CFG, None, transport=httpx.MockTransport(handler)) as client: + req = client.build_request("POST", _STREAM, json=_CONTENTS) + resp = await client.send(req, stream=True) + async for _line in resp.aiter_lines(): + pass + await resp.aclose() + + record = current_last_call() + assert record is not None and record.requested_model == "gemini-2.5-flash" + # The terminal event's cumulative counts win. + assert (record.input_tokens, record.output_tokens, record.total_tokens) == (9, 29, 38) diff --git a/python/tests/unit/test_lastcall.py b/python/tests/unit/test_lastcall.py index 374f3715..aaed3030 100644 --- a/python/tests/unit/test_lastcall.py +++ b/python/tests/unit/test_lastcall.py @@ -432,6 +432,26 @@ def test_parse_usage_chat_completions_shape() -> None: } +def test_parse_usage_gemini_usage_metadata_shape() -> None: + # Gemini's flat camelCase counts (LIVE, #540/#691). totalTokenCount already + # includes the thoughts, so it is taken as reported, never recomputed. + usage = { + "promptTokenCount": 9, + "candidatesTokenCount": 2, + "totalTokenCount": 32, + "cachedContentTokenCount": 4, + "thoughtsTokenCount": 21, + } + assert parse_usage(usage) == { + "input_tokens": 9, + "output_tokens": 2, + "total_tokens": 32, + "cached_tokens": 4, + "cache_write_tokens": None, + "reasoning_tokens": 21, + } + + def test_parse_usage_absent_detail_fields_are_none_not_zero() -> None: # Only the top-level counts present: the detail fields are ABSENT, so None — # distinct from the fixture's present-but-zero 0 (AC2). @@ -459,6 +479,8 @@ def test_usage_mapping_reads_both_shapes_and_rejects_others() -> None: assert usage_mapping({"usage": {"prompt_tokens": 1}}) == {"prompt_tokens": 1} # The Responses API terminal event nests usage under `response`. assert usage_mapping({"response": {"usage": {"input_tokens": 2}}}) == {"input_tokens": 2} + # A Gemini SSE chunk carries usageMetadata at the top level. + assert usage_mapping({"usageMetadata": {"promptTokenCount": 3}}) == {"promptTokenCount": 3} assert usage_mapping({"type": "response.output_text.delta"}) is None assert usage_mapping("not-a-dict") is None diff --git a/python/tests/unit/test_verify_frameworks.py b/python/tests/unit/test_verify_frameworks.py index 8a5e08fc..ef794032 100644 --- a/python/tests/unit/test_verify_frameworks.py +++ b/python/tests/unit/test_verify_frameworks.py @@ -118,9 +118,11 @@ def test_each_framework_checks_a_distribution_its_extra_installs() -> None: # a distribution the extra doesn't ship would read as NOT INSTALLED and exit 0. requirements = [Requirement(r) for r in importlib.metadata.requires("donkey-kit") or []] for key, _, _, _, distribution in vf.FRAMEWORKS: + # A dotted key is a second factory of the same framework (adk.gemini). + framework = key.partition(".")[0] extra = { canonicalize_name(r.name) for r in requirements - if r.marker and r.marker.evaluate({"extra": canonicalize_name(key)}) + if r.marker and r.marker.evaluate({"extra": canonicalize_name(framework)}) } assert canonicalize_name(distribution) in extra, (key, distribution, sorted(extra)) diff --git a/website/components/adapter-error-taxonomy-note.mdx b/website/components/adapter-error-taxonomy-note.mdx index 46c369fb..c12e799a 100644 --- a/website/components/adapter-error-taxonomy-note.mdx +++ b/website/components/adapter-error-taxonomy-note.mdx @@ -8,7 +8,8 @@ Intentional overrides that keep their own inline wording (documented, not silent): • langgraph — prefixes a proxy-surface caveat before this sentence. - • adk, crewai — "surface through ADK's `LiteLlm` model" / "surface through CrewAI". + • adk, crewai — "surface through ADK's `LiteLlm` model" (plus a pointer to the + native-Gemini `classify(exc.response)` section) / "surface through CrewAI". • agent-framework — references the PolicyViolation hierarchy / middleware shape. */} diff --git a/website/content/examples/adk.mdx b/website/content/examples/adk.mdx index e59c7cb5..7ea7dc89 100644 --- a/website/content/examples/adk.mdx +++ b/website/content/examples/adk.mdx @@ -73,8 +73,10 @@ should see:** `APIError 403` and the first line of the proxy's message — **not `PIIDetected`. Without the policy it prints `NO REFUSAL`. - If you need typed refusals, `last_call` or run ids with ADK today, prefer a - framework path where the SDK owns the transport, such as + If you need typed refusals, `last_call` or run ids with ADK, use + `donkey.adk.gemini("gemini-2.5-flash")` on a `Format=Gemini` proxy — the SDK + owns that transport (see [Native Gemini](/frameworks/adk#native-gemini)). + Otherwise prefer a framework path where the SDK owns the transport, such as [OpenAI](/examples/openai) or [LangGraph](/examples/langgraph). A `404` here means the proxy's upstream has no `/chat/completions` route. diff --git a/website/content/examples/gemini.mdx b/website/content/examples/gemini.mdx index 3d13f82a..7ec1737b 100644 --- a/website/content/examples/gemini.mdx +++ b/website/content/examples/gemini.mdx @@ -4,9 +4,11 @@ description: Two plain-httpx scripts against a Format=Gemini proxy — native ge # Gemini -No Gemini adapter ships, so these are plain `httpx` against a proxy -provisioned **`Format=Gemini`** (for example `ddk-gemini-inbound`) with the -same `client_id` / `client_secret` pair. The route is +These scripts show the wire: plain `httpx` against a proxy provisioned +**`Format=Gemini`** (for example `ddk-gemini-inbound`) with the same +`client_id` / `client_secret` pair. For an agent, ADK's native `Gemini` model is +bound to the same proxy by `donkey.adk.gemini("gemini-2.5-flash")` — see +[Native Gemini](/frameworks/adk#native-gemini). The route is `/models/:generateContent`. `DonkeyConfig` still resolves and validates the credentials, and `classify()` still types the errors. diff --git a/website/content/feature-overview.mdx b/website/content/feature-overview.mdx index b1e7d91c..046d80e1 100644 --- a/website/content/feature-overview.mdx +++ b/website/content/feature-overview.mdx @@ -44,6 +44,7 @@ attaches there once, so you never wire it call by call. |---|---|---| | LangGraph | `donkey.langgraph.chat_model("gpt-4o")` | `langchain_openai.ChatOpenAI` | | Google ADK | `donkey.adk.model("gpt-4o")` | `LiteLlm` | +| Google ADK on a `Format=Gemini` proxy | `donkey.adk.gemini("gemini-2.5-flash")` | `google.adk.models.Gemini` | | Strands | `donkey.strands.model("gpt-4o")` | `OpenAIModel` | | MS Agent Framework | `donkey.agent_framework.chat_client("gpt-4o")` | Agent Framework chat client | | LlamaIndex | `donkey.llamaindex.llm("gpt-4o")` | `OpenAILike` | diff --git a/website/content/frameworks/adk.mdx b/website/content/frameworks/adk.mdx index c7316826..dde205e3 100644 --- a/website/content/frameworks/adk.mdx +++ b/website/content/frameworks/adk.mdx @@ -1,5 +1,5 @@ --- -description: Use a governed LiteLlm model in Google's Agent Development Kit (ADK), pointed at your Omni Gateway LLM proxy. +description: Use a governed LiteLlm model, or ADK's native Gemini model on a Format=Gemini proxy, in Google's Agent Development Kit (ADK), pointed at your Omni Gateway LLM proxy. --- import { Tabs } from 'nextra/components' @@ -7,18 +7,28 @@ import OpenAiProxyTs from '../../components/openai-proxy-ts.mdx' # Google ADK -Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy through -ADK's `LiteLlm` model wrapper. The adapter translates the governed connection -into LiteLLM's own model-string and kwarg conventions for you. +Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy in one +of two ways, depending on the proxy's ingress **Format**: + +- `model()` — ADK's `LiteLlm` model wrapper, for a `Format=OpenAI` proxy (the + default). The adapter translates the governed connection into LiteLLM's own + model-string and kwarg conventions for you. +- `gemini()` — ADK's native `Gemini` model, for a `Format=Gemini` proxy. The + SDK's shared HTTP client is injected, so every governed feature works: per-run + correlation, spans, usage and `donkey.last_call`. See + [Native Gemini](#native-gemini). **What you get** -- A native `google.adk.models.lite_llm.LiteLlm`, with the proxy auth and - attribution headers set. -- The `openai/` model prefix and LiteLLM kwarg names handled automatically. -- Supported at `connection_kwargs()`. ADK's `LiteLlm` model makes the HTTP +- A native `google.adk.models.lite_llm.LiteLlm` or `google.adk.models.Gemini`, + with the proxy auth and attribution headers set. +- For `model()`: the `openai/` model prefix and LiteLLM kwarg names handled + automatically. Supported at `connection_kwargs()`. LiteLLM makes the HTTP calls itself, so correlation is per client and `donkey.last_call` is not populated (see [Notes](#notes)). +- For `gemini()`: the shared client injected through + `HttpOptions.httpx_async_client`, round-trip verified live against a + `Format=Gemini` proxy. ## Install @@ -99,20 +109,144 @@ llm = LiteLlm( LiteLLM uses `api_base` and `extra_headers`, not `base_url` / `default_headers` — `connection_kwargs()` already translates for you. +## Native Gemini + +A proxy provisioned with **Format = Gemini** exposes the native Gemini API at +`POST //models/:generateContent` (and +`:streamGenerateContent`). ADK's own `Gemini` model speaks that API, so +`gemini()` returns a native `google.adk.models.Gemini` bound to the proxy, with +the SDK's shared HTTP client injected. + +The three forms mirror `model()`. Pass the bare Gemini model id — there is no +prefix, because the URL path carries the model: + +```python +from donkey_kit import Donkey +from google.adk.models import Gemini + +async with Donkey.from_env() as donkey: + # 1. Off a shared Donkey instance + llm = donkey.adk.gemini("gemini-2.5-flash") + + # 3. Governed kwargs, native constructor + llm = Gemini(model="gemini-2.5-flash", **donkey.adk.gemini_connection_kwargs()) +``` + +```python +# 2. Module-level factory +from donkey_kit.integrations.adk import gemini + +llm = gemini("gemini-2.5-flash") +``` + +**Pointing at the Gemini proxy.** `DONKEY_LLM_PROXY_URL` usually names a +`Format=OpenAI` proxy. If your Gemini proxy is a different one, pass its URL +as `base_url`; the `client_id`/`client_secret` pair must be contracted on that +proxy: + +```python +llm = donkey.adk.gemini("gemini-2.5-flash", base_url="https://…/ddk-gemini-inbound/") +``` + +Any other keyword is passed to ADK's `Gemini` and overrides the governed +default — for example `retry_options`, or your own `client_kwargs`. + +**Manual equivalent.** This is what `gemini_connection_kwargs()` returns: + +```python +from google.adk.models import Gemini + +llm = Gemini( + model="gemini-2.5-flash", + base_url=..., # the Format=Gemini proxy URL + client_kwargs={ + "api_key": ..., # a placeholder; google-genai requires one + "http_options": { + "base_url": ..., # same URL + "api_version": "", # the proxy path has no /v1beta segment + "headers": ..., # client_id / client_secret header pair + "timeout": ..., # milliseconds + "httpx_async_client": ..., # the SDK's shared client + }, + }, +) +``` + +`client_kwargs` replaces ADK's default HTTP options wholesale, so all five +`http_options` keys are passed together. `google-genai` requires an API key and +always sends it as `x-goog-api-key`; the gateway authenticates on the +`client_id`/`client_secret` pair and ignores it. + +**`donkey.last_call`.** Usage (`input_tokens`, `output_tokens`, +`total_tokens`, `cached_tokens`, `reasoning_tokens`) is read from Gemini's +`usageMetadata`, and `requested_model` from the URL path. Gemini's +`total_tokens` includes the thinking tokens it also reports as +`reasoning_tokens`. A `Format=Gemini` proxy is a passthrough, so it sends no +routing headers: `served_provider`, `served_model` and `routing_type` stay +`None`, and `substituted` is `False`. `request_id` and `api_instance_id` are +populated. + +`last_call` is contextvar-scoped, and ADK's `Runner` makes the model call in a +task of its own, so the caller of `runner.run_async(...)` reads `UNOBSERVED`. +Read it in an `after_model_callback`, which runs in the same task as the call: + +```python +from google.adk.agents import LlmAgent + +def record_usage(callback_context, llm_response): + r = donkey.last_call + print(r.requested_model, r.input_tokens, r.output_tokens) + return None # keep the model's response + +agent = LlmAgent( + name="assistant", + model=donkey.adk.gemini("gemini-2.5-flash"), + after_model_callback=record_usage, +) +``` + +When streaming, the callback runs once per partial response; usage lands on the +last one. + +**Errors.** `google-genai` raises its own `google.genai.errors.APIError` +(`ClientError` for a 4xx) and the SDK does not wrap it. The error's `.response` +is the proxy's HTTP response, so `classify()` gives you the typed Donkey error: + +```python +from donkey_kit.core.errors import classify +from google.genai import errors as genai_errors + +try: + async for event in runner.run_async(...): + ... +except genai_errors.APIError as exc: + err = classify(exc.response) # e.g. AuthError on 401, UpstreamRequestError on 404 +``` + +**One `Donkey` per event loop.** The shared HTTP client's connection pool is +bound to the event loop that first uses it. Build the `Donkey` — and the +models — inside the loop that runs them (`async with Donkey.from_env()` in your +`main`), and don't reuse the module-level `gemini()` across separate +`asyncio.run(...)` calls: the second run fails with +`RuntimeError: Event loop is closed`. + ## Notes -- **Correlation IDs are per-client, not per-run.** ADK sends requests through - its built-in LiteLLM model layer rather than the SDK's shared HTTP client, so - the correlation ID is set once per client instead of per `donkey.run()`. Every - governance header is still sent on every request. The conformance suite - checks this as a documented behaviour. -- **`donkey.last_call` is unavailable.** Because the response is handled by LiteLLM, - gateway identity, routing, and usage fields can't be observed. When every - adapter resolved on a `Donkey` is like this one, `donkey.last_call` reports - `status == LastCallStatus.UNAVAILABLE` and `available == False`, and names the - resolved adapters in `surface`. +- **Correlation IDs are per-client, not per-run (`model()` only).** `model()` + sends requests through ADK's built-in LiteLLM model layer rather than the + SDK's shared HTTP client, so the correlation ID is set once per client instead + of per `donkey.run()`. Every governance header is still sent on every request. + The conformance suite checks this as a documented behaviour. `gemini()` uses + the shared client, so its correlation ID is per run. +- **`donkey.last_call` is unavailable (`model()` only).** Because the response + is handled by LiteLLM, gateway identity, routing, and usage fields can't be + observed. When every adapter resolved on a `Donkey` is like this one, + `donkey.last_call` reports `status == LastCallStatus.UNAVAILABLE` and + `available == False`, and names the resolved adapters in `surface`. Once + `gemini()` has been called on a `Donkey`, the ADK adapter observes calls, so an + empty record reads `UNOBSERVED` instead. - `google-adk` requires `litellm>=1.84` as a floor, not a ceiling — pin your own upper bound if you need one. See the [error taxonomy](/errors) for how proxy rejections surface through -ADK's `LiteLlm` model. +ADK's `LiteLlm` model, and [Native Gemini](#native-gemini) for `gemini()`. diff --git a/website/content/frameworks/index.mdx b/website/content/frameworks/index.mdx index e81ce97d..a40b735f 100644 --- a/website/content/frameworks/index.mdx +++ b/website/content/frameworks/index.mdx @@ -35,7 +35,7 @@ object. `chat_model()` → `langchain_openai.ChatOpenAI` - `model()` → `google.adk … LiteLlm` + `model()` → `google.adk … LiteLlm`; `gemini()` → `google.adk.models.Gemini` `model()` → `strands … OpenAIModel` @@ -79,8 +79,8 @@ constructor call the factory makes for you. ## Match the adapter to your proxy's wire format -Every adapter on this page except Anthropic — and the raw `donkey.llm.client()` -— speaks the **OpenAI wire format**. The format your proxy accepts is the +Every adapter on this page except Anthropic and ADK's `gemini()` — and the raw +`donkey.llm.client()` — speaks the **OpenAI wire format**. The format your proxy accepts is the **Format** (OpenAI / Anthropic / Gemini) chosen when the proxy was provisioned. It is a property of the proxy, not an SDK setting, so there is no config field for it: pick the adapter that matches your proxy. @@ -89,7 +89,7 @@ for it: pick the adapter that matches your proxy. |---|---| | **OpenAI** | `donkey.llm.client()` or any framework adapter. Default DDK proxies are `Format=OpenAI`. | | **Anthropic** | `donkey.anthropic.client()` (native `AsyncAnthropic`). The proxy serves the native Messages route at `POST //v1/messages`; OpenAI-shape `/chat/completions` returns 404. | -| **Gemini** | No SDK adapter. Point a native `google-genai` client at `…/models/:generateContent` yourself, or reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). | +| **Gemini** | `donkey.adk.gemini("gemini-2.5-flash")` (ADK's native `Gemini` model — see [Native Gemini](/frameworks/adk#native-gemini)). The proxy serves `POST //models/:generateContent` and `:streamGenerateContent`. There is no standalone `google-genai` adapter; you can also reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). | **Ingress Format is not the same as the upstream provider.** The ingress @@ -131,8 +131,9 @@ HTTP client is also used, which adds per-run correlation IDs and | Anthropic SDK | ✅ | ✅ | Returns a bare `client()`, not a model-bound object — see the [Anthropic page](/frameworks/anthropic). | | LlamaIndex | ✅ | ❌ | Static `default_headers` snapshot: no per-run correlation or `donkey.last_call`. `is_chat_model=True` is forced. | | MS Agent Framework | ✅ | ❌ | Static `default_headers` snapshot: no per-run correlation or `donkey.last_call`. | -| Google ADK | ✅ (`extra_headers`) | ❌ | Calls go through ADK's `LiteLlm` model: correlation is per client and `donkey.last_call` is not populated. | -| CrewAI | ✅ (`extra_headers`) | ❌ | Calls go through CrewAI's native OpenAI provider, which builds its own HTTP client: same behaviour as Google ADK. | +| Google ADK — `model()` | ✅ (`extra_headers`) | ❌ | Calls go through ADK's `LiteLlm` model: correlation is per client and `donkey.last_call` is not populated. | +| Google ADK — `gemini()` | ✅ | ✅ | Via `HttpOptions.httpx_async_client`, on a `Format=Gemini` proxy. | +| CrewAI | ✅ (`extra_headers`) | ❌ | Calls go through CrewAI's native OpenAI provider, which builds its own HTTP client: same behaviour as ADK's `model()`. | See the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/develop/docs/verified-apis.md) for how each constructor signature the adapters depend on is checked. diff --git a/website/content/reference/last-call.mdx b/website/content/reference/last-call.mdx index 3d1635b9..23fee2c9 100644 --- a/website/content/reference/last-call.mdx +++ b/website/content/reference/last-call.mdx @@ -44,7 +44,7 @@ even on a cold read. `donkey.last_call` is populated only when the governed response passes through the SDK's shared httpx client. Adapters that route outside that response path - (ADK and CrewAI, whose framework owns the transport, or the + (ADK's `model()` and CrewAI, whose framework owns the transport, or the `default_headers`-only adapters) report `UNAVAILABLE` with the surface named. See [When `last_call` is unavailable](/telemetry#when-last_call-is-unavailable). diff --git a/website/content/telemetry.mdx b/website/content/telemetry.mdx index 431b74f2..84ef210f 100644 --- a/website/content/telemetry.mdx +++ b/website/content/telemetry.mdx @@ -394,15 +394,19 @@ through every field. `donkey.last_call` is populated only when the governed response passes through the SDK's shared httpx client. Four `connection_kwargs()`-only adapters route -outside that response path: ADK sends requests through LiteLLM and CrewAI -through its native OpenAI provider, while LlamaIndex and Microsoft Agent -Framework receive only `default_headers`. +outside that response path: ADK's `model()` sends requests through LiteLLM and +CrewAI through its native OpenAI provider, while LlamaIndex and Microsoft Agent +Framework receive only `default_headers`. ADK's `gemini()` uses the shared +client, so once it has been called the ADK adapter observes calls (see +[Native Gemini](/frameworks/adk#native-gemini) for reading `last_call` inside an +ADK run). That static snapshot excludes the correlation ID bound later by `donkey.run(id=...)`, so those two adapters also do not propagate the run's correlation ID. When every adapter resolved on a `Donkey` is one of those four, a cold read -reports the limitation explicitly. For a `Donkey` that resolved only ADK: +reports the limitation explicitly. For a `Donkey` that resolved only ADK's +`model()`: ```python r = donkey.last_call diff --git a/website/public/examples/adk.md b/website/public/examples/adk.md index 4a6e3b52..a141ed76 100644 --- a/website/public/examples/adk.md +++ b/website/public/examples/adk.md @@ -66,8 +66,10 @@ except litellm.exceptions.APIError as err: should see:** `APIError 403` and the first line of the proxy's message — **not** `PIIDetected`. Without the policy it prints `NO REFUSAL`. - If you need typed refusals, `last_call` or run ids with ADK today, prefer a - framework path where the SDK owns the transport, such as + If you need typed refusals, `last_call` or run ids with ADK, use + `donkey.adk.gemini("gemini-2.5-flash")` on a `Format=Gemini` proxy — the SDK + owns that transport (see [Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini)). + Otherwise prefer a framework path where the SDK owns the transport, such as [OpenAI](https://donkey-development-kit.github.io/donkey-development-kit/examples/openai.md) or [LangGraph](https://donkey-development-kit.github.io/donkey-development-kit/examples/langgraph.md). A `404` here means the proxy's upstream has no `/chat/completions` route. diff --git a/website/public/examples/gemini.md b/website/public/examples/gemini.md index a1a6a81a..7669cdc5 100644 --- a/website/public/examples/gemini.md +++ b/website/public/examples/gemini.md @@ -1,8 +1,10 @@ # Gemini -No Gemini adapter ships, so these are plain `httpx` against a proxy -provisioned **`Format=Gemini`** (for example `ddk-gemini-inbound`) with the -same `client_id` / `client_secret` pair. The route is +These scripts show the wire: plain `httpx` against a proxy provisioned +**`Format=Gemini`** (for example `ddk-gemini-inbound`) with the same +`client_id` / `client_secret` pair. For an agent, ADK's native `Gemini` model is +bound to the same proxy by `donkey.adk.gemini("gemini-2.5-flash")` — see +[Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini). The route is `/models/:generateContent`. `DonkeyConfig` still resolves and validates the credentials, and `classify()` still types the errors. diff --git a/website/public/feature-overview.md b/website/public/feature-overview.md index b79944e3..f873ef7b 100644 --- a/website/public/feature-overview.md +++ b/website/public/feature-overview.md @@ -29,6 +29,7 @@ attaches there once, so you never wire it call by call. |---|---|---| | LangGraph | `donkey.langgraph.chat_model("gpt-4o")` | `langchain_openai.ChatOpenAI` | | Google ADK | `donkey.adk.model("gpt-4o")` | `LiteLlm` | +| Google ADK on a `Format=Gemini` proxy | `donkey.adk.gemini("gemini-2.5-flash")` | `google.adk.models.Gemini` | | Strands | `donkey.strands.model("gpt-4o")` | `OpenAIModel` | | MS Agent Framework | `donkey.agent_framework.chat_client("gpt-4o")` | Agent Framework chat client | | LlamaIndex | `donkey.llamaindex.llm("gpt-4o")` | `OpenAILike` | diff --git a/website/public/frameworks.md b/website/public/frameworks.md index c1a0da7d..7b154600 100644 --- a/website/public/frameworks.md +++ b/website/public/frameworks.md @@ -25,7 +25,7 @@ object. `chat_model()` → `langchain_openai.ChatOpenAI` - `model()` → `google.adk … LiteLlm` + `model()` → `google.adk … LiteLlm`; `gemini()` → `google.adk.models.Gemini` `model()` → `strands … OpenAIModel` @@ -68,8 +68,8 @@ constructor call the factory makes for you. ## Match the adapter to your proxy's wire format -Every adapter on this page except Anthropic — and the raw `donkey.llm.client()` -— speaks the **OpenAI wire format**. The format your proxy accepts is the +Every adapter on this page except Anthropic and ADK's `gemini()` — and the raw +`donkey.llm.client()` — speaks the **OpenAI wire format**. The format your proxy accepts is the **Format** (OpenAI / Anthropic / Gemini) chosen when the proxy was provisioned. It is a property of the proxy, not an SDK setting, so there is no config field for it: pick the adapter that matches your proxy. @@ -78,7 +78,7 @@ for it: pick the adapter that matches your proxy. |---|---| | **OpenAI** | `donkey.llm.client()` or any framework adapter. Default DDK proxies are `Format=OpenAI`. | | **Anthropic** | `donkey.anthropic.client()` (native `AsyncAnthropic`). The proxy serves the native Messages route at `POST //v1/messages`; OpenAI-shape `/chat/completions` returns 404. | -| **Gemini** | No SDK adapter. Point a native `google-genai` client at `…/models/:generateContent` yourself, or reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). | +| **Gemini** | `donkey.adk.gemini("gemini-2.5-flash")` (ADK's native `Gemini` model — see [Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini)). The proxy serves `POST //models/:generateContent` and `:streamGenerateContent`. There is no standalone `google-genai` adapter; you can also reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). | **Ingress Format is not the same as the upstream provider.** The ingress Format is the wire protocol *your request* speaks to the proxy. The upstream @@ -118,8 +118,9 @@ HTTP client is also used, which adds per-run correlation IDs and | Anthropic SDK | ✅ | ✅ | Returns a bare `client()`, not a model-bound object — see the [Anthropic page](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/anthropic.md). | | LlamaIndex | ✅ | ❌ | Static `default_headers` snapshot: no per-run correlation or `donkey.last_call`. `is_chat_model=True` is forced. | | MS Agent Framework | ✅ | ❌ | Static `default_headers` snapshot: no per-run correlation or `donkey.last_call`. | -| Google ADK | ✅ (`extra_headers`) | ❌ | Calls go through ADK's `LiteLlm` model: correlation is per client and `donkey.last_call` is not populated. | -| CrewAI | ✅ (`extra_headers`) | ❌ | Calls go through CrewAI's native OpenAI provider, which builds its own HTTP client: same behaviour as Google ADK. | +| Google ADK — `model()` | ✅ (`extra_headers`) | ❌ | Calls go through ADK's `LiteLlm` model: correlation is per client and `donkey.last_call` is not populated. | +| Google ADK — `gemini()` | ✅ | ✅ | Via `HttpOptions.httpx_async_client`, on a `Format=Gemini` proxy. | +| CrewAI | ✅ (`extra_headers`) | ❌ | Calls go through CrewAI's native OpenAI provider, which builds its own HTTP client: same behaviour as ADK's `model()`. | See the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/develop/docs/verified-apis.md) for how each constructor signature the adapters depend on is checked. diff --git a/website/public/frameworks/adk.md b/website/public/frameworks/adk.md index 3d647f38..a9bd1dad 100644 --- a/website/public/frameworks/adk.md +++ b/website/public/frameworks/adk.md @@ -1,17 +1,27 @@ # Google ADK -Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy through -ADK's `LiteLlm` model wrapper. The adapter translates the governed connection -into LiteLLM's own model-string and kwarg conventions for you. +Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy in one +of two ways, depending on the proxy's ingress **Format**: + +- `model()` — ADK's `LiteLlm` model wrapper, for a `Format=OpenAI` proxy (the + default). The adapter translates the governed connection into LiteLLM's own + model-string and kwarg conventions for you. +- `gemini()` — ADK's native `Gemini` model, for a `Format=Gemini` proxy. The + SDK's shared HTTP client is injected, so every governed feature works: per-run + correlation, spans, usage and `donkey.last_call`. See + [Native Gemini](#native-gemini). **What you get** -- A native `google.adk.models.lite_llm.LiteLlm`, with the proxy auth and - attribution headers set. -- The `openai/` model prefix and LiteLLM kwarg names handled automatically. -- Supported at `connection_kwargs()`. ADK's `LiteLlm` model makes the HTTP +- A native `google.adk.models.lite_llm.LiteLlm` or `google.adk.models.Gemini`, + with the proxy auth and attribution headers set. +- For `model()`: the `openai/` model prefix and LiteLLM kwarg names handled + automatically. Supported at `connection_kwargs()`. LiteLLM makes the HTTP calls itself, so correlation is per client and `donkey.last_call` is not populated (see [Notes](#notes)). +- For `gemini()`: the shared client injected through + `HttpOptions.httpx_async_client`, round-trip verified live against a + `Format=Gemini` proxy. ## Install @@ -100,20 +110,144 @@ llm = LiteLlm( LiteLLM uses `api_base` and `extra_headers`, not `base_url` / `default_headers` — `connection_kwargs()` already translates for you. +## Native Gemini + +A proxy provisioned with **Format = Gemini** exposes the native Gemini API at +`POST //models/:generateContent` (and +`:streamGenerateContent`). ADK's own `Gemini` model speaks that API, so +`gemini()` returns a native `google.adk.models.Gemini` bound to the proxy, with +the SDK's shared HTTP client injected. + +The three forms mirror `model()`. Pass the bare Gemini model id — there is no +prefix, because the URL path carries the model: + +```python +from donkey_kit import Donkey +from google.adk.models import Gemini + +async with Donkey.from_env() as donkey: + # 1. Off a shared Donkey instance + llm = donkey.adk.gemini("gemini-2.5-flash") + + # 3. Governed kwargs, native constructor + llm = Gemini(model="gemini-2.5-flash", **donkey.adk.gemini_connection_kwargs()) +``` + +```python +# 2. Module-level factory +from donkey_kit.integrations.adk import gemini + +llm = gemini("gemini-2.5-flash") +``` + +**Pointing at the Gemini proxy.** `DONKEY_LLM_PROXY_URL` usually names a +`Format=OpenAI` proxy. If your Gemini proxy is a different one, pass its URL +as `base_url`; the `client_id`/`client_secret` pair must be contracted on that +proxy: + +```python +llm = donkey.adk.gemini("gemini-2.5-flash", base_url="https://…/ddk-gemini-inbound/") +``` + +Any other keyword is passed to ADK's `Gemini` and overrides the governed +default — for example `retry_options`, or your own `client_kwargs`. + +**Manual equivalent.** This is what `gemini_connection_kwargs()` returns: + +```python +from google.adk.models import Gemini + +llm = Gemini( + model="gemini-2.5-flash", + base_url=..., # the Format=Gemini proxy URL + client_kwargs={ + "api_key": ..., # a placeholder; google-genai requires one + "http_options": { + "base_url": ..., # same URL + "api_version": "", # the proxy path has no /v1beta segment + "headers": ..., # client_id / client_secret header pair + "timeout": ..., # milliseconds + "httpx_async_client": ..., # the SDK's shared client + }, + }, +) +``` + +`client_kwargs` replaces ADK's default HTTP options wholesale, so all five +`http_options` keys are passed together. `google-genai` requires an API key and +always sends it as `x-goog-api-key`; the gateway authenticates on the +`client_id`/`client_secret` pair and ignores it. + +**`donkey.last_call`.** Usage (`input_tokens`, `output_tokens`, +`total_tokens`, `cached_tokens`, `reasoning_tokens`) is read from Gemini's +`usageMetadata`, and `requested_model` from the URL path. Gemini's +`total_tokens` includes the thinking tokens it also reports as +`reasoning_tokens`. A `Format=Gemini` proxy is a passthrough, so it sends no +routing headers: `served_provider`, `served_model` and `routing_type` stay +`None`, and `substituted` is `False`. `request_id` and `api_instance_id` are +populated. + +`last_call` is contextvar-scoped, and ADK's `Runner` makes the model call in a +task of its own, so the caller of `runner.run_async(...)` reads `UNOBSERVED`. +Read it in an `after_model_callback`, which runs in the same task as the call: + +```python +from google.adk.agents import LlmAgent + +def record_usage(callback_context, llm_response): + r = donkey.last_call + print(r.requested_model, r.input_tokens, r.output_tokens) + return None # keep the model's response + +agent = LlmAgent( + name="assistant", + model=donkey.adk.gemini("gemini-2.5-flash"), + after_model_callback=record_usage, +) +``` + +When streaming, the callback runs once per partial response; usage lands on the +last one. + +**Errors.** `google-genai` raises its own `google.genai.errors.APIError` +(`ClientError` for a 4xx) and the SDK does not wrap it. The error's `.response` +is the proxy's HTTP response, so `classify()` gives you the typed Donkey error: + +```python +from donkey_kit.core.errors import classify +from google.genai import errors as genai_errors + +try: + async for event in runner.run_async(...): + ... +except genai_errors.APIError as exc: + err = classify(exc.response) # e.g. AuthError on 401, UpstreamRequestError on 404 +``` + +**One `Donkey` per event loop.** The shared HTTP client's connection pool is +bound to the event loop that first uses it. Build the `Donkey` — and the +models — inside the loop that runs them (`async with Donkey.from_env()` in your +`main`), and don't reuse the module-level `gemini()` across separate +`asyncio.run(...)` calls: the second run fails with +`RuntimeError: Event loop is closed`. + ## Notes -- **Correlation IDs are per-client, not per-run.** ADK sends requests through - its built-in LiteLLM model layer rather than the SDK's shared HTTP client, so - the correlation ID is set once per client instead of per `donkey.run()`. Every - governance header is still sent on every request. The conformance suite - checks this as a documented behaviour. -- **`donkey.last_call` is unavailable.** Because the response is handled by LiteLLM, - gateway identity, routing, and usage fields can't be observed. When every - adapter resolved on a `Donkey` is like this one, `donkey.last_call` reports - `status == LastCallStatus.UNAVAILABLE` and `available == False`, and names the - resolved adapters in `surface`. +- **Correlation IDs are per-client, not per-run (`model()` only).** `model()` + sends requests through ADK's built-in LiteLLM model layer rather than the + SDK's shared HTTP client, so the correlation ID is set once per client instead + of per `donkey.run()`. Every governance header is still sent on every request. + The conformance suite checks this as a documented behaviour. `gemini()` uses + the shared client, so its correlation ID is per run. +- **`donkey.last_call` is unavailable (`model()` only).** Because the response + is handled by LiteLLM, gateway identity, routing, and usage fields can't be + observed. When every adapter resolved on a `Donkey` is like this one, + `donkey.last_call` reports `status == LastCallStatus.UNAVAILABLE` and + `available == False`, and names the resolved adapters in `surface`. Once + `gemini()` has been called on a `Donkey`, the ADK adapter observes calls, so an + empty record reads `UNOBSERVED` instead. - `google-adk` requires `litellm>=1.84` as a floor, not a ceiling — pin your own upper bound if you need one. See the [error taxonomy](https://donkey-development-kit.github.io/donkey-development-kit/errors.md) for how proxy rejections surface through -ADK's `LiteLlm` model. +ADK's `LiteLlm` model, and [Native Gemini](#native-gemini) for `gemini()`. diff --git a/website/public/llms-full.txt b/website/public/llms-full.txt index a83bf680..42abccd0 100644 --- a/website/public/llms-full.txt +++ b/website/public/llms-full.txt @@ -583,6 +583,7 @@ attaches there once, so you never wire it call by call. |---|---|---| | LangGraph | `donkey.langgraph.chat_model("gpt-4o")` | `langchain_openai.ChatOpenAI` | | Google ADK | `donkey.adk.model("gpt-4o")` | `LiteLlm` | +| Google ADK on a `Format=Gemini` proxy | `donkey.adk.gemini("gemini-2.5-flash")` | `google.adk.models.Gemini` | | Strands | `donkey.strands.model("gpt-4o")` | `OpenAIModel` | | MS Agent Framework | `donkey.agent_framework.chat_client("gpt-4o")` | Agent Framework chat client | | LlamaIndex | `donkey.llamaindex.llm("gpt-4o")` | `OpenAILike` | @@ -728,7 +729,7 @@ object. `chat_model()` → `langchain_openai.ChatOpenAI` - `model()` → `google.adk … LiteLlm` + `model()` → `google.adk … LiteLlm`; `gemini()` → `google.adk.models.Gemini` `model()` → `strands … OpenAIModel` @@ -771,8 +772,8 @@ constructor call the factory makes for you. ## Match the adapter to your proxy's wire format -Every adapter on this page except Anthropic — and the raw `donkey.llm.client()` -— speaks the **OpenAI wire format**. The format your proxy accepts is the +Every adapter on this page except Anthropic and ADK's `gemini()` — and the raw +`donkey.llm.client()` — speaks the **OpenAI wire format**. The format your proxy accepts is the **Format** (OpenAI / Anthropic / Gemini) chosen when the proxy was provisioned. It is a property of the proxy, not an SDK setting, so there is no config field for it: pick the adapter that matches your proxy. @@ -781,7 +782,7 @@ for it: pick the adapter that matches your proxy. |---|---| | **OpenAI** | `donkey.llm.client()` or any framework adapter. Default DDK proxies are `Format=OpenAI`. | | **Anthropic** | `donkey.anthropic.client()` (native `AsyncAnthropic`). The proxy serves the native Messages route at `POST //v1/messages`; OpenAI-shape `/chat/completions` returns 404. | -| **Gemini** | No SDK adapter. Point a native `google-genai` client at `…/models/:generateContent` yourself, or reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). | +| **Gemini** | `donkey.adk.gemini("gemini-2.5-flash")` (ADK's native `Gemini` model — see [Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini)). The proxy serves `POST //models/:generateContent` and `:streamGenerateContent`. There is no standalone `google-genai` adapter; you can also reach Gemini as an upstream provider behind an OpenAI-format proxy (see below). | **Ingress Format is not the same as the upstream provider.** The ingress Format is the wire protocol *your request* speaks to the proxy. The upstream @@ -821,8 +822,9 @@ HTTP client is also used, which adds per-run correlation IDs and | Anthropic SDK | ✅ | ✅ | Returns a bare `client()`, not a model-bound object — see the [Anthropic page](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/anthropic.md). | | LlamaIndex | ✅ | ❌ | Static `default_headers` snapshot: no per-run correlation or `donkey.last_call`. `is_chat_model=True` is forced. | | MS Agent Framework | ✅ | ❌ | Static `default_headers` snapshot: no per-run correlation or `donkey.last_call`. | -| Google ADK | ✅ (`extra_headers`) | ❌ | Calls go through ADK's `LiteLlm` model: correlation is per client and `donkey.last_call` is not populated. | -| CrewAI | ✅ (`extra_headers`) | ❌ | Calls go through CrewAI's native OpenAI provider, which builds its own HTTP client: same behaviour as Google ADK. | +| Google ADK — `model()` | ✅ (`extra_headers`) | ❌ | Calls go through ADK's `LiteLlm` model: correlation is per client and `donkey.last_call` is not populated. | +| Google ADK — `gemini()` | ✅ | ✅ | Via `HttpOptions.httpx_async_client`, on a `Format=Gemini` proxy. | +| CrewAI | ✅ (`extra_headers`) | ❌ | Calls go through CrewAI's native OpenAI provider, which builds its own HTTP client: same behaviour as ADK's `model()`. | See the [verification ledger](https://github.com/Donkey-Development-Kit/donkey-development-kit/blob/develop/docs/verified-apis.md) for how each constructor signature the adapters depend on is checked. @@ -1089,18 +1091,28 @@ Source: https://donkey-development-kit.github.io/donkey-development-kit/framewor # Google ADK -Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy through -ADK's `LiteLlm` model wrapper. The adapter translates the governed connection -into LiteLLM's own model-string and kwarg conventions for you. +Google's Agent Development Kit (ADK) reaches the Agent Fabric LLM proxy in one +of two ways, depending on the proxy's ingress **Format**: + +- `model()` — ADK's `LiteLlm` model wrapper, for a `Format=OpenAI` proxy (the + default). The adapter translates the governed connection into LiteLLM's own + model-string and kwarg conventions for you. +- `gemini()` — ADK's native `Gemini` model, for a `Format=Gemini` proxy. The + SDK's shared HTTP client is injected, so every governed feature works: per-run + correlation, spans, usage and `donkey.last_call`. See + [Native Gemini](#native-gemini). **What you get** -- A native `google.adk.models.lite_llm.LiteLlm`, with the proxy auth and - attribution headers set. -- The `openai/` model prefix and LiteLLM kwarg names handled automatically. -- Supported at `connection_kwargs()`. ADK's `LiteLlm` model makes the HTTP +- A native `google.adk.models.lite_llm.LiteLlm` or `google.adk.models.Gemini`, + with the proxy auth and attribution headers set. +- For `model()`: the `openai/` model prefix and LiteLLM kwarg names handled + automatically. Supported at `connection_kwargs()`. LiteLLM makes the HTTP calls itself, so correlation is per client and `donkey.last_call` is not populated (see [Notes](#notes)). +- For `gemini()`: the shared client injected through + `HttpOptions.httpx_async_client`, round-trip verified live against a + `Format=Gemini` proxy. ## Install @@ -1189,23 +1201,147 @@ llm = LiteLlm( LiteLLM uses `api_base` and `extra_headers`, not `base_url` / `default_headers` — `connection_kwargs()` already translates for you. +## Native Gemini + +A proxy provisioned with **Format = Gemini** exposes the native Gemini API at +`POST //models/:generateContent` (and +`:streamGenerateContent`). ADK's own `Gemini` model speaks that API, so +`gemini()` returns a native `google.adk.models.Gemini` bound to the proxy, with +the SDK's shared HTTP client injected. + +The three forms mirror `model()`. Pass the bare Gemini model id — there is no +prefix, because the URL path carries the model: + +```python +from donkey_kit import Donkey +from google.adk.models import Gemini + +async with Donkey.from_env() as donkey: + # 1. Off a shared Donkey instance + llm = donkey.adk.gemini("gemini-2.5-flash") + + # 3. Governed kwargs, native constructor + llm = Gemini(model="gemini-2.5-flash", **donkey.adk.gemini_connection_kwargs()) +``` + +```python +# 2. Module-level factory +from donkey_kit.integrations.adk import gemini + +llm = gemini("gemini-2.5-flash") +``` + +**Pointing at the Gemini proxy.** `DONKEY_LLM_PROXY_URL` usually names a +`Format=OpenAI` proxy. If your Gemini proxy is a different one, pass its URL +as `base_url`; the `client_id`/`client_secret` pair must be contracted on that +proxy: + +```python +llm = donkey.adk.gemini("gemini-2.5-flash", base_url="https://…/ddk-gemini-inbound/") +``` + +Any other keyword is passed to ADK's `Gemini` and overrides the governed +default — for example `retry_options`, or your own `client_kwargs`. + +**Manual equivalent.** This is what `gemini_connection_kwargs()` returns: + +```python +from google.adk.models import Gemini + +llm = Gemini( + model="gemini-2.5-flash", + base_url=..., # the Format=Gemini proxy URL + client_kwargs={ + "api_key": ..., # a placeholder; google-genai requires one + "http_options": { + "base_url": ..., # same URL + "api_version": "", # the proxy path has no /v1beta segment + "headers": ..., # client_id / client_secret header pair + "timeout": ..., # milliseconds + "httpx_async_client": ..., # the SDK's shared client + }, + }, +) +``` + +`client_kwargs` replaces ADK's default HTTP options wholesale, so all five +`http_options` keys are passed together. `google-genai` requires an API key and +always sends it as `x-goog-api-key`; the gateway authenticates on the +`client_id`/`client_secret` pair and ignores it. + +**`donkey.last_call`.** Usage (`input_tokens`, `output_tokens`, +`total_tokens`, `cached_tokens`, `reasoning_tokens`) is read from Gemini's +`usageMetadata`, and `requested_model` from the URL path. Gemini's +`total_tokens` includes the thinking tokens it also reports as +`reasoning_tokens`. A `Format=Gemini` proxy is a passthrough, so it sends no +routing headers: `served_provider`, `served_model` and `routing_type` stay +`None`, and `substituted` is `False`. `request_id` and `api_instance_id` are +populated. + +`last_call` is contextvar-scoped, and ADK's `Runner` makes the model call in a +task of its own, so the caller of `runner.run_async(...)` reads `UNOBSERVED`. +Read it in an `after_model_callback`, which runs in the same task as the call: + +```python +from google.adk.agents import LlmAgent + +def record_usage(callback_context, llm_response): + r = donkey.last_call + print(r.requested_model, r.input_tokens, r.output_tokens) + return None # keep the model's response + +agent = LlmAgent( + name="assistant", + model=donkey.adk.gemini("gemini-2.5-flash"), + after_model_callback=record_usage, +) +``` + +When streaming, the callback runs once per partial response; usage lands on the +last one. + +**Errors.** `google-genai` raises its own `google.genai.errors.APIError` +(`ClientError` for a 4xx) and the SDK does not wrap it. The error's `.response` +is the proxy's HTTP response, so `classify()` gives you the typed Donkey error: + +```python +from donkey_kit.core.errors import classify +from google.genai import errors as genai_errors + +try: + async for event in runner.run_async(...): + ... +except genai_errors.APIError as exc: + err = classify(exc.response) # e.g. AuthError on 401, UpstreamRequestError on 404 +``` + +**One `Donkey` per event loop.** The shared HTTP client's connection pool is +bound to the event loop that first uses it. Build the `Donkey` — and the +models — inside the loop that runs them (`async with Donkey.from_env()` in your +`main`), and don't reuse the module-level `gemini()` across separate +`asyncio.run(...)` calls: the second run fails with +`RuntimeError: Event loop is closed`. + ## Notes -- **Correlation IDs are per-client, not per-run.** ADK sends requests through - its built-in LiteLLM model layer rather than the SDK's shared HTTP client, so - the correlation ID is set once per client instead of per `donkey.run()`. Every - governance header is still sent on every request. The conformance suite - checks this as a documented behaviour. -- **`donkey.last_call` is unavailable.** Because the response is handled by LiteLLM, - gateway identity, routing, and usage fields can't be observed. When every - adapter resolved on a `Donkey` is like this one, `donkey.last_call` reports - `status == LastCallStatus.UNAVAILABLE` and `available == False`, and names the - resolved adapters in `surface`. +- **Correlation IDs are per-client, not per-run (`model()` only).** `model()` + sends requests through ADK's built-in LiteLLM model layer rather than the + SDK's shared HTTP client, so the correlation ID is set once per client instead + of per `donkey.run()`. Every governance header is still sent on every request. + The conformance suite checks this as a documented behaviour. `gemini()` uses + the shared client, so its correlation ID is per run. +- **`donkey.last_call` is unavailable (`model()` only).** Because the response + is handled by LiteLLM, gateway identity, routing, and usage fields can't be + observed. When every adapter resolved on a `Donkey` is like this one, + `donkey.last_call` reports `status == LastCallStatus.UNAVAILABLE` and + `available == False`, and names the resolved adapters in `surface`. Once + `gemini()` has been called on a `Donkey`, the ADK adapter observes calls, so an + empty record reads `UNOBSERVED` instead. - `google-adk` requires `litellm>=1.84` as a floor, not a ceiling — pin your own upper bound if you need one. See the [error taxonomy](https://donkey-development-kit.github.io/donkey-development-kit/errors.md) for how proxy rejections surface through -ADK's `LiteLlm` model. +ADK's `LiteLlm` model, and [Native Gemini](#native-gemini) for `gemini()`. --- @@ -2887,15 +3023,19 @@ through every field. `donkey.last_call` is populated only when the governed response passes through the SDK's shared httpx client. Four `connection_kwargs()`-only adapters route -outside that response path: ADK sends requests through LiteLLM and CrewAI -through its native OpenAI provider, while LlamaIndex and Microsoft Agent -Framework receive only `default_headers`. +outside that response path: ADK's `model()` sends requests through LiteLLM and +CrewAI through its native OpenAI provider, while LlamaIndex and Microsoft Agent +Framework receive only `default_headers`. ADK's `gemini()` uses the shared +client, so once it has been called the ADK adapter observes calls (see +[Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini) for reading `last_call` inside an +ADK run). That static snapshot excludes the correlation ID bound later by `donkey.run(id=...)`, so those two adapters also do not propagate the run's correlation ID. When every adapter resolved on a `Donkey` is one of those four, a cold read -reports the limitation explicitly. For a `Donkey` that resolved only ADK: +reports the limitation explicitly. For a `Donkey` that resolved only ADK's +`model()`: ```python r = donkey.last_call @@ -8258,8 +8398,10 @@ except litellm.exceptions.APIError as err: should see:** `APIError 403` and the first line of the proxy's message — **not** `PIIDetected`. Without the policy it prints `NO REFUSAL`. - If you need typed refusals, `last_call` or run ids with ADK today, prefer a - framework path where the SDK owns the transport, such as + If you need typed refusals, `last_call` or run ids with ADK, use + `donkey.adk.gemini("gemini-2.5-flash")` on a `Format=Gemini` proxy — the SDK + owns that transport (see [Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini)). + Otherwise prefer a framework path where the SDK owns the transport, such as [OpenAI](https://donkey-development-kit.github.io/donkey-development-kit/examples/openai.md) or [LangGraph](https://donkey-development-kit.github.io/donkey-development-kit/examples/langgraph.md). A `404` here means the proxy's upstream has no `/chat/completions` route. @@ -8380,9 +8522,11 @@ Source: https://donkey-development-kit.github.io/donkey-development-kit/examples # Gemini -No Gemini adapter ships, so these are plain `httpx` against a proxy -provisioned **`Format=Gemini`** (for example `ddk-gemini-inbound`) with the -same `client_id` / `client_secret` pair. The route is +These scripts show the wire: plain `httpx` against a proxy provisioned +**`Format=Gemini`** (for example `ddk-gemini-inbound`) with the same +`client_id` / `client_secret` pair. For an agent, ADK's native `Gemini` model is +bound to the same proxy by `donkey.adk.gemini("gemini-2.5-flash")` — see +[Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini). The route is `/models/:generateContent`. `DonkeyConfig` still resolves and validates the credentials, and `classify()` still types the errors. @@ -8974,7 +9118,7 @@ even on a cold read. `donkey.last_call` is populated only when the governed response passes through the SDK's shared httpx client. Adapters that route outside that response path - (ADK and CrewAI, whose framework owns the transport, or the + (ADK's `model()` and CrewAI, whose framework owns the transport, or the `default_headers`-only adapters) report `UNAVAILABLE` with the surface named. See [When `last_call` is unavailable](https://donkey-development-kit.github.io/donkey-development-kit/telemetry.md#when-last_call-is-unavailable). diff --git a/website/public/llms.txt b/website/public/llms.txt index fc64a2a2..8bd143c6 100644 --- a/website/public/llms.txt +++ b/website/public/llms.txt @@ -12,7 +12,7 @@ - [Overview](https://donkey-development-kit.github.io/donkey-development-kit/frameworks.md): Governed model access from eight agent frameworks. Each adapter returns the framework's own native model or client object, pointed at your Omni Gateway LLM proxy. - [LangGraph](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/langgraph.md): Use a governed LangChain ChatOpenAI in LangGraph, pointed at your Omni Gateway LLM proxy with full header and transport injection, typed refusals, and conformance testing. -- [Google ADK](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md): Use a governed LiteLlm model in Google's Agent Development Kit (ADK), pointed at your Omni Gateway LLM proxy. +- [Google ADK](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md): Use a governed LiteLlm model, or ADK's native Gemini model on a Format=Gemini proxy, in Google's Agent Development Kit (ADK), pointed at your Omni Gateway LLM proxy. - [Strands](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/strands.md): Use a governed OpenAIModel in Strands Agents, pointed at your Omni Gateway LLM proxy with full header and transport injection. - [MS Agent Framework](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/agent-framework.md): Use a governed OpenAIChatClient in Microsoft Agent Framework, pointed at your Omni Gateway LLM proxy, with policy middleware that ends a run cleanly on a governance rejection. - [OpenAI Agents SDK](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/openai.md): Use a governed OpenAIChatCompletionsModel in the OpenAI Agents SDK, backed by a pre-built AsyncOpenAI client pointed at your Omni Gateway LLM proxy. diff --git a/website/public/reference/last-call.md b/website/public/reference/last-call.md index 81f1084d..96e96781 100644 --- a/website/public/reference/last-call.md +++ b/website/public/reference/last-call.md @@ -37,7 +37,7 @@ even on a cold read. `donkey.last_call` is populated only when the governed response passes through the SDK's shared httpx client. Adapters that route outside that response path - (ADK and CrewAI, whose framework owns the transport, or the + (ADK's `model()` and CrewAI, whose framework owns the transport, or the `default_headers`-only adapters) report `UNAVAILABLE` with the surface named. See [When `last_call` is unavailable](https://donkey-development-kit.github.io/donkey-development-kit/telemetry.md#when-last_call-is-unavailable). diff --git a/website/public/telemetry.md b/website/public/telemetry.md index 17397592..4f67c82d 100644 --- a/website/public/telemetry.md +++ b/website/public/telemetry.md @@ -353,15 +353,19 @@ through every field. `donkey.last_call` is populated only when the governed response passes through the SDK's shared httpx client. Four `connection_kwargs()`-only adapters route -outside that response path: ADK sends requests through LiteLLM and CrewAI -through its native OpenAI provider, while LlamaIndex and Microsoft Agent -Framework receive only `default_headers`. +outside that response path: ADK's `model()` sends requests through LiteLLM and +CrewAI through its native OpenAI provider, while LlamaIndex and Microsoft Agent +Framework receive only `default_headers`. ADK's `gemini()` uses the shared +client, so once it has been called the ADK adapter observes calls (see +[Native Gemini](https://donkey-development-kit.github.io/donkey-development-kit/frameworks/adk.md#native-gemini) for reading `last_call` inside an +ADK run). That static snapshot excludes the correlation ID bound later by `donkey.run(id=...)`, so those two adapters also do not propagate the run's correlation ID. When every adapter resolved on a `Donkey` is one of those four, a cold read -reports the limitation explicitly. For a `Donkey` that resolved only ADK: +reports the limitation explicitly. For a `Donkey` that resolved only ADK's +`model()`: ```python r = donkey.last_call