Skip to content

Proxy OpenAI requests: one request path, PII masking on the proxy - #5

Merged
Marketen merged 8 commits into
mainfrom
fix/stream-hardening
Sep 29, 2026
Merged

Marketen merged 8 commits into
mainfrom
fix/stream-hardening

Conversation

@Marketen

@Marketen Marketen commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

The gateway proxies OpenAI requests instead of translating them

Clients kept hitting random errors ("malformed tool call", empty replies, lost reasoning). Streaming has needed about ten fixes of this kind since the gateway was split out. The cause was architectural: every request and every stream chunk was parsed into a narrow internal model and then rebuilt, so anything that model didn't cover was lost:

  • only the first tool call of each chunk, so parallel tool calls broke;
  • anything after the first finish chunk;
  • a mid-stream provider error, which ended as a clean [DONE];
  • data: lines without a space;
  • images in requests;
  • reasoning from every provider except DeepSeek;
  • every request field it didn't model.

Now, for OpenAI-compatible providers (Novita, DeepSeek, Tinfoil, Phala, MiniMax, DappNode, OpenAI), the client's request goes upstream as sent and the provider's response comes back byte for byte. The gateway only reads what it needs from the response: usage and finish reason for billing, and errors for fallback before the first byte.

Everything the gateway still does that a pure proxy wouldn't

Request body (openai.PrepareProxyBody):

  1. model becomes the provider's model name.
  2. provider_options, a gateway-only field, is removed.
  3. Streams get stream_options.include_usage: true so usage can be billed; other stream_options are kept. Non-stream requests drop stream_options.
  4. max_tokens / max_completion_tokens is clamped to the model's maximum, and renamed to max_tokens for DeepSeek only, because DeepSeek silently ignores max_completion_tokens (with a limit of 40, it wrote 765 and 3,906 tokens).
  5. The developer role becomes system, for every provider except OpenAI: DeepSeek returns 422 for it, Kimi K3 returns 400 on both Novita and Tinfoil, and GLM 5.3 on Novita ignores it.
  6. DeepSeek: an assistant tool-call message without reasoning_content gets "", which thinking mode requires (it returns 400 otherwise).
  7. The JSON is re-encoded: key order and whitespace may differ, values don't, and numbers are exact.

Novita retries:
8. On a generic 400 with a trace id (on tool requests), or a 429 server_overload, the same body is retried (at most 2 and 1 times).

Removed after testing each provider directly (2026-09-29):

  • Novita's max_tokens rename: Novita honors max_completion_tokens on every production model.
  • content: null on assistant tool-call messages: Novita and DeepSeek accept them without content.
  • Novita's downgrade retry: no Novita model rejects the fields it dropped.

Checks before forwarding (the request is rejected, never changed):
9. Authentication, balance reservation, model lookup and router choice (nexus/auto).
10. Rejected fields and values: n other than 1, both token-limit fields at once, functions/function_call/best_of, empty messages, a tool message without tool_call_id, unknown roles, and tools, parallel tool calls, response_format or streaming on models whose catalog entry doesn't support them.
11. Bodies over 10 MB are refused, with a misleading "invalid JSON body" error.

Response:

  • A choice's role is sent only in its first chunk. Novita repeats it on every chunk for some models, and the OpenAI SDKs' stream helpers join the repeats into "assistantassistant…", which then fails on the next turn.
  1. The top-level model field becomes the public model id. It's spliced in place; every other byte is unchanged.
  2. Usage-only chunks (choices: []) are billed but not forwarded when the client didn't ask for usage, so the client gets the stream it asked for.
  3. data: [DONE] is appended when the provider finished (a finish reason arrived) but didn't send it.
  4. If the upstream connection breaks mid-response without a finish reason, an error event (stream_interrupted) is appended instead of ending silently.
  5. A : keepalive comment is sent after 15 s of silence (long reasoning), so proxies don't time out.
  6. If the provider fails before its first output (HTTP error, error event, empty stream), the configured fallback provider is used; the client never sees the failed attempt.
  7. Provider HTTP errors become gateway error JSON: 401/403, 400 and other 5xx become 502, 429 stays 429, 503 stays 503. A non-stream 200 whose body is an error becomes a 502.
  8. The provider's response headers aren't forwarded; the gateway sets its own.

PII masking also runs on the proxy path (see the next section), so there is no longer a translating path.

Tests

  • Unit: request preparation (images, unknown fields, exact numbers, each edit above); relay fidelity (byte-exact apart from model, packed tool calls, reasoning, data: without a space, comments); billing from the passing stream; fallback before first output; error after output (forwarded, billed as a failure); empty stream; client leaving before and after the finish reason; hidden usage chunks; non-stream bodies; PII/Anthropic routing; [DONE] and interrupted-stream framing.
  • End to end: a local gateway and metering built from source, with the TEE gateway's providers (Novita, DeepSeek, Tinfoil) and their production keys. The Docker gateway's other providers (Phala, MiniMax, DappNode, OpenAI, Anthropic) and PII-masked keys were not tested end to end here. The official OpenAI Python SDK ran plain chat, streaming (with and without usage), parallel tool calls (streamed and non-streamed), a tool-result round trip, and JSON output where the catalog allows it. All 15 production models passed every supported test (Novita: GLM 5.2/5.3/5.3-flash, Kimi K2.6/K2.7-code/K3, MiniMax M2.7/M3, Qwen 3.8 27B; DeepSeek: v4.1-flash, v4-pro; Tinfoil: GLM 5.3/5.3-flash, Kimi K3, Gemma 4 31B). Every request was billed with usage, and every Tinfoil request stored a verified transport proof.

Part 2 (from #6): one request path, with PII masking on the proxy

Why

After #5, requests could still run through two code paths: the new proxy, and the old translating path, which rebuilds the request and every stream chunk from an internal model. The translating path was kept only for PII masking and the Anthropic adapter. PII masking is not rare: in production, 120 of 245 active keys use it (balanced 90, high 20, low 10). So half of the keys still went through the code that caused the stream bugs, and it made up about a third of the codebase.

What changes

One path for every request. Every request is authenticated, routed, validated, masked if its key asks for it, reserved, forwarded, and metered from the response. The provider's response is relayed byte for byte except for the model field and, for masked keys, the restored text.

PII masking now works on the raw OpenAI JSON (services/pii.go):

  • Request: the text the model receives is masked in place:
    • message content, either a string or the text of each content part;
    • tool results, as JSON, masking only string values;
    • reasoning_content/reasoning sent back to the model;
    • tool-call arguments, as JSON when valid, otherwise as plain text;
    • the user field.
  • Not changed: images, tool definitions and every other field stay as sent. Before, multi-part content was flattened into one string, which dropped images for masked keys.
  • Responses:
    • content, reasoning_content/reasoning and tool arguments are restored as events pass.
    • Text that may end in a placeholder split across events is held back. It's released with that choice's finish reason, or in a final event before [DONE] if the provider never sends one.
    • Events with no generated text pass through byte for byte.
    • Non-stream bodies are restored the same way.
  • Kept as they were:
    • fail-closed and fail-open handling;
    • the per-key modes;
    • bracketless placeholder aliases;
    • logging of placeholders the model altered;
    • removal of known originals from error logs.

Removed:

  • The translating path:
    • Generate/StreamGenerate;
    • the stream parser and its diagnostics;
    • the request rebuild and the response/chunk rebuild;
    • the usage-tracking, buffered and PII stream wrappers.
  • The Anthropic adapter: no active model uses it.
  • Unused types: the internal stream/output/tool-call types and the response DTOs.
  • Leftovers in pkg/domain from when the control plane lived here: promos, usage queries, catalog and pricing types, and unused error constructors. No other repo imports this package.

The request mapper now only validates the body and reads what routing, feature checks and metering need:

  • the model, the stream flag and the output limit;
  • the features the request uses;
  • each message's role and text, and the tool names (the router's contract).

The validation rules are unchanged, except that response_format: {"type":"text"} no longer counts as structured output.

Size: production Go code was 12.2k lines on main, 13.2k with #5, and is 9.8k with this PR.

Tests (production mapping, PII on)

Unit:

  • request mapper validation and summary;
  • the proxy relay (unchanged from Proxy OpenAI requests: one request path, PII masking on the proxy #5);
  • the Novita retry policy, ported to the proxy methods;
  • Tinfoil proof evidence on Stream and Complete;
  • service routing, balance reservation, fallback and cancellation.
  • PII masking:
    • every masked field, with images, tool definitions and exact numbers left untouched;
    • invalid JSON arguments masked as plain text;
    • key off, filter off, nothing found, fail closed and fail open;
    • placeholders split across content, reasoning and tool-argument events, and bracketless aliases;
    • held text released at the finish reason or before [DONE];
    • byte-exact events without text;
    • unresolved-placeholder logging and error sanitizing.

End to end:

  • Setup:
    • a local gateway and metering built from source;
    • the local model catalog set to production's mapping: DappNode, DeepSeek, MiniMax with its Novita fallback, Novita and Tinfoil;
    • the production provider keys;
    • the real Presidio analyzer (small spaCy model) with the PII filter on, as in the Docker gateway.
  • The suite ran the official OpenAI Python SDK on every model twice: once with a plain key and once with a balanced PII key. It covered:
    • chat;
    • streaming, with and without usage;
    • parallel tool calls, streamed and non-streamed;
    • a tool-result round trip;
    • JSON output where the catalog allows it.
  • PII key only:
    • an echo test (the client must get its real email back);
    • a tool call whose arguments must contain the real email.
  • Results:
Plain key PII key (balanced)
All 15 production models every test passes except one every test passes except one
The exception minimax/minimax-m3 JSON mode (below) same
  • Masking and billing:
    • masking reached the provider (the model saw a placeholder, and the client got the email back);
    • metering recorded usage for all 262 requests, with no failures;
    • all 64 Tinfoil requests stored verified transport proofs.

Found, not changed: MiniMax's own API puts the model's reasoning inside content as <think>…</think>. That breaks JSON mode for minimax/minimax-m3. Its Novita fallback returns reasoning separately, so the response shape depends on which provider served it. The old gateway behaved the same way. The fix is a MiniMax request option (reasoning_split: true). That's a product choice, so it isn't in this PR.

🤖 Generated with Claude Code

Marketen and others added 2 commits September 28, 2026 23:04
The OpenAI-compatible stream parser forwarded only the first tool call of
each chunk, so providers that pack parallel calls lost the rest (clients
saw "malformed tool call"). The handler also discarded output after the
first finish_reason, turned mid-stream provider errors into a clean
[DONE], and data: lines without a space were ignored.

Normalize every stream to the OpenAI shape: all calls forwarded with
contiguous indexes, one id and name per call (synthesized id if missing),
"{}" for calls without arguments, text kept alongside calls, a finish
reason always present, usage-only chunks no longer ending the stream,
and empty or error streams reported as provider errors so fallback runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Requests and stream chunks were parsed into a narrow domain model and
rebuilt, losing whatever it didn't cover: tool calls after the first in a
chunk, output after the first finish chunk, images, reasoning from every
provider but DeepSeek, and unknown request fields.

For OpenAI-compatible providers (all of production) the client's body now
goes upstream with a few documented edits (model name, usage for billing,
token limit clamp, provider compatibility shims), and the provider's
response comes back byte for byte except the model field. The gateway reads
usage, finish reason, and errors from the passing stream for metering and
fallback. The translating path remains only for PII masking with the filter
on and for Anthropic, keeps its fidelity fixes, and no longer rewrites tool
calls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Marketen Marketen changed the title fix: never drop or corrupt streamed tool calls Proxy OpenAI requests instead of translating them Sep 29, 2026
Comment thread apps/gateway/internal/application/services/proxy.go Fixed
Marketen and others added 2 commits September 29, 2026 10:35
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every request now takes the proxy path: the translating path (request
rebuild, canonical stream events, per-chunk re-encoding) is removed. PII
masking, the reason that path was kept, now works on the raw OpenAI JSON:
the LLM-bound text of the request is masked in place, and placeholders in
the response are restored as events pass, holding back text that may end
in a split placeholder.

Also removed: the unused Anthropic adapter, the response DTOs and mapper,
the canonical stream/output/tool types, and control-plane leftovers in
pkg/domain (promos, usage queries, catalog/pricing types). The request
mapper now only validates and reads what routing, feature checks and
metering need. Production code: 13.2k -> 9.8k lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One request path: PII masking on the proxy, remove the translating code
@Marketen Marketen changed the title Proxy OpenAI requests instead of translating them Proxy OpenAI requests: one request path, PII masking on the proxy Sep 29, 2026
Marketen and others added 3 commits September 29, 2026 11:28
Verified against each provider's API on 2026-09-29:
- Novita honors max_completion_tokens on every production model, so the
  gateway no longer renames it to max_tokens (DeepSeek still needs it: it
  silently ignores max_completion_tokens).
- Novita and DeepSeek accept assistant tool-call messages without content,
  so content:null is no longer added.
- No Novita model rejects parallel_tool_calls, store, service_tier, user,
  or tool_choice=auto, so the downgrade retry that dropped them is removed.
The Novita same-body retries (generic 400 with trace id, 429 overload),
the DeepSeek reasoning_content fill, and developer-as-system stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The removed per-stream diagnostics logged these; keep them on the one
end-of-request log line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Novita repeats "role":"assistant" on every chunk for some models (MiniMax
M2.7/M3 always, GLM 5.3-flash on tool calls, GLM 5.3 sometimes). OpenAI
sends it once, and the OpenAI SDKs' stream helpers join repeated fields,
so agents replayed the message with role "assistantassistant..." and the
next request was rejected. Later repeats are dropped; other events pass
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Marketen
Marketen merged commit ede2192 into main Sep 29, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants