Proxy OpenAI requests: one request path, PII masking on the proxy - #5
Merged
Merged
Conversation
The OpenAI-compatible stream parser forwarded only the first tool call of
each chunk, so providers that pack parallel calls lost the rest (clients
saw "malformed tool call"). The handler also discarded output after the
first finish_reason, turned mid-stream provider errors into a clean
[DONE], and data: lines without a space were ignored.
Normalize every stream to the OpenAI shape: all calls forwarded with
contiguous indexes, one id and name per call (synthesized id if missing),
"{}" for calls without arguments, text kept alongside calls, a finish
reason always present, usage-only chunks no longer ending the stream,
and empty or error streams reported as provider errors so fallback runs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Requests and stream chunks were parsed into a narrow domain model and rebuilt, losing whatever it didn't cover: tool calls after the first in a chunk, output after the first finish chunk, images, reasoning from every provider but DeepSeek, and unknown request fields. For OpenAI-compatible providers (all of production) the client's body now goes upstream with a few documented edits (model name, usage for billing, token limit clamp, provider compatibility shims), and the provider's response comes back byte for byte except the model field. The gateway reads usage, finish reason, and errors from the passing stream for metering and fallback. The translating path remains only for PII masking with the filter on and for Anthropic, keeps its fidelity fixes, and no longer rewrites tool calls. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every request now takes the proxy path: the translating path (request rebuild, canonical stream events, per-chunk re-encoding) is removed. PII masking, the reason that path was kept, now works on the raw OpenAI JSON: the LLM-bound text of the request is masked in place, and placeholders in the response are restored as events pass, holding back text that may end in a split placeholder. Also removed: the unused Anthropic adapter, the response DTOs and mapper, the canonical stream/output/tool types, and control-plane leftovers in pkg/domain (promos, usage queries, catalog/pricing types). The request mapper now only validates and reads what routing, feature checks and metering need. Production code: 13.2k -> 9.8k lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One request path: PII masking on the proxy, remove the translating code
Verified against each provider's API on 2026-09-29: - Novita honors max_completion_tokens on every production model, so the gateway no longer renames it to max_tokens (DeepSeek still needs it: it silently ignores max_completion_tokens). - Novita and DeepSeek accept assistant tool-call messages without content, so content:null is no longer added. - No Novita model rejects parallel_tool_calls, store, service_tier, user, or tool_choice=auto, so the downgrade retry that dropped them is removed. The Novita same-body retries (generic 400 with trace id, 429 overload), the DeepSeek reasoning_content fill, and developer-as-system stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The removed per-stream diagnostics logged these; keep them on the one end-of-request log line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Novita repeats "role":"assistant" on every chunk for some models (MiniMax M2.7/M3 always, GLM 5.3-flash on tool calls, GLM 5.3 sometimes). OpenAI sends it once, and the OpenAI SDKs' stream helpers join repeated fields, so agents replayed the message with role "assistantassistant..." and the next request was rejected. Later repeats are dropped; other events pass unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The gateway proxies OpenAI requests instead of translating them
Clients kept hitting random errors ("malformed tool call", empty replies, lost reasoning). Streaming has needed about ten fixes of this kind since the gateway was split out. The cause was architectural: every request and every stream chunk was parsed into a narrow internal model and then rebuilt, so anything that model didn't cover was lost:
[DONE];data:lines without a space;Now, for OpenAI-compatible providers (Novita, DeepSeek, Tinfoil, Phala, MiniMax, DappNode, OpenAI), the client's request goes upstream as sent and the provider's response comes back byte for byte. The gateway only reads what it needs from the response: usage and finish reason for billing, and errors for fallback before the first byte.
Everything the gateway still does that a pure proxy wouldn't
Request body (
openai.PrepareProxyBody):modelbecomes the provider's model name.provider_options, a gateway-only field, is removed.stream_options.include_usage: trueso usage can be billed; otherstream_optionsare kept. Non-stream requests dropstream_options.max_tokens/max_completion_tokensis clamped to the model's maximum, and renamed tomax_tokensfor DeepSeek only, because DeepSeek silently ignoresmax_completion_tokens(with a limit of 40, it wrote 765 and 3,906 tokens).developerrole becomessystem, for every provider except OpenAI: DeepSeek returns 422 for it, Kimi K3 returns 400 on both Novita and Tinfoil, and GLM 5.3 on Novita ignores it.reasoning_contentgets"", which thinking mode requires (it returns 400 otherwise).Novita retries:
8. On a generic 400 with a trace id (on tool requests), or a 429
server_overload, the same body is retried (at most 2 and 1 times).Removed after testing each provider directly (2026-09-29):
max_tokensrename: Novita honorsmax_completion_tokenson every production model.content: nullon assistant tool-call messages: Novita and DeepSeek accept them withoutcontent.Checks before forwarding (the request is rejected, never changed):
9. Authentication, balance reservation, model lookup and router choice (
nexus/auto).10. Rejected fields and values:
nother than 1, both token-limit fields at once,functions/function_call/best_of, emptymessages, a tool message withouttool_call_id, unknown roles, and tools, parallel tool calls,response_formator streaming on models whose catalog entry doesn't support them.11. Bodies over 10 MB are refused, with a misleading "invalid JSON body" error.
Response:
roleis sent only in its first chunk. Novita repeats it on every chunk for some models, and the OpenAI SDKs' stream helpers join the repeats into"assistantassistant…", which then fails on the next turn.modelfield becomes the public model id. It's spliced in place; every other byte is unchanged.choices: []) are billed but not forwarded when the client didn't ask for usage, so the client gets the stream it asked for.data: [DONE]is appended when the provider finished (a finish reason arrived) but didn't send it.stream_interrupted) is appended instead of ending silently.: keepalivecomment is sent after 15 s of silence (long reasoning), so proxies don't time out.PII masking also runs on the proxy path (see the next section), so there is no longer a translating path.
Tests
model, packed tool calls, reasoning,data:without a space, comments); billing from the passing stream; fallback before first output; error after output (forwarded, billed as a failure); empty stream; client leaving before and after the finish reason; hidden usage chunks; non-stream bodies; PII/Anthropic routing;[DONE]and interrupted-stream framing.Part 2 (from #6): one request path, with PII masking on the proxy
Why
After #5, requests could still run through two code paths: the new proxy, and the old translating path, which rebuilds the request and every stream chunk from an internal model. The translating path was kept only for PII masking and the Anthropic adapter. PII masking is not rare: in production, 120 of 245 active keys use it (balanced 90, high 20, low 10). So half of the keys still went through the code that caused the stream bugs, and it made up about a third of the codebase.
What changes
One path for every request. Every request is authenticated, routed, validated, masked if its key asks for it, reserved, forwarded, and metered from the response. The provider's response is relayed byte for byte except for the
modelfield and, for masked keys, the restored text.PII masking now works on the raw OpenAI JSON (
services/pii.go):content, either a string or thetextof each content part;reasoning_content/reasoningsent back to the model;arguments, as JSON when valid, otherwise as plain text;userfield.content,reasoning_content/reasoningand tool arguments are restored as events pass.[DONE]if the provider never sends one.Removed:
Generate/StreamGenerate;pkg/domainfrom when the control plane lived here: promos, usage queries, catalog and pricing types, and unused error constructors. No other repo imports this package.The request mapper now only validates the body and reads what routing, feature checks and metering need:
The validation rules are unchanged, except that
response_format: {"type":"text"}no longer counts as structured output.Size: production Go code was 12.2k lines on
main, 13.2k with #5, and is 9.8k with this PR.Tests (production mapping, PII on)
Unit:
StreamandComplete;[DONE];End to end:
balancedPII key. It covered:minimax/minimax-m3JSON mode (below)Found, not changed: MiniMax's own API puts the model's reasoning inside
contentas<think>…</think>. That breaks JSON mode forminimax/minimax-m3. Its Novita fallback returns reasoning separately, so the response shape depends on which provider served it. The old gateway behaved the same way. The fix is a MiniMax request option (reasoning_split: true). That's a product choice, so it isn't in this PR.🤖 Generated with Claude Code