One request path: PII masking on the proxy, remove the translating code - #6
Merged
Merged
Conversation
Every request now takes the proxy path: the translating path (request rebuild, canonical stream events, per-chunk re-encoding) is removed. PII masking, the reason that path was kept, now works on the raw OpenAI JSON: the LLM-bound text of the request is masked in place, and placeholders in the response are restored as events pass, holding back text that may end in a split placeholder. Also removed: the unused Anthropic adapter, the response DTOs and mapper, the canonical stream/output/tool types, and control-plane leftovers in pkg/domain (promos, usage queries, catalog/pricing types). The request mapper now only validates and reads what routing, feature checks and metering need. Production code: 13.2k -> 9.8k lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #5. It removes the second request path that #5 left in place.
Why
After #5, requests could still run through two code paths: the new proxy, and the old translating path, which rebuilds the request and every stream chunk from an internal model. The translating path was kept only for PII masking and the Anthropic adapter. PII masking is not rare: in production, 120 of 245 active keys use it (balanced 90, high 20, low 10). So half of the keys still went through the code that caused the stream bugs, and it made up about a third of the codebase.
What changes
One path for every request. Every request is authenticated, routed, validated, masked if its key asks for it, reserved, forwarded, and metered from the response. The provider's response is relayed byte for byte except for the
modelfield and, for masked keys, the restored text.PII masking now works on the raw OpenAI JSON (
services/pii.go):content, either a string or thetextof each content part;reasoning_content/reasoningsent back to the model;arguments, as JSON when valid, otherwise as plain text;userfield.content,reasoning_content/reasoningand tool arguments are restored as events pass.[DONE]if the provider never sends one.Removed:
Generate/StreamGenerate;pkg/domainfrom when the control plane lived here: promos, usage queries, catalog and pricing types, and unused error constructors. No other repo imports this package.The request mapper now only validates the body and reads what routing, feature checks and metering need:
The validation rules are unchanged, except that
response_format: {"type":"text"}no longer counts as structured output.Size: production Go code was 12.2k lines on
main, 13.2k with #5, and is 9.8k with this PR.Tests
Unit:
StreamandComplete;[DONE];End to end:
balancedPII key. It covered:minimax/minimax-m3JSON mode (below)Found, not changed: MiniMax's own API puts the model's reasoning inside
contentas<think>…</think>. That breaks JSON mode forminimax/minimax-m3. Its Novita fallback returns reasoning separately, so the response shape depends on which provider served it. The old gateway behaved the same way. The fix is a MiniMax request option (reasoning_split: true). That's a product choice, so it isn't in this PR.🤖 Generated with Claude Code