Skip to content

One request path: PII masking on the proxy, remove the translating code - #6

Merged
Marketen merged 1 commit into
fix/stream-hardeningfrom
refactor/one-path
Sep 29, 2026
Merged

Marketen merged 1 commit into
fix/stream-hardeningfrom
refactor/one-path

Conversation

@Marketen

Copy link
Copy Markdown
Contributor

Stacked on #5. It removes the second request path that #5 left in place.

Why

After #5, requests could still run through two code paths: the new proxy, and the old translating path, which rebuilds the request and every stream chunk from an internal model. The translating path was kept only for PII masking and the Anthropic adapter. PII masking is not rare: in production, 120 of 245 active keys use it (balanced 90, high 20, low 10). So half of the keys still went through the code that caused the stream bugs, and it made up about a third of the codebase.

What changes

One path for every request. Every request is authenticated, routed, validated, masked if its key asks for it, reserved, forwarded, and metered from the response. The provider's response is relayed byte for byte except for the model field and, for masked keys, the restored text.

PII masking now works on the raw OpenAI JSON (services/pii.go):

  • Request: the text the model receives is masked in place:
    • message content, either a string or the text of each content part;
    • tool results, as JSON, masking only string values;
    • reasoning_content/reasoning sent back to the model;
    • tool-call arguments, as JSON when valid, otherwise as plain text;
    • the user field.
  • Not changed: images, tool definitions and every other field stay as sent. Before, multi-part content was flattened into one string, which dropped images for masked keys.
  • Responses:
    • content, reasoning_content/reasoning and tool arguments are restored as events pass.
    • Text that may end in a placeholder split across events is held back. It's released with that choice's finish reason, or in a final event before [DONE] if the provider never sends one.
    • Events with no generated text pass through byte for byte.
    • Non-stream bodies are restored the same way.
  • Kept as they were:
    • fail-closed and fail-open handling;
    • the per-key modes;
    • bracketless placeholder aliases;
    • logging of placeholders the model altered;
    • removal of known originals from error logs.

Removed:

  • The translating path:
    • Generate/StreamGenerate;
    • the stream parser and its diagnostics;
    • the request rebuild and the response/chunk rebuild;
    • the usage-tracking, buffered and PII stream wrappers.
  • The Anthropic adapter: no active model uses it.
  • Unused types: the internal stream/output/tool-call types and the response DTOs.
  • Leftovers in pkg/domain from when the control plane lived here: promos, usage queries, catalog and pricing types, and unused error constructors. No other repo imports this package.

The request mapper now only validates the body and reads what routing, feature checks and metering need:

  • the model, the stream flag and the output limit;
  • the features the request uses;
  • each message's role and text, and the tool names (the router's contract).

The validation rules are unchanged, except that response_format: {"type":"text"} no longer counts as structured output.

Size: production Go code was 12.2k lines on main, 13.2k with #5, and is 9.8k with this PR.

Tests

Unit:

  • request mapper validation and summary;
  • the proxy relay (unchanged from Proxy OpenAI requests: one request path, PII masking on the proxy #5);
  • the Novita retry policy, ported to the proxy methods;
  • Tinfoil proof evidence on Stream and Complete;
  • service routing, balance reservation, fallback and cancellation.
  • PII masking:
    • every masked field, with images, tool definitions and exact numbers left untouched;
    • invalid JSON arguments masked as plain text;
    • key off, filter off, nothing found, fail closed and fail open;
    • placeholders split across content, reasoning and tool-argument events, and bracketless aliases;
    • held text released at the finish reason or before [DONE];
    • byte-exact events without text;
    • unresolved-placeholder logging and error sanitizing.

End to end:

  • Setup:
    • a local gateway and metering built from source;
    • the local model catalog set to production's mapping: DappNode, DeepSeek, MiniMax with its Novita fallback, Novita and Tinfoil;
    • the production provider keys;
    • the real Presidio analyzer (small spaCy model) with the PII filter on, as in the Docker gateway.
  • The suite ran the official OpenAI Python SDK on every model twice: once with a plain key and once with a balanced PII key. It covered:
    • chat;
    • streaming, with and without usage;
    • parallel tool calls, streamed and non-streamed;
    • a tool-result round trip;
    • JSON output where the catalog allows it.
  • PII key only:
    • an echo test (the client must get its real email back);
    • a tool call whose arguments must contain the real email.
  • Results:
Plain key PII key (balanced)
All 15 production models every test passes except one every test passes except one
The exception minimax/minimax-m3 JSON mode (below) same
  • Masking and billing:
    • masking reached the provider (the model saw a placeholder, and the client got the email back);
    • metering recorded usage for all 262 requests, with no failures;
    • all 64 Tinfoil requests stored verified transport proofs.

Found, not changed: MiniMax's own API puts the model's reasoning inside content as <think>…</think>. That breaks JSON mode for minimax/minimax-m3. Its Novita fallback returns reasoning separately, so the response shape depends on which provider served it. The old gateway behaved the same way. The fix is a MiniMax request option (reasoning_split: true). That's a product choice, so it isn't in this PR.

🤖 Generated with Claude Code

Every request now takes the proxy path: the translating path (request
rebuild, canonical stream events, per-chunk re-encoding) is removed. PII
masking, the reason that path was kept, now works on the raw OpenAI JSON:
the LLM-bound text of the request is masked in place, and placeholders in
the response are restored as events pass, holding back text that may end
in a split placeholder.

Also removed: the unused Anthropic adapter, the response DTOs and mapper,
the canonical stream/output/tool types, and control-plane leftovers in
pkg/domain (promos, usage queries, catalog/pricing types). The request
mapper now only validates and reads what routing, feature checks and
metering need. Production code: 13.2k -> 9.8k lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant