feat: add LiteLLM messages and agents integrations - #102
andrewklatzke wants to merge 6 commits into
Conversation
Add in-process messages and Agents SDK handlers so AI Configs can route across LiteLLM-supported providers with tools, streaming, and graph support.
Register catalog metadata, document the server SDK peer, emit native graph telemetry, and capture streamed content on the root span.
Start tracing after graph setup and always end the graph span across success, failures, result processing errors, and cancellation.
mypy cannot resolve a base class read directly off a dynamically imported module, so bind the name first as the OpenAI adapter does. Co-authored-by: Cursor <cursoragent@cursor.com>
…-lite-llm # Conflicts: # .release-please-manifest.json
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 3683c97. Configure here.
| model_factory=model_factory, capture_content=capture_content | ||
| ), | ||
| **kwargs, | ||
| ).invoke(user_input, context, variables) |
There was a problem hiding this comment.
Convenience wrappers drop conversation history
Medium Severity
litellm_agents and litellm_messages pop variables for invoke but leave history in kwargs for config. history is an invoke argument, so it is either rejected or never reaches the handler, and conversation history passed to the convenience APIs is lost.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 3683c97. Configure here.
| if usage_reported: | ||
| root.set_attribute("gen_ai.response.model", _model_name(config_value)) | ||
| set_usage_span_attributes(root, total) | ||
| end_span_once(root, ended, abandoned=True) |
There was a problem hiding this comment.
Stream cancel missing cancelled marker
Medium Severity
The messages stream finally treats every unfinished run as abandoned and never sets cancelled. CancelledError is not caught, so a cancelled stream is recorded as launchdarkly.stream.abandoned without launchdarkly.run.cancelled. The LiteLLM graph span has the same gap: cancel ends the span with no status and no cancelled marker.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 3683c97. Configure here.
jeffdupont
left a comment
There was a problem hiding this comment.
Reviewed with the 1.0 freeze in mind, alongside launchdarkly/js-ai-sdk#75 and the spec in launchdarkly/ai-sdks-monorepo#18. Tests pass at 3683c97 (uv run pytest: 1365 passed, 11 skipped, exit 0; make typecheck clean). The branch is 29 commits behind main (fee904a) and conflicts in .release-please-manifest.json and uv.lock. I resolved those by hand in a scratch tree and the merged result passes too (1445 passed, mypy clean). LiteLLM isn't in the GA plan, so the questions here are whether it's safe to ship and whether it ships as 1.0 or as experimental.
Four things I'd like settled before merge:
- An AI Config can send the customer's provider key to any host. Both handlers pass
model.parametersto LiteLLM, and LiteLLM takes connection settings as call arguments. I reproduced it at 3683c97: withANTHROPIC_API_KEYset in the process andmodel.parameters = {api_base: "http://127.0.0.1:<port>"}, bothcreate_litellm_messages_handler()andcreate_litellm_agents_handler()sent their request to my local listener withx-api-key: <the customer's key>. The agents model factory also readsbase_urlandapi_keyfrom the config on purpose (handler.py:117,native_graph.py:45). It's the same kind of problem as claude-agents on #107, and worse here: Python runs LiteLLM in-process, so no proxy sits in between. I'd use an allowlist per handler, as #107 does, and never forward connection or credential keys. Details inline. - The native graph emits the retired graph event. Since #114 (AIC-3211), every native adapter emits
$ld:ai:graph:nodeonce per node and no longer emits$ld:ai:graph:path.TESTING.md§2.2 onmainsays so.to_litellm_agentswas copied before that change. After a merge it will be the only adapter still sending$ld:ai:graph:path, and the only one sending no node events. The merge doesn't conflict there and the tests still pass, so nothing will catch it. - Python and JS don't match (js #75). The transport difference is by design and the spec explains it (A.13). These other differences aren't explained:
- Native graph telemetry: here you get a span and graph events. JS
toLiteLLMAgentsemits neither and takes nocontext. - Agent parameters: here, known snake_case keys go into
ModelSettingsand the rest intoextra_args. JS passes the raw bag asmodelSettings, wheremax_tokensandtop_pare silently ignored. So one config gets a token limit in Python and none in JS. - Framework tracing: Python leaves the Agents SDK's own tracing on. JS turns it off for the whole process.
- Factory name:
create_litellm_agents_handlerhere,createLiteLLMAgentHandlerin JS. The other agent packages use the same name in both languages. - Default
gen_ai.provider.namewhen the config has no provider:litellmhere,openaiin JS. That's a span attribute Monitoring reads.
- Native graph telemetry: here you get a span and graph events. JS
- I think this should ship as experimental. Both packages start at 0.1.0 with
bump-minor-pre-major, so they can stay 0.x when everything else goes to 1.0. The lifecycle spec (monorepo #33) only covers client features exported fromlaunchdarkly_ai_server.experimental, not whole provider packages, so it needs a line for this case. My suggestion: keep both LiteLLM packages out of the 1.0 cut, say "experimental" in their READMEs, add them to the promotion register, and promote them once the two languages match.
Smaller notes:
Runner.runis called withoutRunConfig(tracing_disabled=True). The Agents SDK's default exporter posts tohttps://api.openai.com/v1/traces/ingest(agents/tracing/processors.py:34). So ifOPENAI_API_KEYis also set, a run that LiteLLM routes to Anthropic still sends its trace, prompts included, to OpenAI. I took that from the SDK source and haven't watched it happen at runtime. Spec §2.x.2 asks for suppression. Per-runtracing_disabled(agents/run_config.py:257) does that without touching the customer's own Agents usage.- Messages streaming still records a cancelled stream as abandoned, with no
launchdarkly.run.cancelled(Bugbot's last finding). The agents stream catchesCancelledError; the messages stream doesn't. - Once #121 lands, native graph spans get
set_ld_span_attributes, andto_litellm_agentswill need the same change. - Content capture runs
str(message["content"])(litellm-messages/.../handler.py:362), so multimodal input shows up on the span as a Python list repr.
Bugbot's earlier findings (catalog entries, README install line, graph span leak, root content on stream) are fixed at this head. The history one in the convenience wrappers matches every other provider's wrapper, so it's not specific to this PR.
| model = config_value.get("model") or {} | ||
| raw_parameters = model.get("parameters") | ||
| parameters = dict(raw_parameters) if isinstance(raw_parameters, dict) else {} | ||
| for field in _OWNED_PARAMETERS: |
There was a problem hiding this comment.
This passes the whole model.parameters bag to litellm.acompletion, except the six handler-owned keys. LiteLLM reads connection settings from those same keyword arguments. I reproduced it at 3683c97: with ANTHROPIC_API_KEY set and parameters: {api_base: "http://127.0.0.1:<port>"}, the request went to the local listener with x-api-key set to the customer's key. Anyone who can edit an AI Config can collect the provider key that way.
I'd turn this into an allowlist of model settings (temperature, top_p, max_tokens, max_completion_tokens, stop, seed, presence_penalty, frequency_penalty, tool_choice, parallel_tool_calls, reasoning_effort, ...). That would keep out api_base, base_url, api_key, api_version, custom_llm_provider, extra_headers, headers, aws_*, vertex_* and the rest. A test that sends a config with api_base / api_key and checks that neither reaches completion would keep it that way. JS #75 strips api_key and base_url but not api_base.
| def _default_model_factory(name: str, parameters: dict[str, Any]) -> Any: | ||
| return LitellmModel( | ||
| model=name, | ||
| base_url=parameters.get("base_url"), |
There was a problem hiding this comment.
The default factory takes base_url and api_key straight from the AI Config. So a config chooses where the agent's requests go, and with which key. I reproduced the end-to-end run too: create_litellm_agents_handler() with api_base in the parameters sent the customer's ANTHROPIC_API_KEY to the local listener. native_graph.py:45 has the same factory.
Could these come from handler options or the environment only (for example create_litellm_agents_handler(base_url=..., api_key=...)), never from model.parameters? That's also what JS does: baseURL / apiKey are handler options there.
| key: value | ||
| for key, value in parameters.items() | ||
| if key | ||
| not in _MODEL_SETTING_FIELDS | {"max_turns", "base_url", "api_key", "model"} |
There was a problem hiding this comment.
Every key that isn't a known ModelSettings field goes into extra_args. LitellmModel adds extra_args to the litellm.acompletion call unchanged (agents/extensions/models/litellm_model.py:526). So even with the factory fixed, api_base, custom_llm_provider or cloud credentials in the config still reach LiteLLM. With {api_base: ...}, _model_settings(...).extra_args is {'api_base': ...}. Separately, extra_headers, extra_query and extra_body (lines 51-53) are allowlisted on purpose; those are the keys #107 excludes. I'd drop the catch-all extra_args and those three keys.
| client.track( | ||
| "$ld:ai:graph:total_tokens", ld_context, root_td, total_tokens | ||
| ) | ||
| client.track("$ld:ai:graph:path", ld_context, root_td, len(path)) |
There was a problem hiding this comment.
$ld:ai:graph:path was retired by #114 (AIC-3211). The other native adapters on main emit $ld:ai:graph:node from on_agent_start (line 202 here), with nodeKey and a 0-based index, and they no longer track path. TESTING.md §2.2 now requires that, and the openai-agents tests assert $ld:ai:graph:path is not tracked. The merge won't conflict here, so this file would quietly keep the old event contract. Copying the current openai-agents/native_graph.py hooks over would fix it.
| system_instructions=instructions, | ||
| messages=[text_message("user", str(prompt))], | ||
| ) | ||
| result = await importlib.import_module("agents").Runner.run( |
There was a problem hiding this comment.
Nothing turns off the Agents SDK's own tracing here, and its default exporter sends traces to OpenAI (agents/tracing/processors.py:34). With LiteLLM routing to another provider, that's a second, unexpected destination for prompts whenever OPENAI_API_KEY is set. I haven't checked this at runtime. Spec §2.x.2 asks for it to be suppressed. run_config=RunConfig(tracing_disabled=True) on this call and on run_streamed keeps it per run. JS uses the process-wide setTracingDisabled(true), which I've commented on in #75.


Summary
launchdarkly-ai-litellm-messagesusing in-process LiteLLM completionslaunchdarkly-ai-litellm-agentsusing the OpenAI Agents SDK LiteLLM modelUsage
Python runs LiteLLM in-process, so callers set the upstream provider credentials required by the evaluated model; no LiteLLM proxy is required.
Test plan
uv run pytestuv run mypyacross all package sourcesuv run ruff check .uv run ruff format --check .Made with Cursor