Skip to content

feat: add LiteLLM messages and agents integrations - #102

Open
andrewklatzke wants to merge 6 commits into
mainfrom
aklatzke/AIC-3410/add-lite-llm
Open

andrewklatzke wants to merge 6 commits into
mainfrom
aklatzke/AIC-3410/add-lite-llm

Conversation

@andrewklatzke

Copy link
Copy Markdown
Contributor

Summary

  • add launchdarkly-ai-litellm-messages using in-process LiteLLM completions
  • add launchdarkly-ai-litellm-agents using the OpenAI Agents SDK LiteLLM model
  • support tools, streaming, structured output, conversation history, native graphs, telemetry, and release configuration
  • document installation, provider credentials, examples, and wildcard-handler behavior

Usage

from launchdarkly_ai_litellm_messages import litellm_messages

result = await litellm_messages(
    "launch-darkly-documentation-summarizer-messages",
    "How do AI Configs work?",
    {"kind": "user", "key": "user-123"},
)
print(result["response"])

Python runs LiteLLM in-process, so callers set the upstream provider credentials required by the evaluated model; no LiteLLM proxy is required.

Test plan

  • uv run pytest
  • uv run mypy across all package sources
  • uv run ruff check .
  • uv run ruff format --check .
  • LiteLLM messages and agents integration scenarios

Made with Cursor

Add in-process messages and Agents SDK handlers so AI Configs can route across LiteLLM-supported providers with tools, streaming, and graph support.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/litellm-agents/README.md
Register catalog metadata, document the server SDK peer, emit native graph telemetry, and capture streamed content on the root span.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Start tracing after graph setup and always end the graph span across success, failures, result processing errors, and cancellation.
andrewklatzke and others added 3 commits September 21, 2026 11:59
mypy cannot resolve a base class read directly off a dynamically imported module, so bind the name first as the OpenAI adapter does.

Co-authored-by: Cursor <cursoragent@cursor.com>
…-lite-llm

# Conflicts:
#	.release-please-manifest.json

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 3683c97. Configure here.

model_factory=model_factory, capture_content=capture_content
),
**kwargs,
).invoke(user_input, context, variables)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Convenience wrappers drop conversation history

Medium Severity

litellm_agents and litellm_messages pop variables for invoke but leave history in kwargs for config. history is an invoke argument, so it is either rejected or never reaches the handler, and conversation history passed to the convenience APIs is lost.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 3683c97. Configure here.

if usage_reported:
root.set_attribute("gen_ai.response.model", _model_name(config_value))
set_usage_span_attributes(root, total)
end_span_once(root, ended, abandoned=True)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stream cancel missing cancelled marker

Medium Severity

The messages stream finally treats every unfinished run as abandoned and never sets cancelled. CancelledError is not caught, so a cancelled stream is recorded as launchdarkly.stream.abandoned without launchdarkly.run.cancelled. The LiteLLM graph span has the same gap: cancel ends the span with no status and no cancelled marker.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 3683c97. Configure here.

@jeffdupont jeffdupont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed with the 1.0 freeze in mind, alongside launchdarkly/js-ai-sdk#75 and the spec in launchdarkly/ai-sdks-monorepo#18. Tests pass at 3683c97 (uv run pytest: 1365 passed, 11 skipped, exit 0; make typecheck clean). The branch is 29 commits behind main (fee904a) and conflicts in .release-please-manifest.json and uv.lock. I resolved those by hand in a scratch tree and the merged result passes too (1445 passed, mypy clean). LiteLLM isn't in the GA plan, so the questions here are whether it's safe to ship and whether it ships as 1.0 or as experimental.

Four things I'd like settled before merge:

  1. An AI Config can send the customer's provider key to any host. Both handlers pass model.parameters to LiteLLM, and LiteLLM takes connection settings as call arguments. I reproduced it at 3683c97: with ANTHROPIC_API_KEY set in the process and model.parameters = {api_base: "http://127.0.0.1:<port>"}, both create_litellm_messages_handler() and create_litellm_agents_handler() sent their request to my local listener with x-api-key: <the customer's key>. The agents model factory also reads base_url and api_key from the config on purpose (handler.py:117, native_graph.py:45). It's the same kind of problem as claude-agents on #107, and worse here: Python runs LiteLLM in-process, so no proxy sits in between. I'd use an allowlist per handler, as #107 does, and never forward connection or credential keys. Details inline.
  2. The native graph emits the retired graph event. Since #114 (AIC-3211), every native adapter emits $ld:ai:graph:node once per node and no longer emits $ld:ai:graph:path. TESTING.md §2.2 on main says so. to_litellm_agents was copied before that change. After a merge it will be the only adapter still sending $ld:ai:graph:path, and the only one sending no node events. The merge doesn't conflict there and the tests still pass, so nothing will catch it.
  3. Python and JS don't match (js #75). The transport difference is by design and the spec explains it (A.13). These other differences aren't explained:
    • Native graph telemetry: here you get a span and graph events. JS toLiteLLMAgents emits neither and takes no context.
    • Agent parameters: here, known snake_case keys go into ModelSettings and the rest into extra_args. JS passes the raw bag as modelSettings, where max_tokens and top_p are silently ignored. So one config gets a token limit in Python and none in JS.
    • Framework tracing: Python leaves the Agents SDK's own tracing on. JS turns it off for the whole process.
    • Factory name: create_litellm_agents_handler here, createLiteLLMAgentHandler in JS. The other agent packages use the same name in both languages.
    • Default gen_ai.provider.name when the config has no provider: litellm here, openai in JS. That's a span attribute Monitoring reads.
  4. I think this should ship as experimental. Both packages start at 0.1.0 with bump-minor-pre-major, so they can stay 0.x when everything else goes to 1.0. The lifecycle spec (monorepo #33) only covers client features exported from launchdarkly_ai_server.experimental, not whole provider packages, so it needs a line for this case. My suggestion: keep both LiteLLM packages out of the 1.0 cut, say "experimental" in their READMEs, add them to the promotion register, and promote them once the two languages match.

Smaller notes:

  • Runner.run is called without RunConfig(tracing_disabled=True). The Agents SDK's default exporter posts to https://api.openai.com/v1/traces/ingest (agents/tracing/processors.py:34). So if OPENAI_API_KEY is also set, a run that LiteLLM routes to Anthropic still sends its trace, prompts included, to OpenAI. I took that from the SDK source and haven't watched it happen at runtime. Spec §2.x.2 asks for suppression. Per-run tracing_disabled (agents/run_config.py:257) does that without touching the customer's own Agents usage.
  • Messages streaming still records a cancelled stream as abandoned, with no launchdarkly.run.cancelled (Bugbot's last finding). The agents stream catches CancelledError; the messages stream doesn't.
  • Once #121 lands, native graph spans get set_ld_span_attributes, and to_litellm_agents will need the same change.
  • Content capture runs str(message["content"]) (litellm-messages/.../handler.py:362), so multimodal input shows up on the span as a Python list repr.

Bugbot's earlier findings (catalog entries, README install line, graph span leak, root content on stream) are fixed at this head. The history one in the convenience wrappers matches every other provider's wrapper, so it's not specific to this PR.

model = config_value.get("model") or {}
raw_parameters = model.get("parameters")
parameters = dict(raw_parameters) if isinstance(raw_parameters, dict) else {}
for field in _OWNED_PARAMETERS:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This passes the whole model.parameters bag to litellm.acompletion, except the six handler-owned keys. LiteLLM reads connection settings from those same keyword arguments. I reproduced it at 3683c97: with ANTHROPIC_API_KEY set and parameters: {api_base: "http://127.0.0.1:<port>"}, the request went to the local listener with x-api-key set to the customer's key. Anyone who can edit an AI Config can collect the provider key that way.

I'd turn this into an allowlist of model settings (temperature, top_p, max_tokens, max_completion_tokens, stop, seed, presence_penalty, frequency_penalty, tool_choice, parallel_tool_calls, reasoning_effort, ...). That would keep out api_base, base_url, api_key, api_version, custom_llm_provider, extra_headers, headers, aws_*, vertex_* and the rest. A test that sends a config with api_base / api_key and checks that neither reaches completion would keep it that way. JS #75 strips api_key and base_url but not api_base.

def _default_model_factory(name: str, parameters: dict[str, Any]) -> Any:
return LitellmModel(
model=name,
base_url=parameters.get("base_url"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The default factory takes base_url and api_key straight from the AI Config. So a config chooses where the agent's requests go, and with which key. I reproduced the end-to-end run too: create_litellm_agents_handler() with api_base in the parameters sent the customer's ANTHROPIC_API_KEY to the local listener. native_graph.py:45 has the same factory.

Could these come from handler options or the environment only (for example create_litellm_agents_handler(base_url=..., api_key=...)), never from model.parameters? That's also what JS does: baseURL / apiKey are handler options there.

key: value
for key, value in parameters.items()
if key
not in _MODEL_SETTING_FIELDS | {"max_turns", "base_url", "api_key", "model"}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every key that isn't a known ModelSettings field goes into extra_args. LitellmModel adds extra_args to the litellm.acompletion call unchanged (agents/extensions/models/litellm_model.py:526). So even with the factory fixed, api_base, custom_llm_provider or cloud credentials in the config still reach LiteLLM. With {api_base: ...}, _model_settings(...).extra_args is {'api_base': ...}. Separately, extra_headers, extra_query and extra_body (lines 51-53) are allowlisted on purpose; those are the keys #107 excludes. I'd drop the catch-all extra_args and those three keys.

client.track(
"$ld:ai:graph:total_tokens", ld_context, root_td, total_tokens
)
client.track("$ld:ai:graph:path", ld_context, root_td, len(path))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

$ld:ai:graph:path was retired by #114 (AIC-3211). The other native adapters on main emit $ld:ai:graph:node from on_agent_start (line 202 here), with nodeKey and a 0-based index, and they no longer track path. TESTING.md §2.2 now requires that, and the openai-agents tests assert $ld:ai:graph:path is not tracked. The merge won't conflict here, so this file would quietly keep the old event contract. Copying the current openai-agents/native_graph.py hooks over would fix it.

system_instructions=instructions,
messages=[text_message("user", str(prompt))],
)
result = await importlib.import_module("agents").Runner.run(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing turns off the Agents SDK's own tracing here, and its default exporter sends traces to OpenAI (agents/tracing/processors.py:34). With LiteLLM routing to another provider, that's a second, unexpected destination for prompts whenever OPENAI_API_KEY is set. I haven't checked this at runtime. Spec §2.x.2 asks for it to be suppressed. run_config=RunConfig(tracing_disabled=True) on this call and on run_streamed keeps it per run. JS uses the process-wide setTracingDisabled(true), which I've commented on in #75.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants