From bb3a20943fe6614937e6c2892aa400692d11022c Mon Sep 17 00:00:00 2001 From: nicholascole Date: Tue, 8 Sep 2026 15:20:17 -0700 Subject: [PATCH 01/20] docs: agents parity design specs and implementation plan Adds the Ruby agents parity one-pager (mirror of the Confluence plan), the tools/streaming/teams/secrets/testing sub-specs, and the implementation plan that reconciles them against the Python SDK and the server-side agent API. Co-Authored-By: Claude Fable 5.1 --- docs/design/AGENTS_IMPLEMENTATION_PLAN.md | 616 +++++++++++++++++++ docs/design/AGENTS_PARITY_ONEPAGER.md | 580 +++++++++++++++++ docs/design/AGENT_SECRETS.md | 134 ++++ docs/design/AGENT_STREAMING.md | 153 +++++ docs/design/AGENT_TEAMS.md | 68 ++ docs/design/AGENT_TESTING.md | 109 ++++ docs/design/AGENT_TOOLS_DSL.md | 106 ++++ docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md | 554 +++++++++++++++++ 8 files changed, 2320 insertions(+) create mode 100644 docs/design/AGENTS_IMPLEMENTATION_PLAN.md create mode 100644 docs/design/AGENTS_PARITY_ONEPAGER.md create mode 100644 docs/design/AGENT_SECRETS.md create mode 100644 docs/design/AGENT_STREAMING.md create mode 100644 docs/design/AGENT_TEAMS.md create mode 100644 docs/design/AGENT_TESTING.md create mode 100644 docs/design/AGENT_TOOLS_DSL.md create mode 100644 docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md diff --git a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md new file mode 100644 index 0000000..82ccdf0 --- /dev/null +++ b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md @@ -0,0 +1,616 @@ +# Ruby Agents Parity: Implementation Plan + +Status: proposed, 2026-09-08. Owner: Ruby SDK. + +Source of truth for *what* we build is the Confluence page +[Ruby Agents Parity Plan](https://orkes.atlassian.net/wiki/spaces/ENG/pages/53739522/Ruby+Agents+Parity+Plan) +and its local mirror `docs/design/AGENTS_PARITY_ONEPAGER.md`, plus the sub-specs +`AGENT_TOOLS_DSL.md`, `AGENT_STREAMING.md`, `AGENT_TEAMS.md`, `AGENT_SECRETS.md`, +`AGENT_TESTING.md`, `SDK_AGENT_TESTING_STRATEGY_V2.md`. This document is *how* and *in what +order*, reconciled against two things the design docs were written before we had verified: + +- the Python SDK agents package (`python-sdk/src/conductor/ai/agents`), the parity source; and +- the server implementation (`conductor/agentspan`, branch `feature/llm_mock_impl`), the wire contract. + +Section 2 lists every place the verified contract changes the design. Everything after that is +the work, split into PR-sized slices with acceptance criteria. + +--- + +## 1. Goal and scope + +Ship `Conductor::Agents` in `conductor_ruby`: define an agent in Ruby, serialize it to the same +`agentConfig` Python sends, start it on the server, run the tool workers locally, stream the +result. The three examples in the one-pager (tools, streaming + approval, team + secret) must run +unchanged against a Conductor server that has `conductor.integrations.ai.enabled=true`. + +### In scope (parity surface) + +| Area | Ruby API (from design docs) | Python source | +|---|---|---| +| Definition | `Agent`, `ToolDef`, `ToolType`, `Tools` DSL (`tool def`, `describe`, `requires_approval`, `[]`), `Guardrail`/`RegexGuardrail`/`LlmGuardrail`, `Termination::*`, `Handoff::*`, `CallbackHandler`, `ConversationMemory`, `PromptTemplate` | `agent.py`, `tool.py`, `guardrail.py`, `termination.py`, `handoff.py`, `callback.py`, `memory.py` | +| Serialization | `ConfigSerializer.serialize(agent)` | `config_serializer.py` | +| Transport | `AgentResourceApi`, `AgentClient`, `OrkesClients#get_agent_client`, `SseClient` | `orkes_agent_client.py` | +| Runtime | `AgentRuntime` (`call_sync`, `call_async`, `deploy`, `serve`, `shutdown`), `AgentConfig.from_env`, `Execution`, `ApprovalRequest`, `ToolRegistry`, `Dispatch`, `Secrets` | `runtime/runtime.py`, `_dispatch.py`, `tool_registry.py`, `config.py`, `result.py` | +| Sugar | `add_tool`/`add_agent`/`hands_off_to`/`redact`/`stop_when`/`stop_after`/`on_approval`, `secret()`/`secrets_env()` auto-declaration | Ruby-only | +| Tests | contract tests (schema + 19 golden configs), runtime tests (replay), CI job | `tests/unit/ai`, `examples/agents/_configs` | +| Docs | README section, `docs/agents/`, three runnable examples, CHANGELOG | `docs/agents/` | + +### Out of scope for this plan (Python has them; deliberately deferred) + +Framework agents (`framework`/`rawConfig`: OpenAI Agents, LangGraph, ADK, Claude Agent SDK), +skills, `claude-code` pseudo-provider, `plan_execute` strategy (planner/fallback/plannerContext), +local code execution and CLI tools, OCG retrieval agent, schedules, semantic memory, +`openai_compat.Runner`, OpenTelemetry tracing, liveness monitor / worker restarter, +`scatter_gather`, `a >> b` sequential operator, `@agent`-on-methods (`Agent.from_instance`), +`prefill_tools`, `output_type` (structured output via schema), `router` strategy with a worker +router function, `gate`, `allowed_transitions`, `masked_fields`, `introduction`, +`include_contents`, `thinking_config`, `reasoning_effort`, `context_window_budget`. + +The serializer will be written so any of these can be added as one field + one test later; none +of them changes the architecture. The Confluence page already marks the frameworks row N/A. + +--- + +## 2. Verified contract vs. design docs: what changes + +Each item below is a decision. Where the design docs and the verified contract disagree, the +contract wins and the design docs get a one-line update in Phase 4. + +### 2.1 `agentConfig` wire format (from Python `config_serializer.py`, server `AgentConfig.java`) + +- Keys are camelCase. Every `nil` is dropped. Leaf agents omit `strategy`; it is emitted only + when the agent has sub-agents. `external` and `maxTurns` and `timeoutSeconds` are always + emitted. `approvalRequired`, `stateful`, `enablePlanning` are emitted only when `true`. +- **Credentials asymmetry**: agent-level credentials are top-level `"credentials": [...]`; + tool-level credentials are nested `"config": { "credentials": [...] }`. The server reads + `tool.config.credentials` and falls back to the agent list. +- `strategy` values on the wire are lowercase snake_case: `handoff sequential parallel router + round_robin random swarm manual plan_execute`. Default `handoff`. +- Termination: `{"type": "text_mention"|"stop_message"|"max_message"|"token_usage"|"and"|"or", ...}`. + Handoffs: `{"target", "type": "on_tool_result"|"on_text_mention"|"on_condition", ...}`. + Guardrails: `{"name","position","onFail","maxRetries","guardrailType": "regex"|"llm"|"custom"|"external", ...}`. + Callbacks: `[{"position": "before_model", "taskName": "_before_model"}]`. + Memory: `{"messages": [...], "maxMessages": n}`. Prompt template instructions: + `{"type": "prompt_template", "name", "variables", "version"}`. +- Agent name regex `^[a-zA-Z_][a-zA-Z0-9_-]*$`, `maxTurns >= 1`, duplicate sub-agent names rejected. +- There is **no server-side JSON schema** for `agentConfig`. Python's + `docs/agents/reference/agent-schema.json` (Draft 2020-12, `additionalProperties: false` at the + root) is the contract artifact. We vendor it and the 19 golden configs from + `python-sdk/examples/agents/_configs/`. + +### 2.2 Server requires `model` on every agent, including team parents + +`ModelParser.parse` runs unconditionally on every compile path and raises on a missing or +slash-less model. The one-pager's `team = Agent.new(name: 'bug_desk')` would be rejected. +Python only auto-inherits for `parallel`. + +**Decision D1**: `ConfigSerializer` fills a missing parent `model` from the first sub-agent that +has one, for every multi-agent strategy. A leaf agent with no model raises +`Conductor::Agents::ConfigurationError` at serialize time (unless `external: true`). The +examples stay as written. + +### 2.3 Approval is a `waiting` event plus a `respond` call; `ApprovalRequest` is Ruby-only + +Python has no `ApprovalRequest` and no `on_approval`. On the server, `approvalRequired: true` +on any tool inserts a `HUMAN` task (`_approval_human`, status `IN_PROGRESS`). The SDK +sees SSE `waiting` with `pendingTool = {taskRefName, toolCalls: [{name, args}], response_schema, +...}`. One HUMAN task gates the whole batch of tool calls, so `tool_name`/`parameters` in +`pendingTool` are `null` and `toolCalls` is populated. The reply is +`POST /api/agent/{executionId}/respond` with `{"approved": true}` or +`{"approved": false, "reason": "..."}`. A rejection ends the execution **COMPLETED** with +`output.finishReason == "rejected"` and `output.rejectionReason` set. + +**Decision D2**: `ApprovalRequest` wraps one `waiting` event. It exposes `execution_id`, +`task_ref_name`, `tool_calls` (array of `ToolCall(name:, arguments:)`), and, as sugar, `tool_name` +and argument accessors (`request.amount`) taken from the first tool call. `approve` and +`reject(reason)` post to `respond`. If no `on_approval` block is registered the request is +parked on `execution.pending` and nothing is sent. `finish_reason` becomes `:rejected` from +`output.finishReason`. + +### 2.4 Tool task input carries server-injected keys at the top level + +`inputData` for a worker tool is the LLM's arguments flattened at the top level **plus** +`method` (tool name, always), `_agent_state`, `_agent_tool_name`, and `_allowed_commands` for +`cli` tools. Verified in `conductor-mocks/mocks/agent/tool_happy_path/mappings/04_*.json`. + +**Decision D7**: `Dispatch#coerce_args` strips those four keys before binding keyword arguments, +rejects missing required arguments with `FAILED_WITH_TERMINAL_ERROR` (not a Ruby +`ArgumentError` mid-call; see 2.9 on `city: String` defaults), coerces `String -> Integer/Float/ +Boolean` and `String <-> JSON` for array/object params, returns a Hash as-is and wraps any other +result as `{"result" => value}`. A `_state_updates` key in the result passes through (server +merges it into `_agent_state`). Non-JSON-serializable results fail the task with a clear reason. + +### 2.5 `requiredWorkers` and system workers + +`POST /agent/start` returns `{executionId, agentName, requiredWorkers: [String]}`. The list is +flat task names: every `toolType: worker` tool **by its own name, no prefix**, plus +compiler-generated SIMPLE tasks the SDK must serve: `_termination` whenever +`termination` is set, custom guardrail `taskName`s, callback `_` tasks, +`on_condition` handoff `taskName`s, `stopWhen.taskName`. `_handoff_check`, +`_check_transfer` and `_transfer_to_` are compiled INLINE on this server and +are **not** in the list (Python still registers them for older servers; we do not). + +**Decision**: `ToolRegistry#register_system_workers(required_workers)` serves exactly the +compiler-generated names it recognises by suffix, with Ruby ports of Python's +`TerminationEntry` (`{should_continue, reason}`), `GuardrailEntry` (`{passed, message, on_fail, +fixed_output, guardrail_name, should_continue}`), `CallbackEntry`, and the `on_condition` +handoff worker. An unrecognised name logs a warning listing it, because a task with no worker +sits SCHEDULED forever. Since `stop_when`/`stop_after` serialize to `termination`, the +termination worker is needed for the team example and is Phase 2, not later. + +### 2.6 SSE framing and reconnect (server `AgentStreamRegistry`) + +`GET /api/agent/stream/{executionId}`, `Accept: text/event-stream`. Server writes `:connected` +first, then `id:\nevent:\ndata:\n\n`, ids start at 1 per execution, a +`:heartbeat` comment every 15 s, emitter never times out, stream is closed after `done` or +`error`. Replay buffer of 200 events for 5 minutes after completion. `Last-Event-ID` is an HTTP +request header bound to a Java `Long`: send a bare integer; when absent the server replays from +0, so connecting after start loses nothing. Child sub-workflow events are aliased onto the +parent stream. Event types: `thinking tool_call tool_result handoff waiting guardrail_pass +guardrail_fail error done` plus `context_condensed subagent_start subagent_stop`. `done.output` +is the workflow output `{result, finishReason, context, rejectionReason}`. + +**Decision D3**: `SseClient` reconnects with `Last-Event-ID` after any drop with a fixed 1 s +backoff, raises `SseUnavailableError` if the first connect fails, and the runtime falls back to +polling `GET /agent/{id}/status` (fields `isComplete`, `isWaiting`, `output`, `pendingTool`, +`reasonForIncompletion`) every 0.5 s. This is simpler than Python's task-graph synthesis and +loses only `partial_text` in fallback mode. `partial_text` is the concatenation of `thinking` and +`message` `content` fields; there is no separate delta event. + +### 2.7 SSE cannot go through the existing `RestClient` + +`RestClient` has a 120 s total timeout, retry middleware, and buffers bodies. `SseClient` opens +its own long-lived connection: plain `Net::HTTP` with `read_body` streaming (no adapter +uncertainty, no new gem), read timeout disabled, headers from `ApiClient#get_authentication_headers` +so token refresh stays in one place. Faraday `on_data` through `faraday-net_http_persistent 2.3.1` +is the alternative; spike both in Phase 2 slice 1 and keep whichever streams the recorded +scenario end to end. + +### 2.8 Secrets ride on `runtimeMetadata`; both Ruby models lack the field + +`TaskDef.runtimeMetadata` is `List` (names). `Task.runtimeMetadata` is +`Map` (values), filled at poll time by `RuntimeMetadataResolver` (secret store, +then server env with `CONDUCTOR_SECRET_` / `CONDUCTOR_ENV_` prefixes), silently omitted on miss, +never persisted. Needs conductor-oss with PR #1255 (3.32.0-rc.8+). Neither field exists in +`lib/conductor/http/models/task.rb` or `task_def.rb` today. + +**Decision D5**: add both fields. `Secrets.secret(name)` reads +`TaskContext.current.task.runtime_metadata[name]`, then `ENV[name]`, then raises +`CredentialNotFoundError`. `TaskContext` is stored in `Thread.current[]`, which is fiber-local in +Ruby, so this already satisfies the "fiber-local, never writes ENV" rule with no new binding +mechanism. `secrets_env(*names)` returns a Hash for `system`/`spawn`/`Open3`. + +### 2.9 `tool def` type inference needs the AST, and so does secret scanning + +`Method#parameters` returns kinds and names only, never default values, so `city: String` vs +`units: 'metric'` is indistinguishable by reflection. Both type inference and `secret('...')` +scanning therefore use `RubyVM::AbstractSyntaxTree.of(method)` (MRI 2.6+, works when the source +file is on disk; in irb/`eval` needs `RubyVM.keep_script_lines = true`, Ruby 3.2+). Prism ships +only with Ruby 3.3, and the gem targets Ruby >= 3.0, so no parser dependency. + +**Decision D9**: `Tools#get_json_schema_def` maps `String/Integer/Float/Hash` class defaults to +required typed properties, literal defaults to optional properties with `default`, `[String]` to +`array` with `items`, `%w[a b]` to `enum` (optional, default first), `true/false` to boolean. +When the AST is unavailable the fallback is Python's behaviour: every keyword param gets `{}` +and `:keyreq` params are required; secret declaration then needs the explicit +`add_tool ..., credentials:`. Positional parameters raise at `tool def`. A one-line spike +proving `AbstractSyntaxTree.of` works for a top-level `def` and for a method in a module is the +first task of Phase 1. + +Runtime consequence: because `city: String` has the class as its Ruby default, the worker must +never call the method with a required argument missing (it would receive the `String` class). +That is why D7 validates required arguments before the call. + +### 2.10 Tool task definitions + +Python registers each tool with `retryCount 2, retryDelaySeconds 2, retryLogic LINEAR_BACKOFF, +timeoutSeconds 0, responseTimeoutSeconds 10, timeoutPolicy RETRY, runtimeMetadata [names]`, +`overwrite_task_def: true`, `worker_id "agent-sdk"`. The server also registers a TaskDef for +every required worker (with `responseTimeoutSeconds 3600`); ours overwrites it, exactly as the +recorded scenario shows (`02_put_api_metadata_taskdefs.json`). + +**Decision D6**: same defaults, via `Worker.define(name, register_task_def: true, +overwrite_task_def: true, task_def_template: TaskDef.new(...))`. `TaskDefinitionRegistrar# +build_task_definition` already honours `task_def_template`; it must stop overriding +`timeout_seconds`/`response_timeout_seconds` when the template sets them (it currently uses +`||=`, which is fine for non-nil values but `timeout_seconds: 0` is truthy in Ruby, so no change +needed; add a test). + +### 2.11 MCP discovery is server-side; drop `McpDiscovery` + +The compiler emits `LIST_MCP_TOOLS` and an LLM filter chain before the agent loop for every +`toolType: mcp` tool, and dispatches `CALL_MCP_TOOL` at runtime. Python's `mcp_discovery.py` +served its now-legacy local-compile path. **Decision D4**: no `McpDiscovery` class; `ToolDef.mcp` +serializes `{server_url, headers, tool_names, max_tools}` and nothing else happens client-side. +The one-pager's runtime diagram loses one box. + +### 2.12 Stateful runs, domains, sessions + +`AgentStartRequest.runId` makes the server map every required worker to that task domain. Python +sends `runId = uuid4().hex` only when the agent or any tool is `stateful`, then polls with +`domain == run_id`. `sessionId` is passed straight through. **Decision**: same. `Agent#stateful` +and `ToolDef#stateful` exist from Phase 1; domain wiring is a Phase 2 task with one test. + +### 2.13 Finish reasons and statuses + +Server `finishReason` is uppercase (`STOP TOOL_CALLS MAX_TOKENS CONTENT_FILTER LENGTH`) except +the literal `"rejected"`. Execution status is the Conductor workflow status (`RUNNING COMPLETED +FAILED TIMED_OUT TERMINATED PAUSED`); `executionId == workflowId`. **Decision D8**: +`Execution#finish_reason` is `:stop | :tool_calls | :length | :content_filter | :rejected | +:error | :cancelled | :timeout`, derived the way Python's `_derive_finish_reason` does. + +### 2.14 `redact` needs verification + +`redact` is specified as `RegexGuardrail(position: :output, on_fail: :fix)`. Whether +`GuardrailCompiler` rewrites matched text for a regex guardrail with `onFail: fix` was not +confirmed. Phase 1 serializes it as specified; Phase 3's replay scenario for the team example +verifies behaviour, and if the server does not rewrite, `redact` switches to `on_fail: :retry` +and the doc is updated. + +### 2.15 Testing infrastructure reality + +`SDK_AGENT_TESTING_STRATEGY_V2.md` (in-server `mockLLM` provider, record mode) is **not +implemented on the server yet**: only fixture primitives exist on `feature/llm_mock_impl` +(`ai/.../testing/LlmFixture*.java`, `llm-fixture.schema.json`); no `MockLLM` provider, no +`conductor.ai.mock-llm.record` property, no test-server task. What exists today is +`conductor-mocks` with one normalized WireMock scenario, `agent/tool_happy_path`, recorded +against a real server for this SDK. + +**Decision**: contract tests need no server and land first. Runtime tests use WireMock replay of +`tool_happy_path` now (it is the only recording, and it covers start, TaskDef PUT, poll with +`runtimeMetadata: []`, update-v2, SSE). Approval, secrets and team scenarios are recorded into +`conductor-mocks` as they become runnable. The mockLLM functional suite is a Phase 3 follow-up +gated on the server; the spec helper is written so `mocks:` (WireMock) and `model: +'mockLLM/'` (real server) coexist. + +### 2.16 Prerequisite already called out in the design: instance-level token cache + +`Configuration` keeps `auth_token`/`token_update_time` as class-level state. Two runtimes with +different credentials in one process (tests, multi-tenant workers) would share one token. Only +`ApiClient` reads it. Move to instance level; keep the class accessors as deprecated shims for +one release. + +### 2.17 Local toolchain + +Ruby is not installed on this machine (`ruby: command not found`; no rbenv/rvm/mise). Docker and +podman are. Either install Ruby 3.3 or run the suite in `ruby:3.3-alpine` (the repo's +`Dockerfile` base). This is the first checklist item in Phase 0. + +--- + +## 3. Target layout + +``` +lib/conductor/agents.rb # require 'conductor/agents'; requires everything below +lib/conductor/agents/ + version.rb? # no: reuse Conductor::VERSION + errors.rb # ConfigurationError, CredentialNotFoundError, AgentApiError, + # AgentNotFoundError, SseUnavailableError, ToolSerializationError + agent.rb # Agent (definition + sugar + call_sync/call_async delegating to runtime) + tool_def.rb # ToolDef, ToolType, PrefillToolCall (minimal), factories http/mcp/human/agent + tools.rb # Tools module: tool, describe, requires_approval, [], schema + secret scan + tools/schema_builder.rb # AST -> JSON schema (D9) + tools/secret_scanner.rb # AST -> literal secret()/secrets_env() names + tools/ruby_llm_adapter.rb # RubyLLM::Tool class -> ToolDef (only if defined?(RubyLLM)) + guardrail.rb # Guardrail, RegexGuardrail, LlmGuardrail, GuardrailResult + termination.rb # Termination::Condition, TextMention, StopMessage, MaxMessage, TokenUsage, And, Or + handoff.rb # Handoff::Condition, OnToolResult, OnTextMention, OnCondition + callback_handler.rb # CallbackHandler + POSITION_TO_METHOD + memory.rb # ConversationMemory + prompt_template.rb # PromptTemplate + config_serializer.rb # ConfigSerializer.serialize(agent) -> Hash (camelCase) + runtime/agent_config.rb # AgentConfig.from_env + runtime/agent_runtime.rb # AgentRuntime: call_sync/call_async/deploy/serve/shutdown + runtime/execution.rb # Execution, ToolCall, TokenUsage + runtime/approval_request.rb # ApprovalRequest + runtime/sse_client.rb # SseClient: each_event(execution_id, last_event_id:), reconnect + runtime/status_poller.rb # polling fallback over /agent/{id}/status + runtime/tool_registry.rb # ToolRegistry: register_tool_workers, register_system_workers + runtime/dispatch.rb # Dispatch.run_tool_task, coerce_args + runtime/system_workers.rb # termination / guardrail / callback / handoff worker bodies + runtime/secrets.rb # Secrets.secret, secrets_env +lib/conductor/http/api/agent_resource_api.rb # AgentResourceApi +lib/conductor/client/agent_client.rb # AgentClient (Hash in / Hash out, like Python) +lib/conductor/orkes/orkes_clients.rb # + get_agent_client +lib/conductor/http/models/task.rb # + runtime_metadata (Hash) +lib/conductor/http/models/task_def.rb # + runtime_metadata (Array) +spec/conductor/agents/** # unit + contract specs (no server) +spec/agents/** # runtime specs (WireMock replay / real server), tagged +spec/fixtures/agents/agent-schema.json # vendored from python-sdk docs/agents/reference +spec/fixtures/agents/configs/*.json # 19 golden configs vendored from python-sdk examples/agents/_configs +examples/agents/{weather,support_approval,bug_desk}.rb +examples/agents/dump_agent_configs.rb # regenerates goldens from Ruby for cross-SDK diff +docs/agents/README.md + concepts/*.md +``` + +Namespacing: definition classes live directly under `Conductor::Agents` so `include +Conductor::Agents` gives `Agent`, `RegexGuardrail`, `Termination`, `Handoff`, and the `tool`, +`describe`, `requires_approval`, `secret`, `secrets_env` methods. `Tools` is included into +`Conductor::Agents` so top-level `include Conductor::Agents` works; `extend +Conductor::Agents::Tools` on a module works because `tool` resolves `method(name)` first and +falls back to `instance_method(name)` + `module_function name`. + +--- + +## 4. Work breakdown + +Each slice is one reviewable PR. Every PR: `bundle exec rubocop`, `bundle exec rspec +spec/conductor/`, coverage not lower than before, `bundle exec ruby -Ilib -e "require +'conductor/agents'"` loads. Sizes: S < 1 day, M 1-3 days, L 3-5 days. + +### Phase 0: prerequisites (no agent code yet) + +**PR 0.1 Toolchain + hygiene (S)** +- Install Ruby 3.3 locally or document `docker run --rm -v $PWD:/app -w /app ruby:3.3 bundle exec rspec`. +- Remove unused `vcr` dev dependency (called out in `AGENT_TESTING.md`); keep `webmock`. +- Add `json_schemer` as a development dependency for contract tests. +- Acceptance: `bundle exec rspec spec/conductor/` green locally. + +**PR 0.2 Instance-level auth token cache (S)** (2.16) +- `Configuration`: `@auth_token`, `@token_update_time` per instance; `update_token`, + `auth_token`, `token_update_time` read instance state. Class-level accessors remain, emit a + deprecation warning once, and read/write a process-wide fallback only when the instance has none. +- `ApiClient` unchanged in behaviour; add a spec that two configurations hold two tokens. +- Acceptance: all existing specs pass; new spec proves isolation. + +**PR 0.3 `runtimeMetadata` on models + registrar (S)** (2.8, 2.10) +- `Task`: `runtime_metadata: 'Hash'` / `:runtimeMetadata`, default `{}`. +- `TaskDef`: `runtime_metadata: 'Array'` / `:runtimeMetadata`, default `[]`. +- `TaskDefinitionRegistrar`: test that `task_def_template` with `timeout_seconds: 0`, + `response_timeout_seconds: 10`, `runtime_metadata: ['GH_TOKEN']` survives `build_task_definition`. +- Acceptance: model round-trip specs; recorded `04_*` poll body deserializes `runtimeMetadata`. + +**PR 0.4 Agent transport (M)** (server §2 of the API report) +- `Http::Api::AgentResourceApi` over `ApiClient#call_api`, `return_type: 'Hash'`: + - `start(body)` POST `/agent/start`; `deploy(body)` POST `/agent/deploy`; `compile(body)` POST `/agent/compile` + - `status(id)` GET `/agent/{executionId}/status`; `execution(id)` GET `/agent/execution/{executionId}` + - `executions(params)` GET `/agent/executions` (`start,size,sort,freeText,status,agentName,sessionId`) + - `respond(id, body)` POST `/agent/{executionId}/respond`; `stop(id)`; `signal(id, message)` + - `pause(id)` PUT `/agent/{executionId}/pause`; `resume(id)` PUT; `cancel(id, reason:)` DELETE `/agent/{executionId}/cancel` + - `list` GET `/agent/list`; `get(name, version:)` GET `/agent/{name}`; `delete(name, version:)` +- `Client::AgentClient` wrapping it with Python's method names (`start_agent`, `deploy_agent`, + `compile_agent`, `get_status`, `get_execution`, `list_executions`, `respond`, `approve`, + `reject`, `send_message`, `stop`, `signal`, `pause`, `resume`, `cancel`). Map 404 to + `AgentNotFoundError`, other `ApiError` to `AgentApiError` carrying the `{error, status}` body. +- `OrkesClients#get_agent_client`. +- Specs mirror `spec/conductor/client/prompt_client_spec.rb` (instance doubles) plus WebMock + specs asserting exact paths and bodies for start/respond/status. +- Acceptance: every endpoint in the table has a spec asserting method + path + body keys. + +### Phase 1: definition layer and serializer (no server) + +**PR 1.1 Spike: AST availability (S, half day, can be a scratch script)** (2.9) +- Prove `RubyVM::AbstractSyntaxTree.of(method)` returns `KW_ARG` defaults and `secret('X')` + call nodes for: a top-level `tool def` in a file, a method in a module with `extend Tools`, a + method defined in irb with `keep_script_lines`. Record the matrix in `tools/schema_builder.rb` + comments. If the top-level-in-file case fails on any supported Ruby (3.0-3.3), the fallback in + D9 becomes the primary path and the DSL doc changes before any more code is written. + +**PR 1.2 ToolDef, ToolType, factories (M)** +- `ToolType` constants: `worker http api mcp human agent_tool generate_image generate_audio + generate_video generate_pdf rag_index rag_search pull_workflow_messages cli`. +- `ToolDef` fields and defaults exactly as Python: `name, description '', input_schema {}, + output_schema {}, func nil, approval_required false, timeout_seconds nil, tool_type 'worker', + config {}, guardrails [], credentials [], stateful false, max_calls nil, retry_count 2, + retry_delay_seconds 2, retry_policy 'linear_backoff'`. +- Factories from the one-pager: `ToolDef.http(name, url, method: 'GET', description:, headers:, + input_schema:, credentials:)`, `ToolDef.mcp(server_url, name: 'mcp_tools', headers:, + tool_names:, max_tools: 64, credentials:)`, `ToolDef.human(name, description:, input_schema:)`, + `ToolDef.agent(agent, name:, description:, retry_count:, retry_delay_seconds:, optional:)`. + `${NAME}` header placeholders must appear in `credentials` or raise. Media/RAG/api factories + are one method each and can ride along if cheap; not required by the examples. +- Specs: config keys per type match Python's `_serialize_tool` (`url method headers accept + contentType`; `server_url headers tool_names max_tools`; `agent` replaced by `agentConfig`). + +**PR 1.3 Tools DSL: `tool def`, `describe`, `requires_approval`, `[]`, schema, secret scan (L)** +- `tool(name)`: resolves the method, builds `ToolDef(name:, description: humanize(name), + input_schema: SchemaBuilder.for(method), credentials: SecretScanner.scan(method), func: ->(**kw) {...})`, + registers it in a per-`self` registry (`Tools#[]`), returns the ToolDef. Positional params raise. +- `describe(name, text)`, `requires_approval(name)` mutate the registered ToolDef. +- `SchemaBuilder` implements the D9 table; `required` only when non-empty; no `$schema` key + (Python emits a bare `{"type":"object","properties":{...},"required":[...]}`). +- `SecretScanner` collects string-literal first arguments of `secret(...)` and every + string-literal argument of `secrets_env(...)`; ignores dynamic names. +- `RubyLlmAdapter`: when `defined?(RubyLLM::Tool)` and `add_tool` receives such a class, build a + `ToolDef` from `.name`, `.description`, `.parameters`-derived schema, `func: ->(**kw) { + klass.new.execute(**kw) }`. +- Specs: one per row of the type table; secret scan positive/negative/dynamic; module-extend + form; RubyLLM adapter behind a stub class. + +**PR 1.4 Guardrails, termination, handoffs, callbacks, memory, prompt template (M)** +- Straight ports with Ruby naming. Validation rules from Python: guardrail `position` in + `input|output`, `on_fail` in `retry|raise|fix|human`, `human` illegal on `input`, + `max_retries >= 1`; `MaxMessage >= 1`; `TokenUsage` needs at least one limit; `&`/`|` + flatten same-type children; `RegexGuardrail(patterns, mode: :block|:allow, message:)`; + `LlmGuardrail(model, policy, max_tokens:)`; `Guardrail.new(name:)` with no block is external. +- `ConversationMemory` with `add_user_message` etc. and `_trim` semantics (keep system messages). +- Specs per class including combinator flattening. + +**PR 1.5 Agent (M)** (2.1, 2.2, 2.12) +- Constructor keywords: `name:, model: nil, instructions: '', tools: [], agents: [], strategy: + :handoff, router: nil, output_type: nil, guardrails: [], memory: nil, termination: nil, + handoffs: [], callbacks: [], credentials: [], max_turns: 25, max_tokens: nil, + timeout_seconds: 0, temperature: nil, stateful: false, metadata: nil, description: nil, + external: false`. +- Validation: name regex, strategy enum (symbols, serialized lowercase), `max_turns >= 1`, + duplicate sub-agent names, `router` required for `:router`. +- Sugar: `add_tool(tool, credentials: nil)` accepts `Symbol` (looks up `Tools` registry of the + caller and of `Conductor::Agents`), `ToolDef`, a `Tools`-extended module (adds all), a + RubyLLM class; `add_tools(*)`, `add_agent(s)`, `hands_off_to(agent, on:)` -> + `Handoff::OnTextMention`, `redact(words)` -> `RegexGuardrail(position: :output, on_fail: :fix, + mode: :block, name: "#{name}_redact")`, `stop_when(text)` -> `Termination::TextMention`, + `stop_after(messages:)` -> `Termination::MaxMessage` (combined with `|` when both set), + `on_approval(&block)`, `strategy=`. +- `call_sync(prompt, session_id: nil)` and `call_async(prompt, session_id: nil, &on_done)` + delegate to `Conductor::Agents.runtime` (Phase 2); until then they raise `NotImplementedError` + with a pointer. +- Specs: validation matrix; sugar produces the parity objects listed in `AGENT_TEAMS.md`. + +**PR 1.6 ConfigSerializer + contract tests (L)** (2.1, 2.2) +- Emission order and conditions per Python's `serialize`, `_serialize_tool`, + `_serialize_guardrail`, `_serialize_termination`, `_serialize_handoff`, `_serialize_memory`, + plus D1 (parent model inheritance) and the `ConfigurationError` for model-less leaves. +- Callbacks serialize to `{position, taskName: "#{name}_#{position}"}` for every + `CallbackHandler` method a handler overrides. +- Vendor `spec/fixtures/agents/agent-schema.json` and the 19 `configs/*.json`. Port + `dump_agent_configs.py` to `examples/agents/dump_agent_configs.rb` so the Ruby definitions of + the same 19 agents live in the repo; the spec compares `serialize(agent)` to each golden file + (`sort_keys`-insensitive Hash equality, model string parametrised by env like Python's dump + script). Validate every serialized config against the schema; assert an unknown root key fails. +- Acceptance: 19/19 golden matches, schema valid, and `examples/agents/dump_agent_configs.rb` + regenerates byte-identical JSON (sorted keys, 2-space indent) for cross-SDK diffs. + +### Phase 2: runtime and transport + +**PR 2.1 Spike + SseClient (M)** (2.6, 2.7) +- Implement over `Net::HTTP` streaming first; if `Faraday` `on_data` via the persistent adapter + streams the recorded `06_get_api_agent_stream_exec_1.json` body (WireMock dribbles it over ~3 s) + with less code, switch. Parser handles `:` comments (heartbeat), `id:`, `event:`, `data:` + (multi-line joined by `\n`), blank-line boundaries, JSON parse failure -> `{"content" => raw}`. +- `each_event(execution_id, last_event_id: nil)` is an Enumerator that stops after `done`/`error`, + reconnects with `Last-Event-ID: ` after a drop, raises `SseUnavailableError` on + first-connect failure or 15 s of heartbeat-only silence before the first real event. +- Specs with WebMock streaming bodies (WebMock can serve a String body; chunk pacing is not + needed for parsing tests) covering comments, multi-line data, reconnect id, terminal events. + +**PR 2.2 Secrets + Dispatch (M)** (2.4, 2.8) +- `Secrets.secret(name)` / `secrets_env(*names)` per D5; `CredentialNotFoundError` message names + the tool and the missing key. +- `Dispatch.run_tool_task(task, tool_def)` per D7; returns a `TaskResult`, sets `worker_id + 'agent-sdk'`, `FAILED` with `reason_for_incompletion` on exceptions, `FAILED_WITH_TERMINAL_ERROR` + on missing required args, missing credentials, or `ToolSerializationError`. Circuit breaker + (10 consecutive failures per tool) is optional; include only if trivial. +- Specs: strips injected keys (uses the recorded `04_*` input verbatim), coercions, missing + required arg, Hash vs scalar result, `_state_updates` passthrough, secret from + `runtime_metadata` vs ENV vs missing, `secrets_env` returns only requested names. + +**PR 2.3 ToolRegistry + system workers (M)** (2.5, 2.10) +- `register_tool_workers(tools, agent_name, domain)` builds one `Worker.define` per `worker`/`cli` + tool with the D6 TaskDef template and `Dispatch` body; server-side tool types are skipped. +- `register_system_workers(required_workers, agent)` matches `_termination`, + `_` callbacks, custom guardrail names, `on_condition` handoff task names; + bodies in `system_workers.rb` are ports of Python's `TerminationEntry`, `GuardrailEntry`, + `CallbackEntry`, `HandoffCondition`. Unknown names warn. +- Workers run on a dedicated `TaskHandler` owned by the runtime with `poll_interval` / + `thread_count` from `AgentConfig`, `register_task_definitions: true`. +- Specs: worker set for the three examples; TaskDef body equals the recorded `02_*` PUT body + (with `runtimeMetadata: []`); termination worker returns `should_continue: false` on + `TextMention` match; unknown required worker logs. + +**PR 2.4 AgentRuntime, Execution, ApprovalRequest (L)** (2.3, 2.6, 2.12, 2.13) +- `AgentConfig.from_env`: `CONDUCTOR_AGENT_WORKER_POLL_INTERVAL` (100 ms), + `CONDUCTOR_AGENT_WORKER_THREADS` (1), `CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER` (false), + `CONDUCTOR_AGENT_STREAMING_ENABLED` (true); Python boolean parsing (`true/1/yes/on`). +- `AgentRuntime#call_async(agent, prompt, session_id:, &on_done)`: serialize; `run_id = + SecureRandom.hex` when stateful; `POST /agent/start` with `{agentConfig, prompt, sessionId, + runId?}`; register tool + system workers from `requiredWorkers` under `domain = run_id`; start + the handler; create `Execution`; spawn the stream thread (`SseClient`, falling back to + `StatusPoller` on `SseUnavailableError` or when streaming is disabled); return immediately. + `call_sync` is `call_async(...).result`. +- Stream thread: `thinking`/`message` append `partial_text`; `tool_call`/`tool_result` pair into + `execution.tool_calls`; `waiting` with `pendingTool.toolCalls` builds an `ApprovalRequest` and + invokes `agent.on_approval` (else parks it on `execution.pending`); `done` sets `result = + output['result']`, `finish_reason` per D8, fetches `tokenUsage` from `GET /agent/execution/{id}` + (recursing `tasks[].subWorkflowId`), marks done, fires `on_done`; `error` sets `:error`. All + callbacks are wrapped so a user exception never kills the stream thread. +- `Execution`: `execution_id, done?, waiting?, result (blocks with optional timeout), + partial_text, finish_reason, tool_calls, token_usage, pending, pause, resume, cancel(reason:), + stop, signal(text)`; `Execution.find(id)` builds from `get_status` + `get_execution`. +- `deploy(*agents)` -> `POST /agent/deploy` each, returns agent names; `serve(*agents)` deploys, + registers workers, blocks until `INT`/`TERM`, then `shutdown`; `shutdown` stops the handler and + stream threads. +- `Conductor::Agents.runtime` memoised default (`Configuration.new` from env), + `Conductor::Agents.configure(configuration:, agent_config:)`. +- Specs (WebMock, no WireMock): start body equals `serialize(agent)` + prompt; workers + registered for `requiredWorkers`; approval approve/reject post the exact bodies; rejection -> + `:rejected`; fallback poller path; `on_done` fires once; user exception in `on_approval` is + logged not raised. + +### Phase 3: end-to-end tests and CI + +**PR 3.1 WireMock replay of `agent/tool_happy_path` (M)** (2.15) +- `spec/agents/agents_helper.rb`: `mocks: 'agent/tool_happy_path'` metadata boots + `wiremock/wiremock:3x` with the scenario dir mounted (Docker via `docker`/`podman` CLI; skip + with a clear message when neither exists or `CONDUCTOR_MOCKS_DIR` is unset), points + `CONDUCTOR_SERVER_URL` at it, asserts `GET /__admin/requests/unmatched` is empty in `after`. +- The weather spec from `AGENT_TESTING.md` passes end to end: start, TaskDef PUT, poll, tool + runs in-process, update-v2, SSE `done`, `call_sync` returns the recorded text. +- CI: new `agents-replay` job that checks out `conductor-oss/conductor-mocks` and runs + `bundle exec rspec spec/agents`. Runs on every PR including forks (no secrets). + +**PR 3.2 Record the remaining scenarios (M, blocked on a server with AI enabled and keys)** +- `approval_approve`, `approval_reject`, `secrets_runtime_metadata`, `team_handoff` using the + three examples, recorded with WireMock's snapshot recorder and `scripts/normalize.py` per the + `conductor-mocks` README; PRs to `conductor-mocks`; specs here tagged with each scenario. + `team_handoff` also settles 2.14 (`redact` semantics). + +**PR 3.3 mockLLM functional suite (L, blocked on server)** +- When `MockLLM` lands in conductor-oss: `spec/agents/functional/*_spec.rb` for the six-scenario + catalog in `SDK_AGENT_TESTING_STRATEGY_V2.md`, asserting persisted LLM task input and tool + tasks via `WorkflowClient#get_workflow`, plus a CI job that builds the pinned server. Not a + blocker for releasing the feature; replay covers the wire contract until then. + +### Phase 4: docs, examples, release + +**PR 4.1 Examples + docs (M)** +- `examples/agents/weather.rb`, `support_approval.rb`, `bug_desk.rb` exactly as in the one-pager, + plus `dump_agent_configs.rb` from Phase 1. +- `docs/agents/README.md` and `concepts/{tools,streaming-hitl,teams,secrets}.md` adapted from the + four design sub-specs; README gets an "Agents" section after "LLM/AI Tasks"; `AGENTS.md` and + `DESIGN.md` get the new layer in the architecture diagrams and directory listing. +- Update the design docs for the decisions in section 2 (D1 model inheritance, D2 approval batch + semantics, D4 no `McpDiscovery`, D9 AST fallback, 2.15 testing reality) and refresh the + Confluence page from the one-pager (read the live page immediately before writing; it is + edited concurrently). +- `CHANGELOG.md` Unreleased: Added `Conductor::Agents`, `AgentClient`, `runtimeMetadata` + fields; Changed instance-level token cache; Removed `vcr`. + +**PR 4.2 Release (S)** +- Version bump (minor), gem build, tag. `conductor/agents` stays an explicit require; `require + 'conductor'` does not load it. + +--- + +## 5. Sequencing and parallelism + +``` +0.1 -> 0.2 -> 0.3 -> 0.4 ----------------------------------. + \ \ + 1.1 -> 1.2 -> 1.3 -> 1.4 -> 1.5 -> 1.6 -> 2.3 -> 2.4 -> 3.1 -> 4.1 -> 4.2 + 2.1 --' \ + 2.2 --' '-> 3.2 (server + keys) + '-> 3.3 (server mockLLM) +``` + +Phase 0 and Phase 1 are independent after 0.3 and can run in parallel. 2.1 and 2.2 depend only +on Phase 0 and 1.2. The critical path is 1.1 -> 1.3 -> 1.6 -> 2.4 -> 3.1. + +Rough total: Phase 0 ~3 days, Phase 1 ~8 days, Phase 2 ~9 days, Phase 3.1 ~2 days, Phase 4 ~3 +days. About five weeks for one engineer to a releasable feature with replay-tested wire contract; +3.2 and 3.3 follow as the server pieces land. + +--- + +## 6. Risks and open questions + +| # | Risk | Mitigation | +|---|---|---| +| R1 | `RubyVM::AbstractSyntaxTree.of` unavailable or lossy on some supported Ruby (3.0-3.3) or non-MRI | PR 1.1 spike first; D9 fallback (`{}` schema, explicit `credentials:`) is always available; document limits | +| R2 | Server rejects team parent without `model` | D1 inherits from first child at serialize time; contract test | +| R3 | `redact` (`regex` + `fix`) may not rewrite on the server | Verify in 3.2; fall back to `retry` | +| R4 | SSE through WireMock is buffered; without `chunkedDribbleDelay` `done` can arrive before the tool executes | `normalize.py` already paces the body; assert unmatched requests empty | +| R5 | `mockLLM` functional testing not available on the server yet | Replay now, functional later (3.3); spec helper supports both | +| R6 | `runtimeMetadata` needs conductor-oss >= 3.32.0-rc.8 (PR #1255); older servers omit it | `secret()` falls back to `ENV`; document minimum version | +| R7 | OSS vs Orkes model string semantics differ (provider type key vs integration name) | Docs say "left side is the integration name"; on OSS that is the provider key with `*_API_KEY` env; note both | +| R8 | Long-lived SSE thread plus worker threads plus user callbacks: exceptions in user blocks | Wrap every user callback; `execution.error` captures; never let the stream thread die silently | +| R9 | Class-level token cache change alters behaviour for users relying on cross-instance sharing | Deprecated shim for one release; changelog entry | +| R10 | `Last-Event-ID` must be an integer; a stray string 400s the reconnect | Parser stores `id` as Integer; reconnect sends `to_s` of it only | +| R11 | Server auto-registers TaskDefs with `responseTimeoutSeconds 3600`; ours overwrites with 10 | Same as Python; document that lease extension keeps long tools alive (follow-up: `lease_extend_enabled`) | + +Open questions for the server team (none block Phases 0-2): + +1. Confirm `regex` guardrail with `onFail: fix` rewrites content (2.14). +2. Timeline for `MockLLM` provider and the test-server entry point (2.15). +3. Whether `_termination` is emitted for `termination` on sub-agents too, so the + registry can register it per agent in a team. diff --git a/docs/design/AGENTS_PARITY_ONEPAGER.md b/docs/design/AGENTS_PARITY_ONEPAGER.md new file mode 100644 index 0000000..fde5846 --- /dev/null +++ b/docs/design/AGENTS_PARITY_ONEPAGER.md @@ -0,0 +1,580 @@ +# Ruby Agents Parity Plan + +Port `python-sdk/src/conductor/ai/agents` → `conductor_ruby` as `Conductor::Agents`. Same `agentConfig` on the wire; server compiles, SDK serializes + runs workers. + +## Classes + +### Definition (serialized to agentConfig) + +```mermaid +classDiagram + direction LR + class Agent { + +new(name:, model:, instructions:, tools: [], agents: [], **opts) + +String name + +String model "provider/model" + +String|PromptTemplate instructions + +List~ToolDef~ tools + +List~Agent~ agents + +Symbol strategy + +Agent|Proc router + +Hash output_type + +List~Guardrail~ guardrails + +ConversationMemory memory + +TerminationCondition termination + +List~HandoffCondition~ handoffs + +List~CallbackHandler~ callbacks + +List~String~ credentials + +Integer max_turns + +Boolean stateful + +add_tool(tool, credentials: nil) add_tools(*tools) + +add_agent(agent) add_agents(*agents) same as agents: + +hands_off_to(agent, on:) + +redact(words) stop_when(text) stop_after(messages:) + +call_sync(prompt, session_id:) String + +call_async(prompt, session_id:, &on_done) Execution + +on_approval() &block + } + class ConfigSerializer { + +serialize(agent) Hash + } + class ToolDef { + +String name + +String description + +Hash input_schema + +Hash output_schema + +Proc func + +ToolType tool_type + +Hash config per type + +Boolean approval_required + +List~String~ credentials + +Integer retry_count + +call(**args) PrefillToolCall + +http(name, url, method) ToolDef$ + +mcp(server_url) ToolDef$ + +human(name) ToolDef$ + +agent(agent) ToolDef$ + } + class Tools { + <> + +tool(method_name) ToolDef + +requires_approval(name) + +describe(name, text) + +[](name) ToolDef + -get_json_schema_def(method) Hash + -scan_secrets(method) List~String~ literal secret() names + } + class ToolType { + <> + worker + http + api + mcp + human + agent_tool + generate_image + generate_audio + generate_video + generate_pdf + rag_index + rag_search + pull_workflow_messages + } + class Guardrail { + +String name + +Symbol position input|output + +Symbol on_fail retry|raise|fix|human + +Integer max_retries + +Proc func + } + class RegexGuardrail + class LlmGuardrail + class TerminationCondition { + <> + +should_terminate(ctx) TerminationResult + +&(other) And + +|(other) Or + } + class TextMention + class StopMessage + class MaxMessage + class TokenUsage + class HandoffCondition { + +String target + +should_handoff(ctx) Boolean + } + class OnToolResult + class OnTextMention + class OnCondition + class CallbackHandler { + +on_agent_start() on_agent_end() + +on_model_start() on_model_end() + +on_tool_start() on_tool_end() + } + class ConversationMemory { + +List~Hash~ messages + +Integer max_messages + } + class PromptTemplate { + +String name + +Hash variables + +Integer version + } + + Agent "1" o-- "*" ToolDef : tools, shared by name + Agent "1" o-- "*" Agent : agents + Agent "1" o-- "*" CallbackHandler : callbacks + Agent "1" *-- "*" Guardrail : guardrails + Agent "1" *-- "0..1" TerminationCondition : termination + Agent "1" *-- "*" HandoffCondition : handoffs + Agent "1" *-- "0..1" ConversationMemory : memory + Agent --> "0..1" PromptTemplate : instructions + Agent --> "0..1" Agent : router + ConfigSerializer ..> Agent : reads + Tools ..> ToolDef : creates + ToolDef --> "1" ToolType : tool_type, one of these enum values + ToolDef "1" *-- "*" Guardrail : tool guardrails + Guardrail <|-- RegexGuardrail + Guardrail <|-- LlmGuardrail + TerminationCondition <|-- TextMention + TerminationCondition <|-- StopMessage + TerminationCondition <|-- MaxMessage + TerminationCondition <|-- TokenUsage + HandoffCondition <|-- OnToolResult + HandoffCondition <|-- OnTextMention + HandoffCondition <|-- OnCondition +``` + +### Runtime + transport + +```mermaid +classDiagram + direction LR + class AgentRuntime { + +call_sync(agent, prompt) String + +call_async(agent, prompt, &on_done) Execution + +deploy(*agents) List~String~ + +serve(*agents) + +shutdown() + } + class AgentConfig { + +Integer worker_poll_interval_ms + +Integer worker_thread_count + +Boolean auto_register_integrations + +Boolean streaming_enabled + +from_env() AgentConfig$ + } + class Execution { + +String execution_id + +done() Boolean + +waiting() Boolean + +result() String blocks until done + +String partial_text so far, non-blocking + +Symbol finish_reason + +List~ToolCall~ tool_calls + +TokenUsage token_usage + +pause() resume() cancel() + +find(execution_id) Execution$ + } + class ApprovalRequest { + +String execution_id + +String tool_name + +Hash arguments + +approve() + +reject(reason) + } + class ToolRegistry { + +register_tool_workers(tools, agent_name, domain) + +register_system_workers(required_workers) + } + class Dispatch { + +run_tool_task(task, tool_def) Hash + -coerce_args(input_data) + -bind_secrets(task, tool_def) + } + class Secrets { + <> + +secret(name) String + +secrets_env(*names) Hash + } + class SseClient { + +each_event(execution_id) Enumerator + -fallback_polling() + } + class McpDiscovery { + +expand(mcp_tool) List~ToolDef~ + } + class AgentClient { + +start_agent(payload) Hash + +deploy_agent(payload) Hash + +compile_agent(payload) Hash + +get_status(id) Hash + +get_execution(id) Hash + +list_executions(params) Hash + +respond(id, body) + +stop(id) signal(id, msg) + +stream_sse(id, last_event_id) Enumerator + } + class AgentResourceApi { + +start(body) POST agent-start + +deploy(body) POST agent-deploy + +compile(body) POST agent-compile + +status(id) GET agent-id-status + +execution(id) GET agent-execution-id + +executions(params) GET agent-executions + +respond(id, body) POST agent-id-respond + +stop(id) POST agent-id-stop + +signal(id, msg) POST agent-id-signal + +stream(id) GET agent-stream-id SSE + } + class Agent { + <> + } + class ConfigSerializer { + <> + } + class TaskHandler { + <> + } + class WorkflowClient { + <> + } + class ApiClient { + <> + } + class OrkesClients { + <> + +get_agent_client() AgentClient + } + + AgentRuntime *-- AgentConfig + AgentRuntime *-- AgentClient + AgentRuntime *-- ToolRegistry + AgentRuntime *-- SseClient + AgentRuntime ..> ConfigSerializer : agentConfig + AgentRuntime ..> Agent : reads on_approval + AgentRuntime ..> McpDiscovery : uses + AgentRuntime ..> Execution : creates + Execution --> AgentClient : pause, cancel + ApprovalRequest --> AgentClient : respond + SseClient ..> Execution : appends partial_text, fires on_done + SseClient ..> ApprovalRequest : creates on waiting + ToolRegistry --> TaskHandler : Worker.define + ToolRegistry ..> Dispatch : worker body + Dispatch ..> Secrets : binds per task + McpDiscovery ..> WorkflowClient : ephemeral LIST_MCP_TOOLS workflow + AgentClient *-- AgentResourceApi + AgentResourceApi --> ApiClient + OrkesClients ..> AgentClient : creates +``` + +## Examples + +### 1. Tools + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } +end + +agent = Agent.new( + name: 'weather', + model: 'openai/gpt-4o', + instructions: 'Answer weather questions.' +) +agent.add_tool :get_weather + +puts agent.call_sync('Weather in Lisbon?') +``` + +`tool def` marks a method as a tool. Types come from the keyword defaults: `city: String` = required string, `units: 'metric'` = optional with default. Description = humanized method name (`describe :get_weather, '...'` to override). `RubyLLM::Tool` classes work as-is: `agent.add_tool Weather`. Spec: `docs/design/AGENT_TOOLS_DSL.md`. + +```mermaid +sequenceDiagram + actor You + participant SDK as Ruby SDK + participant Server as Conductor Server + participant LLM as OpenAI + + You->>+SDK: agent.call_sync("Weather in Lisbon?") (blocks until done) + SDK->>+Server: POST /agent/start (agentConfig, prompt) + Server-->>-SDK: executionId, requiredWorkers [get_weather] + SDK->>SDK: start worker for get_weather + SDK->>Server: GET /agent/stream (SSE) + Server->>+LLM: instructions + prompt + tool schema + LLM-->>-Server: call get_weather(city: "Lisbon") + Server->>SDK: task get_weather + SDK->>SDK: get_weather(city: "Lisbon") runs + SDK-->>Server: result + Server->>+LLM: tool result + LLM-->>-Server: "Sunny, 21°C in Lisbon" + Server-->>SDK: SSE text + Server-->>SDK: SSE done + SDK-->>-You: "Sunny, 21°C in Lisbon" +``` + +### 2. Streaming + approval + +```ruby +tool def issue_refund(order_id: String, amount: Float) + Billing.refund(order_id, amount) +end +requires_approval :issue_refund + +agent = Agent.new( + name: 'support', + model: 'anthropic/claude-sonnet-4-5', + instructions: 'Help with orders.' +) +agent.add_tool :issue_refund + +agent.on_approval do |request| + request.amount < 100 ? request.approve : request.reject('Needs a manager') +end + +# blocking +answer = agent.call_sync('Refund order A-1029, it arrived broken') + +# non-blocking, callback when finished +agent.call_async('Refund order A-1029, it arrived broken') do |answer| + Mailer.send(customer, answer) +end + +# non-blocking, poll it yourself +execution = agent.call_async('Refund order A-1029, it arrived broken') +execution.done? # false until finished +execution.partial_text # what has streamed so far +execution.result # blocks for the answer +execution.finish_reason # :stop | :rejected +``` + +`call_sync` blocks and returns the answer. `call_async` returns an `Execution` immediately; SSE runs on a background thread, and the block (if given) runs with the answer when done. Spec: `docs/design/AGENT_STREAMING.md`. + +```mermaid +sequenceDiagram + actor You + participant SDK as Ruby SDK + participant Server as Conductor Server + participant LLM as Anthropic + + You->>+SDK: agent.call_async("Refund order A-1029") with on_done block + SDK->>+Server: POST /agent/start + Server-->>-SDK: executionId, requiredWorkers [issue_refund] + SDK->>SDK: start worker for issue_refund + SDK->>Server: GET /agent/stream (SSE, background thread) + SDK-->>-You: Execution (returns immediately) + + Note over You,SDK: Everything below runs on background thread + + Server->>+LLM: prompt + tool schema + LLM-->>-Server: call issue_refund(order_id: "A-1029", amount: 49.0) + Server->>Server: requires approval → pause + + Server-->>+SDK: SSE waiting (tool, arguments) + SDK->>+You: on_approval(request) + You-->>-SDK: request.approve + SDK->>+Server: POST /agent/id/respond approved + Server-->>-SDK: ok + deactivate SDK + + Server->>+SDK: task issue_refund + SDK->>SDK: issue_refund runs + SDK-->>-Server: result + + Server->>+LLM: tool result + LLM-->>-Server: "Refunded $49." + + Server-->>+SDK: SSE text + SDK->>SDK: execution.partial_text appended + deactivate SDK + + Server-->>+SDK: SSE done + SDK->>SDK: execution.done? = true + SDK->>+You: on_done("Refunded $49.") + You-->>-SDK: block returns + deactivate SDK + + You->>+SDK: execution.finish_reason + SDK-->>-You: :stop +``` + +Had `on_approval` called `request.reject`, the server skips the tool and finishes with `finish_reason == :rejected`. Same flow with `call_sync`: the first activation bar simply stays open until done and the answer is the return value. + +### 3. Team + secret + +```ruby +tool def create_issue(title: String, body: '') + Github.create_issue(title, body, token: secret('GH_TOKEN')) +end + +triage = Agent.new( + name: 'triage', + model: 'openai/gpt-4o-mini', + instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.' +) + +filer = Agent.new( + name: 'filer', + model: 'anthropic/claude-sonnet-4-5', + instructions: 'File the bug as a GitHub issue.' +) +filer.add_tool :create_issue + +triage.hands_off_to filer, on: 'ACTIONABLE' + +team = Agent.new(name: 'bug_desk') +team.add_agent triage +team.add_agent filer + +puts team.call_sync(File.read('report.md')) +``` + +Make the agent, then give it things. `secret('GH_TOKEN')` in the tool body both reads the secret and declares it — the SDK scans for it at `tool def` and the server attaches the value to each task. Optional: `filer.redact %w[password api_key]`, `filer.stop_when 'ISSUE_FILED'`, `filer.stop_after messages: 12`, `team.strategy = :sequential`. Every line is sugar over the Python-parity objects (`Handoff::OnTextMention`, `RegexGuardrail`, `Termination::*`). Spec: `docs/design/AGENT_TEAMS.md`. + +```mermaid +sequenceDiagram + actor You + participant SDK as Ruby SDK + participant Server as Conductor Server + participant Triage as triage (gpt-4o-mini) + participant Filer as filer (claude) + + You->>+SDK: team.call_sync(report) (blocks until done) + SDK->>+Server: POST /agent/start (team agentConfig: triage, filer) + Server-->>-SDK: executionId, requiredWorkers [create_issue] + SDK->>SDK: start worker for create_issue, TaskDef.runtime_metadata = [GH_TOKEN] + SDK->>Server: GET /agent/stream (SSE) + Server->>+Triage: report + Triage-->>-Server: "... ACTIONABLE" + Server->>Server: hands_off_to filer matched + Server->>+Filer: conversation so far + create_issue schema + Filer-->>-Server: call create_issue(title, body) + Server->>Server: resolve GH_TOKEN from secret store + Server->>SDK: task create_issue, runtime_metadata GH_TOKEN=ghp_... + SDK->>SDK: secret("GH_TOKEN") → Github.create_issue + SDK-->>Server: issue url + Server->>+Filer: tool result + Filer-->>-Server: "Filed: github.com/.../issues/42" + Server-->>SDK: SSE text + Server-->>SDK: SSE done + SDK-->>-You: "Filed: github.com/.../issues/42" +``` + +## Secrets + +```mermaid +classDiagram + direction LR + class Integration { + <> + holds LLM provider key + model "openai/gpt-4o" → integration "openai" + } + class Agent { + +String model "openai/gpt-4o" + +List~String~ credentials inherited by sub-agents and tools + } + class Tools { + -scan_secrets(method) literal secret() names + } + class ToolDef { + +List~String~ credentials + } + class TaskDef { + +Hash runtime_metadata names only + } + class Task { + +Hash runtime_metadata name → plaintext, wire-only + } + class Dispatch { + -bind_secrets(task, tool_def) + } + class Secrets { + <> + +secret(name) String + +secrets_env(*names) Hash for system, spawn, Open3 + falls back to ENV + } + class CredentialNotFoundError + + Agent ..> Integration : model prefix names it + Tools ..> ToolDef : scan_secrets at tool def + Agent ..> ToolDef : credentials copied down at serialize + ToolDef ..> TaskDef : credentials copied to runtime_metadata at register + TaskDef ..> Task : server fills values at poll + Dispatch ..> Task : reads runtime_metadata + Dispatch ..> Secrets : binds for this call + Secrets ..> CredentialNotFoundError : missing everywhere +``` + +| Secret | Held by | You write | +|---|---|---| +| Conductor auth | your env | `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET` (existing `Configuration`, token cache moves to instance level) | +| LLM provider key | server Integration | `model: 'openai/gpt-4o'` — `openai` is the integration name. SDK never sees the key. | +| Tool credential | server secret store | `secret('GH_TOKEN')` in the tool body — read and declaration in one | + +### Flow: tool credential + +```mermaid +sequenceDiagram + participant SDK as Ruby SDK + participant Server as Conductor Server + participant Store as Secret Store + + Note over SDK,Server: at startup + SDK->>+Server: register TaskDef create_issue, runtimeMetadata [GH_TOKEN] + Server-->>-SDK: ok + + Note over Server: LLM calls create_issue + Server->>+Store: get GH_TOKEN + Store-->>-Server: ghp_... + Server->>+SDK: task create_issue, runtimeMetadata GH_TOKEN=ghp_... + SDK->>SDK: bind runtimeMetadata for this call (fiber-local) + SDK->>SDK: create_issue runs, secret("GH_TOKEN") → ghp_... + SDK-->>-Server: result (runtimeMetadata dropped) +``` + +### Use a secret in a tool + +```ruby +tool def create_issue(title: String, body: '') + Github.create_issue(title, body, token: secret('GH_TOKEN')) +end +``` + +That's it. `secret('GH_TOKEN')` is the read *and* the declaration: at `tool def` the SDK parses the method body, collects every literal `secret('...')`, and puts the names in the tool's contract (`TaskDef.runtimeMetadata`, `tool.config.credentials`). Python makes you write the list; Ruby reads it off the code. Same wire contract. + +### When the name isn't a literal + +```ruby +filer.add_tool :create_issue, credentials: ['GH_TOKEN'] # this tool +filer = Agent.new(..., credentials: ['GH_TOKEN']) # everything under this agent +``` + +### Tool that shells out + +```ruby +tool def gh_create_issue(title: String) + system(secrets_env('GH_TOKEN'), 'gh', 'issue', 'create', '--title', title) +end +``` + +`secrets_env('GH_TOKEN')` is `{ 'GH_TOKEN' => 'ghp_...' }` for this call, and declares the same way. Ruby's `system` / `spawn` / `Open3` take an env hash as the first argument, so only the child process sees it. We never write to `ENV`. + +### No secret store on the server? + +`secret('GH_TOKEN')` falls back to `ENV['GH_TOKEN']`. + +## Frameworks + +No official Ruby SDK exists for any of these. Not supported. + +| Framework | Ruby | +|---|---| +| OpenAI Agents SDK | N/A | +| Anthropic Claude Agent SDK | N/A | +| LangGraph | N/A | +| Google ADK | N/A | diff --git a/docs/design/AGENT_SECRETS.md b/docs/design/AGENT_SECRETS.md new file mode 100644 index 0000000..30b50cc --- /dev/null +++ b/docs/design/AGENT_SECRETS.md @@ -0,0 +1,134 @@ +# Secrets + +Three kinds. Three different owners. None of them live in your code. + +| Secret | Example | Who holds it | You write | +|---|---|---|---| +| Conductor auth | key + secret for the server | your env | `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET` | +| LLM provider key | OpenAI / Anthropic API key | Conductor server, as an Integration | `model: 'openai/gpt-4o'` | +| Tool credential | GitHub token your tool needs | Conductor server, in the secret store | `credentials :create_issue, 'GH_TOKEN'` + `secret('GH_TOKEN')` | + +## 1. Conductor auth + +``` +export CONDUCTOR_SERVER_URL=https://play.orkes.io/api +export CONDUCTOR_AUTH_KEY=... +export CONDUCTOR_AUTH_SECRET=... +``` + +Existing `Configuration`. Nothing new. + +## 2. LLM provider key + +```ruby +agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', instructions: '...') +``` + +`openai` is not a vendor name, it's the **name of an Integration on the server**. Someone +added the OpenAI key there once (UI, or `IntegrationClient#save_integration`). The SDK never +sees the key. Swap `openai/gpt-4o` for `anthropic/claude-sonnet-4-5` and nothing else changes. + +Same as Python. Same as the existing `llm_chat` DSL task (`llmProvider`). + +## 3. Tool credential + +```ruby +tool def create_issue(title: String, body: '') + Github.create_issue(title, body, token: secret('GH_TOKEN')) +end +``` + +That's the whole thing. `secret('GH_TOKEN')` in the body is both the read *and* the +declaration. At `tool def`, the SDK parses the method body, finds every `secret('...')` with a +literal name, and puts those names in the tool's contract (`TaskDef.runtimeMetadata`, +`tool.config.credentials`). The server then attaches the values to each task for that tool. + +Store the value on the server once: `conductor secrets put GH_TOKEN ghp_...`, or the UI, or +`SecretClient#put_secret`. + +When the name isn't a literal, say it when attaching the tool: + +```ruby +filer.add_tool :create_issue, credentials: ['GH_TOKEN'] +``` + +### What happens + +```mermaid +sequenceDiagram + participant SDK as Ruby SDK + participant Server as Conductor Server + participant Store as Secret Store + + Note over SDK: at startup + SDK->>Server: register TaskDef create_issue, runtimeMetadata: [GH_TOKEN] + + Note over Server: LLM calls create_issue + Server->>Store: get GH_TOKEN + Store-->>Server: ghp_... + Server->>+SDK: task create_issue, runtimeMetadata: GH_TOKEN=ghp_... + SDK->>SDK: bind runtimeMetadata for this call (fiber-local) + SDK->>SDK: create_issue runs, secret('GH_TOKEN') → ghp_... + SDK-->>-Server: result, runtimeMetadata dropped +``` + +- Names go up at registration (`TaskDef.runtimeMetadata`). Values come down per task + (`Task.runtimeMetadata`, wire-only, never persisted with the task). +- `secret('X')` reads the map bound for the current call, else `ENV['X']`, else + `CredentialNotFoundError`. + +Same wire contract as Python's `credentials=[...]` + `get_secret()`. Python makes you write +the list; Ruby reads it off the code. + +## The one Ruby difference + +Python also injects secrets into `os.environ` for the duration of the call, so shell-out +tools (`gh`, `aws`) can find them. It can do that because Python workers are separate +processes. + +Ruby workers are threads in one process. Writing `ENV` from a tool would leak secrets into +every other tool running at the same time. So we don't. For subprocesses: + +```ruby +tool def gh_create_issue(title: String) + system(secrets_env('GH_TOKEN'), 'gh', 'issue', 'create', '--title', title) +end +``` + +`secrets_env('GH_TOKEN')` returns `{ 'GH_TOKEN' => 'ghp_...' }` for the current call, and +the literal name is picked up as a declaration the same way `secret()` is. Ruby's `system`, +`spawn`, and `Open3` all accept an env hash as the first argument — the child sees it, nobody +else does. + +## Explicit declaration + +Two cases where the SDK can't read the name off the code: + +```ruby +filer.add_tool :create_issue, credentials: ['GH_TOKEN'] # name is dynamic, or method typed in irb + +filer = Agent.new(..., credentials: ['GH_TOKEN']) # grant to every tool + sub-agent under this agent +``` + +Both add to whatever was detected. Same as Python's `Agent(credentials=[...])`. + +## No secret store on the server? + +`secret('X')` falls back to `ENV['X']`. Always on. + +## Cheat sheet + +```ruby +secret('KEY') # read — and this alone declares it +secrets_env('KEY_A', 'KEY_B') # Hash for system / spawn / Open3 — also declares +agent.add_tool :x, credentials: ['KEY'] # explicit, when the name isn't a literal +Agent.new(..., credentials: ['KEY']) # grant to everything under the agent +# no server value → ENV['KEY'], automatically +``` + +## Decided + +1. `secret('X')` is a bare helper. No `context:` arg. +2. Tool credentials are detected from `secret('LITERAL')` in the tool body at `tool def` (AST scan). Explicit `add_tool ..., credentials:` for dynamic names; agent-level `credentials:` grants to everything under it. +3. `secret('X')` falls back to `ENV['X']` when the server sends nothing. Always on. +4. `Configuration` auth-token cache moves from class level to instance level. Prerequisite. diff --git a/docs/design/AGENT_STREAMING.md b/docs/design/AGENT_STREAMING.md new file mode 100644 index 0000000..0cb8ff1 --- /dev/null +++ b/docs/design/AGENT_STREAMING.md @@ -0,0 +1,153 @@ +# Calling an agent: sync, async, approval + +One idea per step. Each step is a complete program. + +## 1. Sync — call and wait + +```ruby +agent = Agent.new( + name: 'support', + model: 'anthropic/claude-sonnet-4-5', + instructions: 'Help with orders.' +) + +answer = agent.call_sync('What is your return policy?') +puts answer +``` + +`call_sync` blocks until the agent is finished and returns the answer as a String. + +## 2. Async — call and get told when it's done + +```ruby +agent.call_async('What is your return policy?') do |answer| + Mailer.send(customer, answer) +end +``` + +Returns immediately. Your block runs with the answer when the agent finishes. + +## 3. Async — call and check on it yourself + +```ruby +execution = agent.call_async('What is your return policy?') + +execution.done? # false until finished +execution.partial_text # what has streamed in so far +execution.result # blocks until done, returns the answer +execution.cancel # stop it +``` + +Same `call_async`, no block. You get an `Execution` and poll it. + +## 4. Give it a tool + +```ruby +tool def issue_refund(order_id: String, amount: Float) + Billing.refund(order_id, amount) +end + +agent = Agent.new( + name: 'support', + model: 'anthropic/claude-sonnet-4-5', + instructions: 'Help with orders.' +) +agent.add_tool :issue_refund + +puts agent.call_sync('Refund order A-1029, it arrived broken') +``` + +The model calls `issue_refund`, the refund happens, the answer comes back. + +## 5. Make the tool wait for a human + +```ruby +requires_approval :issue_refund +``` + +One line, after the tool. Now the agent stops before running `issue_refund` and waits. + +## 6. Decide + +```ruby +agent.on_approval do |request| + if request.amount < 100 + request.approve + else + request.reject('Needs a manager') + end +end +``` + +`request` has the tool's arguments as methods (`request.amount`, `request.order_id`), plus +`approve` and `reject(reason)`. Works the same with `call_sync` (block runs while you wait) +and `call_async` (block runs on the background thread). + +No `on_approval` registered? The agent waits until something else approves it — a web UI, +another process, `execution.pending.approve`. + +## 7. What happened? + +```ruby +execution = agent.call_async('Refund order A-1029, it arrived broken') +execution.result + +execution.finish_reason # :stop, or :rejected if you rejected the tool +execution.tool_calls # [#] +execution.execution_id # look it up later: Execution.find(id) +``` + +## All together + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def issue_refund(order_id: String, amount: Float) + Billing.refund(order_id, amount) +end +requires_approval :issue_refund + +agent = Agent.new( + name: 'support', + model: 'anthropic/claude-sonnet-4-5', + instructions: 'Help with orders.' +) +agent.add_tool :issue_refund + +agent.on_approval do |request| + request.amount < 100 ? request.approve : request.reject('Needs a manager') +end + +agent.call_async('Refund order A-1029, it arrived broken') do |answer| + Mailer.send(customer, answer) +end +``` + +## Optional + +```ruby +agent.call_sync(question, session_id: 'cust-77') # same conversation across calls + +agent.before_tool_call { |call| log call.name } # hooks +agent.after_tool_result { |result| log result } +``` + +## Fire and forget + +`call_async` with no block, keep the `execution_id`, walk away. One catch: if the agent has +`tool def` tools, they run in *your* process, so it has to stay up. Fire-and-forget only works +when tools are all server-side (http / mcp / human) or you've `deploy`ed the agent and have +workers running elsewhere. Same rule as Python. + +## Under the hood + +`call_async(prompt, &on_done)`: + +1. POST `/agent/start`, register the tool workers the server asks for +2. return an `Execution`; open SSE `/agent/stream/{id}` on a background thread +3. text events append to `execution.partial_text` +4. waiting events build an `ApprovalRequest` and call your `on_approval` block +5. done event sets `done?`, `finish_reason`, `tool_calls`; calls `on_done(answer)` + +`call_sync` = `call_async(prompt).result`. diff --git a/docs/design/AGENT_TEAMS.md b/docs/design/AGENT_TEAMS.md new file mode 100644 index 0000000..535e788 --- /dev/null +++ b/docs/design/AGENT_TEAMS.md @@ -0,0 +1,68 @@ +# Teams + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def create_issue(title: String, body: '') + Github.create_issue(title, body, token: secret('GH_TOKEN')) +end + +triage = Agent.new( + name: 'triage', + model: 'openai/gpt-4o-mini', + instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.' +) + +filer = Agent.new( + name: 'filer', + model: 'anthropic/claude-sonnet-4-5', + instructions: 'File the bug as a GitHub issue.' +) +filer.add_tool :create_issue + +triage.hands_off_to filer, on: 'ACTIONABLE' + +team = Agent.new(name: 'bug_desk') +team.add_agent triage +team.add_agent filer + +puts team.call_sync(File.read('report.md')) +``` + +Make the agent. Then give it things. + +| Line | Does | +|---|---| +| `filer.add_tool :create_issue` | Give filer a tool. | +| `triage.hands_off_to filer, on: 'ACTIONABLE'` | When triage's answer contains ACTIONABLE, filer takes over. | +| `team.add_agent triage` | Give the team a member. First one added starts. `team.add_agents a, b` for several. | +| `secret('GH_TOKEN')` | Reads the secret — and declares it: the SDK sees the literal at `tool def` and tells the server this tool needs GH_TOKEN. Store it once with `conductor secrets put GH_TOKEN ghp_...`. Never in your code or env. | + +`tools:` and `agents:` still work as keyword args in `Agent.new` if you already have the list. +Same object either way. + +## If you need them + +```ruby +filer.redact %w[password api_key] # scrub these from output before anyone sees it +filer.stop_when 'ISSUE_FILED' # stop on this text +filer.stop_after messages: 12 # or after this many messages + +team.strategy = :sequential # one after another (default is handoff) +team.strategy = :parallel # all at once, merged +``` + +## Underneath + +Same Python objects, same `agentConfig`. Sugar only. + +| Sugar | Python-parity object | +|---|---| +| `a.add_tool :x` | appends to `Agent#tools`, same array `tools:` fills | +| `team.add_agent a` | appends to `Agent#agents`, same as `agents:` | +| `a.hands_off_to b, on: 'X'` | `handoffs: [Handoff::OnTextMention.new(target: 'b', text: 'X')]` | +| `a.redact %w[...]` | `guardrails: [RegexGuardrail.new(..., position: :output, on_fail: :fix)]` | +| `a.stop_when 'X'` / `a.stop_after messages: n` | `termination: TextMention \| MaxMessage` | +| `secret('X')` | at `tool def`: AST scan adds `X` to `ToolDef#credentials` → `TaskDef.runtimeMetadata` / `tool.config.credentials`. At run: reads `Task.runtimeMetadata['X']` (fiber-local). Same wire contract as Python `credentials=[...]` + `get_secret` | +| `Agent.new(name: 'bug_desk')` + `add_agent` | `strategy: :handoff` default | diff --git a/docs/design/AGENT_TESTING.md b/docs/design/AGENT_TESTING.md new file mode 100644 index 0000000..c6e78bb --- /dev/null +++ b/docs/design/AGENT_TESTING.md @@ -0,0 +1,109 @@ +# Testing strategy + +How we test the agents SDK. The mocking layer is shared by every Conductor SDK — agentic and +plain-workflow tests alike; Ruby is the first adopter. WireMock records a real server once; +WireMock replays it in CI. No SDK ships a custom mock server. + +## Contract tests — no server + +Serializing a Ruby agent must produce exactly what the server (and Python) expect: + +```ruby +expect(JSONSchemer.schema(AGENT_SCHEMA).valid?(config)).to be true +expect(ConfigSerializer.serialize(agent)).to eq JSON.parse(fixture('05_handoffs.json')) +``` + +Fixtures vendored from python-sdk (`agent-schema.json`, the 19 golden `_configs/*.json`), +plus unit specs for schema generation, secret scanning, and serialization. This is also what +makes shared mocks work: replay matches by verb + path + order, so one recording serves every +SDK precisely because the SDKs send equivalent requests. + +## Runtime tests — record once, replay everywhere + +To create or refresh a recording: + +1. Run a Conductor server locally (docker, `localhost:8080`) with real provider keys + configured as integrations. +2. Run your SDK's test suite as usual, with the two record env vars set. Ruby example: + + ``` + CONDUCTOR_RECORD=1 \ + CONDUCTOR_RECORD_URL=http://localhost:8080 \ + bundle exec rspec spec/agents + ``` + + `CONDUCTOR_RECORD=1` switches the test helper into record mode: it boots WireMock between + the SDK and your server. `CONDUCTOR_RECORD_URL` says where your server is (shown value is + the default). +3. The test executes for real — the server calls the actual LLM, your tool methods run. +4. On green, WireMock has saved every response your server sent. A cleanup script swaps + run-specific values (execution ids, timestamps) for placeholders so the recording replays + for anyone. +5. The cleaned files are a scenario folder — open a PR to + [conductor-mocks](https://github.com/conductor-oss/conductor-mocks) with it. + +``` +mocks/ + agent/tool_happy_path/ + agent/approval_approve/ + agent/approval_reject/ + agent/secrets_runtime_metadata/ + agent/team_handoff/ + agent/mcp_discovery/ + workflow/simple_task/ + workflow/dynamic_fork/ +``` + +Re-recording is manual, and happens when we bump the supported server version — a re-record +is a reviewable diff. + +``` +record spec <-> WireMock proxy <-> your conductor-oss <-> REAL model + \-> normalized mappings -> conductor-mocks + +replay spec <-> wiremock/wiremock container + scenario mappings + no conductor, no LLM, no keys +``` + +In CI, the replayer is WireMock itself. Scenario states keep responses in recorded order; +`POST /tasks` only matches the stub with the recorded body, so a wrong tool result matches +nothing; any unmatched request fails the run. WireMock buffers SSE — fine, because an agent +stream terminates at `done`, so the captured body is complete. Recorded model text is that +day's output, so tests assert structure (`finish_reason`, tool args), not prose. + +### The weather test (Ruby) + +```ruby +RSpec.describe 'weather agent', mocks: 'agent/tool_happy_path' do + tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } + end + + it 'answers with the tool' do + agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', + instructions: 'Answer weather questions.') + agent.add_tool :get_weather + expect(agent.call_sync('Weather in Lisbon?')).to include('Lisbon') + end +end +``` + +The `mocks:` tag boots WireMock with that scenario and points the SDK at it. + +## CI wiring + +Each SDK's workflow checks out conductor-mocks and starts the official image — every push, +fork PRs included: + +```yaml +- uses: actions/checkout@v4 + with: { repository: conductor-oss/conductor-mocks, path: conductor-mocks } +- run: docker run -d -p 8080:8080 + -v $PWD/conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock + wiremock/wiremock:3x +- run: CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/agents +``` + +conductor-mocks PRs are reviewed by the conductor-oss team; nothing replay serves was +invented by hand, and no SDK internals are ever stubbed. (Housekeeping: drop the unused +vcr/webmock dev deps from this repo.) diff --git a/docs/design/AGENT_TOOLS_DSL.md b/docs/design/AGENT_TOOLS_DSL.md new file mode 100644 index 0000000..abeff6b --- /dev/null +++ b/docs/design/AGENT_TOOLS_DSL.md @@ -0,0 +1,106 @@ +# Tools + +## Whole thing, one file + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } +end + +agent = Agent.new( + name: 'weather', + model: 'openai/gpt-4o', + instructions: 'Answer weather questions.' +) +agent.add_tool :get_weather + +puts agent.call_sync('Weather in Lisbon?') +``` + +Run it: `ruby weather.rb`. That's the whole thing. + +## What each line does + +| Line | Does | +|---|---| +| `tool def get_weather(...)` | Makes the method a tool. `tool` sees the name `:get_weather` because `def` returns it (same as `private def`). | +| `city: String` | Required string. | +| `units: 'metric'` | Optional string, default `metric`. | +| `agent.add_tool :get_weather` | Give the agent the tool by name. | +| `agent.call_sync('prompt')` | Run it, wait, return the answer. Server URL / auth come from `CONDUCTOR_*` env vars. | + +The description the LLM sees is the method name: `get_weather` → "Get weather". + +## Types + +Whatever you put as the default is the type: + +| You write | Meaning | +|---|---| +| `city: String` | required string | +| `amount: Float` | required number | +| `count: Integer` | required integer | +| `tags: [String]` | required list of strings | +| `units: 'metric'` | optional, default `"metric"` | +| `limit: 10` | optional, default `10` | +| `verbose: false` | optional, default `false` | +| `units: %w[metric imperial]` | optional, one of these | + +## More than a few tools? Put them in a module + +```ruby +module Weather + extend Conductor::Agents::Tools + + tool def current(city: String) ... end + tool def forecast(city: String, days: 3) ... end +end + +agent.add_tools Weather # both +agent.add_tool Weather[:current] # one +``` + +## Options + +Approval and secrets are one line each, after the tool: + +```ruby +tool def issue_refund(order_id: String, amount: Float) + Stripe.refund(order_id, amount, key: secret('STRIPE_KEY')) +end +requires_approval :issue_refund +``` + +- `requires_approval :name` — a human approves before it runs. See `AGENT_STREAMING.md`. +- `secret('KEY')` — reads a secret. Writing it in the body is also the declaration: the SDK + finds it at `tool def` and tells the server this tool needs `STRIPE_KEY`. See `AGENT_SECRETS.md`. +- `describe :name, '...'` — override the description if the method name isn't enough. + +## Rules + +- Keyword args only (`city:`). Positional args raise at load. +- It's still a normal method: `get_weather(city: 'Lisbon')` works in tests. + +## Already have RubyLLM tools? + +They work as-is: + +```ruby +class Weather < RubyLLM::Tool + desc "Gets current weather for a location" + def execute(latitude:, longitude:) ... end +end + +agent.add_tool Weather # a RubyLLM::Tool class, as-is +``` + +RubyLLM already builds the JSON schema from `execute`'s keyword args and exposes `name` and +`description`; we read those into a `ToolDef` and run `Weather.new.execute(**args)` as the +worker body. Optional dependency — only loaded if `RubyLLM` is defined. + +RubyLLM is not the engine underneath. It runs the LLM loop client-side with your provider key; +Conductor runs it server-side with the key held as an integration. Only the tool class and the +`call_sync` / `call_async` API shape are shared. diff --git a/docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md b/docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md new file mode 100644 index 0000000..3dcc014 --- /dev/null +++ b/docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md @@ -0,0 +1,554 @@ +# SDK Agents Testing Strategy V2 + +> Status: local working draft for rapid iteration. This supersedes the WireMock assumptions in +> the current Confluence draft but does not update Confluence yet. + +## Decision + +SDK agent tests run against a real Conductor server and a deterministic `mockLLM` provider built +directly into the Conductor OSS test runtime. + +There is no WireMock, HTTP recording proxy, replay server, or separate request normalizer. Conductor +OSS owns the small mock provider implementation and its canonical JSON fixtures under +`src/test/resources/llm-mocks/`. Each SDK owns its agent definitions, workers, and assertions. + +The server supports two modes over that one fixture format: + +- **Replay mode** is the default for SDK development and CI. Agents use `mockLLM/`, and + `mockLLM` validates the provider-neutral request before returning the fixture's response. +- **Record mode** is an explicit local authoring mode. A JVM system property names the scenario, + and Conductor records the provider-neutral requests and responses from any real integration other + than `mockLLM` into the canonical fixture. + +Local development and CI use the same path: + +1. Start `conductor-server` from a pinned Conductor OSS revision with the test runtime enabled. +2. Register the no-secret `mockLLM` integration provider. +3. Register an agent whose model points to a named `mockLLM` scenario. +4. Run the agent through the real Conductor APIs and real task queues. +5. Assert the prompt, tool calls, tool results, guardrail behavior, and final execution result. + +Record mode is not part of the normal test run. It exists to create or refresh the single expected +fixture for a scenario, after which that fixture is replayed everywhere through `mockLLM`. + +## Goals + +- Exercise each SDK against the real Conductor server, agent compiler, execution engine, task + queues, and persistence layer. +- Make agent tests deterministic, fast, credential-free, and runnable on fork pull requests. +- Verify the exact SDK/server contract: serialized definitions, prompt wording, tool schemas and + arguments, tool results, guardrail outcomes, and final output. +- Keep the mocking implementation intentionally small and colocated with the server behavior it + replaces. +- Run identical scenarios locally and in CI. +- Make expected model wording, tool calls, and follow-up turns easy to create from a real provider + when a scenario is first authored or intentionally refreshed. + +## Non-goals + +- Testing OpenAI, Anthropic, or another provider's wire protocol. +- Recording or replaying raw provider HTTP traffic. +- Maintaining separate OpenAI, Anthropic, model, or provider-specific mock sets. +- Running record mode in CI. +- Verifying the natural-language quality of a real model response. +- Mocking Conductor's HTTP API, task queue, SSE implementation, or execution engine. +- Using a separate `conductor-mocks` repository for functional-test infrastructure. + +## Test layers + +### 1. Contract tests: SDK only, no server + +Contract tests remain ordinary unit tests. They catch SDK serialization and schema regressions +quickly without booting Conductor. + +Each SDK should assert that its agent configuration: + +- conforms to the published agent schema; +- serializes into the expected Conductor agent/workflow definition; +- generates the expected tool input schema; +- includes guardrails, secrets, handoffs, and other agent settings in the correct shape. + +Ruby example: + +```ruby +expect(JSONSchemer.schema(AGENT_SCHEMA).valid?(config)).to be true +expect(ConfigSerializer.serialize(agent)).to eq JSON.parse(fixture('weather_agent.json')) +``` + +These tests answer: **did the SDK build the correct definition?** + +### 2. Functional tests: SDK plus real Conductor plus `mockLLM` + +Functional tests boot a real server and exercise an agent end to end. Only the external model is +replaced. The `mockLLM` provider returns deterministic chat responses from a named fixture while all +orchestration remains real. + +These tests answer: **does this SDK-defined agent behave correctly when Conductor executes it?** + +### 3. Fixture authoring: real provider through Conductor record mode + +Record mode runs the same server and SDK scenario with a real, non-`mockLLM` integration. Conductor +captures the request and response after provider-specific translation has been removed from the +equation and writes the canonical fixture consumed by `mockLLM`. + +This is a developer workflow, not a test topology or CI dependency. A recorded fixture must be +reviewed and then replayed with `mockLLM` before it is committed. + +## Ownership and layout + +### Conductor OSS + +Conductor OSS owns the provider and the model responses because prompt construction, tool-call +translation, and guardrail orchestration happen in the server. + +Proposed layout (the exact module prefix may change during implementation): + +```text +conductor-oss/ +└── ai/ + └── src/ + └── test/ + ├── java/.../MockLLM.java + ├── java/.../MockLLMRecorder.java + └── resources/ + └── llm-mocks/ + ├── weather_tool_call.json + ├── input_guardrail_block.json + ├── output_guardrail_fix.json + └── team_handoff.json +``` + +`MockLLM` is a normal Conductor AI model implementation available only in the test runtime. It: + +- advertises the provider name `mockLLM`; +- uses the requested model name as the fixture name; +- loads that fixture from `src/test/resources/llm-mocks`; +- finds the single fixture turn whose `expect` block matches the provider-neutral request; +- returns the response through the same Conductor AI model interface as a real provider; and +- performs no network I/O and requires no credentials. + +Replay uses request content rather than a global call counter. That keeps concurrent SDK scenarios +isolated and makes a prompt, tool schema, tool call, or tool-result mismatch fail at the relevant +request rather than shifting every later response. + +A minimal fixture shape could be: + +```json +{ + "schemaVersion": 1, + "scenario": "weather_tool_call", + "turns": [ + { + "expect": { + "messages": [ + { + "role": "system", + "content": "Answer weather questions using the weather tool." + }, + { "role": "user", "content": "Weather in Lisbon?" } + ], + "tools": [ + { + "name": "get_weather", + "description": "Get the weather for a city" + } + ] + }, + "respond": { + "toolCall": { + "id": "call_weather_1", + "name": "get_weather", + "arguments": { "city": "Lisbon", "units": "metric" } + } + } + }, + { + "expect": { + "lastMessage": { + "role": "tool", + "name": "get_weather", + "content": { "temp_c": 21.0, "summary": "Sunny in Lisbon" } + } + }, + "respond": { "text": "Sunny in Lisbon, 21C." } + } + ] +} +``` + +The fixture is a provider-neutral request/response script, not a capture of an HTTP exchange. It +contains only semantic fields required to validate and drive the Conductor behavior under test. +Provider request envelopes, URLs, headers, authentication, completion IDs, timestamps, token usage, +and provider/model names are never written. + +### Record mode in Conductor + +Record mode is enabled when Conductor starts. The proposed JVM properties are: + +```text +-Dconductor.ai.mock-llm.record= +-Dconductor.ai.mock-llm.output-dir= +``` + +The first property enables recording and supplies the canonical scenario name. The output directory +may default to the Conductor OSS test-resource location for a source checkout, but CI and scripts +should pass it explicitly so the destination is unambiguous. + +When record mode is enabled, Conductor decorates its normal AI model invocation at the +provider-neutral boundary: + +1. The agent must reference a real integration such as OpenAI or Anthropic. Selecting `mockLLM` + fails immediately because recording a mock would create a mock of a mock. +2. Conductor builds the normal prompt, messages, tool schemas, and guardrail request. +3. The recorder forwards the call to the selected real provider. +4. The recorder retains only the stable `expect` request fields and semantic `respond` fields. +5. Tool calls execute normally. Their results naturally appear in the next recorded request. +6. Each completed turn is written atomically to `/.json`. + +Recording OpenAI and recording Anthropic both write exactly the same file shape and path. The chosen +provider is merely how a developer generates the desired wording and tool calls. It is not part of +fixture identity, and there is only one supported fixture set. + +Record one scenario at a time. Re-recording intentionally replaces that scenario's canonical file; +the resulting source diff is the review surface. + +### Each SDK + +Each SDK owns: + +- the test agent definition written through that SDK's public API; +- any SDK-hosted tool implementation or worker; +- the input used to start the agent; +- assertions over the registered definition and completed execution; and +- a small test helper that starts or connects to the Conductor test server and registers the + `mockLLM` integration. + +SDK tests do not start a mock HTTP service and do not understand the mock fixture internals. A test +selects a scenario only by using a model such as `mockLLM/weather_tool_call`. + +## Replay mode: one execution model for local development and CI + +```mermaid +sequenceDiagram + actor R as Local developer or SDK CI + participant C as conductor-server + participant M as mockLLM (in-process) + participant S as SDK test + tool workers + + R->>C: Start server with test runtime + C->>C: Load src/test/resources/llm-mocks + R->>C: Register mockLLM integration (no secret) + R->>S: Run SDK agent suite + + S->>C: Register agent using mockLLM/weather_tool_call + S->>C: Start agent with test input + C->>M: Chat request with system prompt, messages, and tool schemas + M-->>C: Scripted get_weather(city: Lisbon) tool call + C-->>S: Queue get_weather task + S->>S: Execute the real SDK tool implementation + S->>C: Complete task with tool result + C->>M: Chat request containing the tool result + M-->>C: Scripted final answer + C-->>S: Complete agent execution + + S->>C: Read execution and task details + S->>S: Assert prompt wording, tool call, tool result, guardrails, and final output + S-->>R: Pass or fail +``` + +This is the only functional-test topology. SSE and polling may both be exercised by SDK tests, but +they are merely two clients of the same real execution; neither needs a separate mock strategy or +sequence diagram. + +## Record mode: local fixture authoring + +```mermaid +sequenceDiagram + actor D as Developer + participant S as SDK scenario + tool workers + participant C as conductor-server + participant R as In-process recorder + participant L as Real non-mockLLM provider + + D->>C: Start with -Dconductor.ai.mock-llm.record=weather_tool_call + D->>S: Run one SDK scenario using OpenAI or Anthropic integration + S->>C: Register and start the agent + C->>R: Provider-neutral prompt, messages, and tool schemas + R->>L: Invoke the configured real provider + L-->>R: Generated tool call + R->>R: Save canonical expect + respond turn + R-->>C: Return the real provider response + C-->>S: Queue the generated tool call + S->>C: Complete the real tool with its result + C->>R: Next request containing that tool result + R->>L: Invoke the same configured provider + L-->>R: Generated final text + R->>R: Save canonical expect + respond turn atomically + R-->>C: Return final text + C-->>S: Complete agent execution + S-->>D: Scenario result; review the fixture diff + + Note over D,L: Provider-specific HTTP is never recorded.
OpenAI and Anthropic produce the same single fixture format. +``` + +After recording, restart Conductor without the record property, change the test agent back to +`mockLLM/weather_tool_call`, and run the same scenario in replay mode. The fixture is ready only when +that deterministic replay and the SDK assertions pass. + +## Guardrail execution + +Guardrails also run in the real server. A scenario supplies deterministic model output only when a +guardrail or the main agent needs an LLM response. + +```mermaid +sequenceDiagram + participant S as SDK test + participant C as conductor-server + participant M as mockLLM (in-process) + + S->>C: Register agent + guardrail using mockLLM scenario + S->>C: Start agent with controlled input + C->>C: Execute the real guardrail path + opt Guardrail requires an LLM decision + C->>M: Guardrail prompt and controlled content + M-->>C: Scripted allow, block, or fix decision + end + C->>C: Continue, stop, or rewrite according to guardrail result + C-->>S: Completed or blocked execution + S->>C: Read execution and guardrail task details + S->>S: Assert prompt text, decision, action, status, and visible output +``` + +For a blocked input guardrail, the test should also assert that the main agent LLM task and tool task +were never scheduled. For a fixing guardrail, it should assert both the original guardrail decision +and the exact rewritten text passed to the next stage. + +## What functional tests assert + +Assertions come from Conductor's persisted execution and task data plus observations made by the +SDK tool worker. They must not depend only on the final answer. + +For each scenario, assert the applicable items: + +1. **Registration** + - the agent definition registered successfully; + - the agent references the expected `mockLLM` integration and scenario; + - generated tool schemas and guardrail configuration match the test. +2. **Prompt construction** + - system/developer instructions contain the required exact wording; + - user messages and relevant conversation history appear in the expected order; + - tool descriptions and schemas exposed to the model are correct; + - runtime-only values are asserted structurally or by stable substring, not as fixed IDs. +3. **Tool execution** + - the expected tool is called exactly once unless the scenario says otherwise; + - arguments match exactly; + - the SDK worker returns the expected result; + - that result is present in the next LLM task input. +4. **Guardrails** + - the expected input/output/tool guardrail runs; + - its decision and configured action are correct; + - blocked paths do not execute forbidden downstream work; + - fixed content, rejection reasons, and final statuses are preserved. +5. **Completion** + - the execution reaches the expected terminal status; + - final text and finish reason match the deterministic fixture; + - no unexpected tool or LLM tasks were scheduled. + +Ruby-style example: + +```ruby +agent = Agent.new( + name: 'weather', + model: 'mockLLM/weather_tool_call', + instructions: 'Answer weather questions using the weather tool.' +) +agent.add_tool :get_weather + +execution = agent.call_async('Weather in Lisbon?') +result = execution.result +details = conductor.workflow_client.get_workflow(execution.execution_id, true) + +expect(recorded_tool_calls).to contain_exactly({ + name: 'get_weather', + arguments: { 'city' => 'Lisbon', 'units' => 'metric' } +}) +expect(llm_task(details).input_data.fetch('messages').to_json) + .to include('Answer weather questions using the weather tool.') +expect(result).to eq('Sunny in Lisbon, 21C.') +``` + +The final helper names will follow each SDK's existing APIs. The important contract is that tests +inspect the real execution rather than a mock server's request log. + +## Initial shared scenario catalog + +All SDKs should implement the same small behavioral matrix. The named model fixture in Conductor OSS +drives the response, while each SDK expresses and verifies the scenario through its own public API. + +| Scenario | Main behavior under test | Required assertions | +|---|---|---| +| `direct_answer` | Agent completes without a tool | Prompt text, final text, no tool tasks | +| `weather_tool_call` | One SDK tool call followed by an answer | Tool schema, exact args, tool result in next prompt, final text | +| `input_guardrail_block` | Input guardrail prevents execution | Guardrail prompt/result, blocked status, no main LLM/tool task | +| `output_guardrail_fix` | Output guardrail rewrites model text | Original output, fix decision, corrected final output | +| `approval_reject` | Tool call waits for and receives rejection | Pending approval, rejection reason, tool body not run, final status | +| `team_handoff` | One agent hands execution to another | Handoff target, prompts for both agents, final owner and output | + +Provider-specific protocol cases do not belong in this suite. Those remain Conductor provider adapter +tests. + +## Local workflow + +Provide one script or Gradle task in Conductor OSS that starts the test server with `mockLLM` on the +classpath. The SDK test helper should support either starting that command or connecting to an +already-running server. + +### Normal replay + +```bash +# terminal 1: conductor-oss +./gradlew + +# terminal 2: ruby-sdk +CONDUCTOR_SERVER_URL=http://localhost:8080/api \ +bundle exec rspec spec/agents +``` + +No LLM API key or mock-service URL is set. The SDK helper registers `mockLLM` idempotently before the +suite and waits for the Conductor health endpoint before running tests. + +### Create or refresh a fixture + +Start the built server with the record properties and configure one real provider integration in +the normal Conductor way: + +```bash +java \ + -Dconductor.ai.mock-llm.record=weather_tool_call \ + -Dconductor.ai.mock-llm.output-dir=/path/to/conductor/ai/src/test/resources/llm-mocks \ + -jar conductor-server.jar +``` + +Then run only the corresponding SDK scenario with its agent temporarily configured for a real +integration: + +```bash +CONDUCTOR_SERVER_URL=http://localhost:8080/api \ +bundle exec rspec spec/agents/weather_tool_call_spec.rb +``` + +Review the resulting `weather_tool_call.json`, stop the recording server, and rerun the scenario +against `mockLLM/weather_tool_call`. The recording invocation requires credentials only for the real +integration chosen by the developer. + +## CI wiring + +Each SDK agent job performs the same operations: + +1. Check out the SDK and a pinned Conductor OSS revision. +2. Build and start `conductor-server` with the test runtime and `mockLLM` fixtures. +3. Wait until the server is healthy. +4. Register the `mockLLM` integration. +5. Install the SDK toolchain and run its agent tests against the server. +6. Upload server logs and execution details on failure. + +Record mode must not be enabled in CI. CI consumes the committed canonical fixtures and never +configures a real LLM integration or provider credential. + +Illustrative workflow: + +```yaml +jobs: + agents: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: actions/checkout@v4 + with: + repository: conductor-oss/conductor + ref: + path: conductor-oss + - run: conductor-oss/gradlew + - run: bundle install + - run: bundle exec rspec spec/agents + env: + CONDUCTOR_SERVER_URL: http://localhost:8080/api +``` + +The final server-start step should run in the background, include a health wait, and guarantee cleanup. +It is shown compactly here because the exact Gradle task is part of the Conductor OSS implementation. + +## Failure behavior + +Failures should be direct and local: + +- an unknown model/scenario fails with the missing fixture path; +- no fixture turn matching the current request fails with a concise summary of message roles and + available selectors; +- malformed fixture JSON fails server startup or the first scenario use; +- enabling record mode while selecting `mockLLM` fails with an explicit unsupported-operation + message; +- record mode refuses to mix multiple scenario names into one output file; +- an SDK assertion shows the actual persisted prompt, tool call, guardrail result, or final output; +- an unexpected external provider selection fails because CI has no credentials and only `mockLLM` + is registered for these tests. + +## Removed from the previous design + +The following components and concepts are deleted: + +- WireMock and its Gradle record/replay modes; +- recorded HTTP mappings and `__files` payloads; +- post-processing raw provider traffic through a request normalizer; +- WireMock scenario state, request logs, and unmatched-request verification; +- `CONDUCTOR_RECORD` and `CONDUCTOR_RECORD_URL`; +- provider/model-specific recording directories; +- a reusable workflow whose primary purpose is to run WireMock; and +- the separate recording lifecycle in `conductor-mocks`. + +The replacement is one in-process replay provider, one in-process provider-neutral recorder, a +single set of deterministic JSON fixtures, and ordinary assertions against a real Conductor +execution. + +## Implementation slices + +1. **Conductor OSS test provider** + - implement `MockLLM` through the existing AI model interface; + - load named fixtures from `src/test/resources/llm-mocks`; + - expose it only in the test server profile/classpath; + - add provider unit tests for text, tool-call, missing-fixture, and concurrent scenario behavior. +2. **Conductor OSS record mode** + - add the `conductor.ai.mock-llm.record` and output-directory JVM properties; + - decorate non-`mockLLM` model calls at the provider-neutral boundary; + - atomically emit the canonical request/response fixture without provider metadata; + - reject `mockLLM` as a recording source; + - test that OpenAI-shaped and Anthropic-shaped adapters produce the same fixture schema. +3. **Conductor OSS test-server entry point** + - add a documented Gradle/script entry point; + - make no-secret `mockLLM` integration registration automatic or idempotent; + - expose a reliable health check. +4. **Ruby SDK first adopter** + - add the server-backed agent spec helper; + - implement the initial scenario catalog; + - assert persisted prompts, tool calls, guardrails, and results; + - add the agent job to CI. +5. **Other SDKs** + - reuse the same fixture names and behavioral assertions; + - implement only the language-specific agent definition, tool worker, and test adapter. + +## Acceptance criteria + +- The full Ruby agent suite passes locally with no provider credentials. +- The same suite passes in CI against a freshly started real Conductor server. +- No WireMock process, dependency, configuration, or mock repository is involved. +- Starting Conductor with `-Dconductor.ai.mock-llm.record=` and a real provider creates or + refreshes that scenario's canonical fixture. +- Recording the same scenario through OpenAI or Anthropic produces the same provider-neutral schema + and the same output path; only one version may be committed. +- Record mode rejects `mockLLM`, and CI never enables record mode. +- At least one test proves a real SDK tool is invoked and its result reaches the next LLM turn. +- At least one test proves a guardrail blocks or fixes content and the forbidden path does not run. +- At least one test asserts stable prompt wording from the persisted LLM task input. +- Tests fail clearly when the SDK changes a prompt, tool schema/argument, tool result, or guardrail + configuration unexpectedly. +- The test server makes zero outbound LLM network calls. From 708d6fd76a7b6e8dfd4572b1a7773f1865a67861 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Tue, 8 Sep 2026 15:20:17 -0700 Subject: [PATCH 02/20] Phase 0: agent transport, runtimeMetadata, instance token cache, update-v2 - Configuration caches the auth token per instance; class-level accessors stay as a deprecated shim that warns once - Task#runtime_metadata (wire-only secret values) and TaskDef#runtime_metadata (declared secret names) plus TaskDef#enforce_schema - TaskResourceApi#update_task_v2; TaskRunner prefers POST /tasks/update-v2 and falls back to POST /tasks once on 404/405 (Python parity); Worker gains the lease_extend_enabled option - AgentResourceApi + AgentClient for /api/agent/* (start, deploy, compile, status, execution, executions, respond/approve/reject, stop, signal, pause, resume, cancel, list/get/delete) and OrkesClients#get_agent_client - AgentApiError / AgentNotFoundError carrying the server's {error, status} body - Dev deps: drop unused vcr, add json_schemer for agent contract tests Co-Authored-By: Claude Fable 5.1 --- Gemfile.lock | 10 +- conductor_ruby.gemspec | 2 +- lib/conductor.rb | 2 + lib/conductor/client/agent_client.rb | 110 ++++++++++++++ lib/conductor/client/task_client.rb | 7 + lib/conductor/configuration.rb | 53 +++++-- lib/conductor/exceptions.rb | 36 +++++ lib/conductor/http/api/agent_resource_api.rb | 138 ++++++++++++++++++ lib/conductor/http/api/task_resource_api.rb | 15 ++ lib/conductor/http/models/task.rb | 12 +- lib/conductor/http/models/task_def.rb | 17 ++- lib/conductor/orkes/orkes_clients.rb | 4 + lib/conductor/worker/task_runner.rb | 22 ++- lib/conductor/worker/worker.rb | 5 +- lib/conductor/worker/worker_config.rb | 3 +- spec/conductor/client/agent_client_spec.rb | 69 +++++++++ .../configuration/token_cache_spec.rb | 29 ++++ .../http/api/agent_resource_api_spec.rb | 96 ++++++++++++ .../conductor/models/runtime_metadata_spec.rb | 53 +++++++ spec/conductor/orkes/orkes_clients_spec.rb | 6 + spec/conductor/worker/task_runner_spec.rb | 5 +- spec/conductor/worker/task_update_v2_spec.rb | 41 ++++++ 22 files changed, 711 insertions(+), 24 deletions(-) create mode 100644 lib/conductor/client/agent_client.rb create mode 100644 lib/conductor/http/api/agent_resource_api.rb create mode 100644 spec/conductor/client/agent_client_spec.rb create mode 100644 spec/conductor/configuration/token_cache_spec.rb create mode 100644 spec/conductor/http/api/agent_resource_api_spec.rb create mode 100644 spec/conductor/models/runtime_metadata_spec.rb create mode 100644 spec/conductor/worker/task_update_v2_spec.rb diff --git a/Gemfile.lock b/Gemfile.lock index 97fbb26..29cb5b7 100644 --- a/Gemfile.lock +++ b/Gemfile.lock @@ -47,10 +47,16 @@ GEM fiber-local (1.1.0) fiber-storage fiber-storage (1.0.1) + hana (1.3.7) hashdiff (1.2.1) io-console (0.8.2) io-event (1.16.0) json (2.7.6) + json_schemer (2.5.0) + bigdecimal + hana (~> 1.3) + regexp_parser (~> 2.0) + simpleidn (~> 0.2) method_source (1.1.0) metrics (0.15.0) net-http-persistent (4.0.8) @@ -105,9 +111,9 @@ GEM rubocop-capybara (~> 2.17) ruby-progressbar (1.13.0) ruby2_keywords (0.0.5) + simpleidn (0.3.0) traces (0.18.2) unicode-display_width (2.6.0) - vcr (6.1.0) webmock (3.26.1) addressable (>= 2.8.0) crack (>= 0.3.2) @@ -120,13 +126,13 @@ PLATFORMS DEPENDENCIES async (~> 2.0) conductor_ruby! + json_schemer (~> 2.0) prometheus-client (~> 4.0) pry (~> 0.14) rake (~> 13.0) rspec (~> 3.0) rubocop (~> 1.0) rubocop-rspec (~> 2.0) - vcr (~> 6.0) webmock (~> 3.0) webrick (~> 1.8) diff --git a/conductor_ruby.gemspec b/conductor_ruby.gemspec index 632163a..e828d15 100644 --- a/conductor_ruby.gemspec +++ b/conductor_ruby.gemspec @@ -41,11 +41,11 @@ Gem::Specification.new do |spec| spec.add_dependency 'json', '>= 2.0' # Development dependencies (alphabetically sorted) + spec.add_development_dependency 'json_schemer', '~> 2.0' spec.add_development_dependency 'pry', '~> 0.14' spec.add_development_dependency 'rake', '~> 13.0' spec.add_development_dependency 'rspec', '~> 3.0' spec.add_development_dependency 'rubocop', '~> 1.0' spec.add_development_dependency 'rubocop-rspec', '~> 2.0' - spec.add_development_dependency 'vcr', '~> 6.0' spec.add_development_dependency 'webmock', '~> 3.0' end diff --git a/lib/conductor.rb b/lib/conductor.rb index 3c8bf3e..4073f36 100644 --- a/lib/conductor.rb +++ b/lib/conductor.rb @@ -69,6 +69,7 @@ require_relative 'conductor/http/api/schema_resource_api' require_relative 'conductor/http/api/integration_resource_api' require_relative 'conductor/http/api/prompt_resource_api' +require_relative 'conductor/http/api/agent_resource_api' # OSS Clients require_relative 'conductor/client/workflow_client' require_relative 'conductor/client/task_client' @@ -80,6 +81,7 @@ require_relative 'conductor/client/schema_client' require_relative 'conductor/client/integration_client' require_relative 'conductor/client/prompt_client' +require_relative 'conductor/client/agent_client' # Orkes-specific models require_relative 'conductor/orkes/models/metadata_tag' require_relative 'conductor/orkes/models/rate_limit_tag' diff --git a/lib/conductor/client/agent_client.rb b/lib/conductor/client/agent_client.rb new file mode 100644 index 0000000..1e5cc66 --- /dev/null +++ b/lib/conductor/client/agent_client.rb @@ -0,0 +1,110 @@ +# frozen_string_literal: true + +require_relative '../exceptions' +require_relative '../http/api/agent_resource_api' + +module Conductor + module Client + # AgentClient - High-level client for the server-side agent runtime. + # + # Mirrors the Python SDK's OrkesAgentClient: hashes in, hashes out, and every + # transport error is re-raised as AgentApiError (AgentNotFoundError on 404). + class AgentClient + attr_reader :agent_api + + # @param api_client [Http::ApiClient] + def initialize(api_client) + @agent_api = Http::Api::AgentResourceApi.new(api_client) + end + + # @param payload [Hash] AgentStartRequest + # @return [Hash] { "executionId", "agentName", "requiredWorkers" } + def start_agent(payload) + wrap { @agent_api.start(payload) } + end + + # @return [Hash] { "agentName", "requiredWorkers" } + def deploy_agent(payload) + wrap { @agent_api.deploy(payload) } + end + + # @return [Hash] { "workflowDef", "requiredWorkers" } + def compile_agent(payload) + wrap { @agent_api.compile(payload) } + end + + def get_status(execution_id) + wrap { @agent_api.status(execution_id) } + end + + def get_execution(execution_id) + wrap { @agent_api.execution(execution_id) } + end + + def list_executions(params = {}) + wrap { @agent_api.executions(params) } + end + + # Respond to a waiting execution. Hashes pass through; anything else is wrapped as + # { "output" => value } like the Python client does. + def respond(execution_id, body) + payload = body.is_a?(Hash) ? body : { 'output' => body } + wrap { @agent_api.respond(execution_id, payload) } + end + + def approve(execution_id) + respond(execution_id, { 'approved' => true }) + end + + def reject(execution_id, reason = '') + respond(execution_id, { 'approved' => false, 'reason' => reason.to_s }) + end + + def send_message(execution_id, message) + respond(execution_id, { 'message' => message.to_s }) + end + + def stop(execution_id) + wrap { @agent_api.stop(execution_id) } + end + + def signal(execution_id, message) + wrap { @agent_api.signal(execution_id, message) } + end + + def pause(execution_id) + wrap { @agent_api.pause(execution_id) } + end + + def resume(execution_id) + wrap { @agent_api.resume(execution_id) } + end + + def cancel(execution_id, reason: nil) + wrap { @agent_api.cancel(execution_id, reason: reason) } + end + + def list_agents + wrap { @agent_api.list } + end + + def get_agent(name, version: nil) + wrap { @agent_api.get_agent(name, version: version) } + end + + def delete_agent(name, version: nil) + wrap { @agent_api.delete(name, version: version) } + end + + private + + def wrap + yield + rescue AgentApiError + raise + rescue ApiError => e + raise AgentApiError.from_api_error(e) + end + end + end +end diff --git a/lib/conductor/client/task_client.rb b/lib/conductor/client/task_client.rb index 4210646..918c1bb 100644 --- a/lib/conductor/client/task_client.rb +++ b/lib/conductor/client/task_client.rb @@ -45,6 +45,13 @@ def update_task(task_result) @task_api.update_task(task_result) end + # Update task status using the v2 endpoint (supports lease extension) + # @param [TaskResult] task_result Task result + # @return [Task, nil] Next task for this worker, if any + def update_task_v2(task_result) + @task_api.update_task_v2(task_result) + end + # Get task details # @param [String] task_id Task ID # @return [Task] Task object diff --git a/lib/conductor/configuration.rb b/lib/conductor/configuration.rb index ac3eddb..62a71fb 100644 --- a/lib/conductor/configuration.rb +++ b/lib/conductor/configuration.rb @@ -5,12 +5,43 @@ module Conductor # Configuration for Conductor client class Configuration - # Class-level auth token cache (shared across instances, like Python SDK) + # Legacy process-wide token cache. Tokens are now cached per Configuration + # instance so that two configurations (different servers or credentials) in + # one process never share a token. The class-level accessors remain for one + # release as a compatibility shim and warn once when used. @auth_token = nil @token_update_time = 0 class << self - attr_accessor :auth_token, :token_update_time + def auth_token + legacy_token_cache_warning + @auth_token + end + + def auth_token=(token) + legacy_token_cache_warning + @auth_token = token + end + + def token_update_time + legacy_token_cache_warning + @token_update_time + end + + def token_update_time=(time) + legacy_token_cache_warning + @token_update_time = time + end + + private + + def legacy_token_cache_warning + return if @legacy_token_cache_warned + + @legacy_token_cache_warned = true + warn '[Conductor] Configuration.auth_token / token_update_time are deprecated: ' \ + 'the auth token is cached per Configuration instance.' + end end attr_accessor :base_url, :server_api_url, :debug, :authentication_settings, @@ -29,6 +60,8 @@ def initialize(base_url: nil, server_api_url: nil, debug: false, @key_file = nil @proxy = nil @auth_token_ttl_min = auth_token_ttl_min + @auth_token = nil + @token_update_time = 0 # Resolve server URL @host = resolve_host(server_api_url, base_url) @@ -50,18 +83,18 @@ def disable_auth! @authentication_settings = nil end + # Cache an auth token on this configuration instance + # @param token [String] JWT returned by the /token endpoint def update_token(token) - self.class.auth_token = token - self.class.token_update_time = (Time.now.to_f * 1000).to_i + @auth_token = token + @token_update_time = (Time.now.to_f * 1000).to_i end - def auth_token - self.class.auth_token - end + # @return [String, nil] The cached auth token for this configuration + attr_reader :auth_token - def token_update_time - self.class.token_update_time - end + # @return [Integer] Epoch milliseconds of the last token update (0 when never set) + attr_reader :token_update_time # Alias for server URL (used in some places) def server_url diff --git a/lib/conductor/exceptions.rb b/lib/conductor/exceptions.rb index 174f76d..952192f 100644 --- a/lib/conductor/exceptions.rb +++ b/lib/conductor/exceptions.rb @@ -1,5 +1,7 @@ # frozen_string_literal: true +require 'json' + module Conductor # Base exception for all Conductor errors class ConductorError < StandardError; end @@ -71,6 +73,40 @@ def build_message end end + # Error returned by the agent REST API (/api/agent/*). The server answers 4xx with + # {"error": "", "status": }; +error+ carries that message when present. + class AgentApiError < ApiError + attr_reader :error + + def initialize(message = nil, status: nil, code: nil, reason: nil, body: nil, headers: nil) + @error = parse_error(body) + super(message || @error, status: status, code: code, reason: reason, body: body, headers: headers) + end + + # Build from a generic ApiError raised by the transport layer + # @param error [ApiError] + # @return [AgentApiError] + def self.from_api_error(error) + klass = error.not_found? ? AgentNotFoundError : AgentApiError + klass.new(error.message, status: error.status, code: error.code, reason: error.reason, + body: error.body, headers: error.headers) + end + + private + + def parse_error(body) + return nil unless body.is_a?(String) && !body.empty? + + data = JSON.parse(body) + data['error'] if data.is_a?(Hash) + rescue JSON::ParserError + nil + end + end + + # Agent, execution, or deployment not found (404 from /api/agent/*) + class AgentNotFoundError < AgentApiError; end + # Non-retryable worker error (terminal failure) class NonRetryableError < ConductorError; end diff --git a/lib/conductor/http/api/agent_resource_api.rb b/lib/conductor/http/api/agent_resource_api.rb new file mode 100644 index 0000000..b8cd3d2 --- /dev/null +++ b/lib/conductor/http/api/agent_resource_api.rb @@ -0,0 +1,138 @@ +# frozen_string_literal: true + +require_relative '../api_client' + +module Conductor + module Http + module Api + # AgentResourceApi - REST bindings for the server-side agent runtime (/api/agent/*) + # + # Every method returns the parsed JSON body as a Hash (or Array) so that callers + # see exactly the keys the server sent (executionId, requiredWorkers, isComplete, ...). + # The SSE stream endpoint is not here: it needs a long-lived streaming connection and + # lives in Conductor::Agents::Runtime::SseClient. + class AgentResourceApi + HASH = 'Hash' + + attr_accessor :api_client + + def initialize(api_client = nil) + @api_client = api_client || ApiClient.new + end + + # Start an agent execution + # @param body [Hash] AgentStartRequest: agentConfig | name, prompt, sessionId, media, context, runId, ... + # @return [Hash] { "executionId", "agentName", "requiredWorkers" } + def start(body) + post('/agent/start', body) + end + + # Register (deploy) an agent definition without starting it + # @param body [Hash] AgentStartRequest with agentConfig + # @return [Hash] { "agentName", "requiredWorkers" } + def deploy(body) + post('/agent/deploy', body) + end + + # Compile an agent config into a workflow definition without registering it + # @param body [Hash] AgentStartRequest with agentConfig + # @return [Hash] { "workflowDef", "requiredWorkers" } + def compile(body) + post('/agent/compile', body) + end + + # Get the status of an execution + # @return [Hash] { "executionId", "status", "isComplete", "isRunning", "isWaiting", "output", "pendingTool", ... } + def status(execution_id) + get('/agent/{executionId}/status', execution_id) + end + + # Get an execution with its tasks and token usage + # @return [Hash] { "executionId", "status", "output", "tokenUsage", "tasks" } + def execution(execution_id) + get('/agent/execution/{executionId}', execution_id) + end + + # List executions + # @param params [Hash] start, size, sort, freeText, status, agentName, sessionId + # @return [Hash] { "totalHits", "results" } + def executions(params = {}) + @api_client.call_api('/agent/executions', 'GET', query_params: params, return_type: HASH, + return_http_data_only: true) + end + + # Respond to a waiting execution (approval, human input, free text) + # @param body [Hash] e.g. { "approved" => true } or { "approved" => false, "reason" => "..." } + def respond(execution_id, body) + post_action(execution_id, 'respond', body) + end + + # Ask the agent loop to stop after the current iteration + def stop(execution_id) + post_action(execution_id, 'stop') + end + + # Inject a signal message into the next LLM turn + def signal(execution_id, message) + post_action(execution_id, 'signal', { 'message' => message }) + end + + # Pause the execution + def pause(execution_id) + @api_client.call_api('/agent/{executionId}/pause', 'PUT', path_params: { executionId: execution_id }, + return_http_data_only: true) + end + + # Resume a paused execution + def resume(execution_id) + @api_client.call_api('/agent/{executionId}/resume', 'PUT', path_params: { executionId: execution_id }, + return_http_data_only: true) + end + + # Cancel (terminate) the execution + def cancel(execution_id, reason: nil) + query = reason ? { reason: reason } : {} + @api_client.call_api('/agent/{executionId}/cancel', 'DELETE', path_params: { executionId: execution_id }, + query_params: query, return_http_data_only: true) + end + + # List deployed agents + # @return [Array] + def list + @api_client.call_api('/agent/list', 'GET', return_type: 'Array', return_http_data_only: true) + end + + # Get a deployed agent definition by name + # @return [Hash] the agentConfig as deployed + def get_agent(name, version: nil) + query = version ? { version: version } : {} + @api_client.call_api('/agent/{name}', 'GET', path_params: { name: name }, query_params: query, + return_type: HASH, return_http_data_only: true) + end + + # Delete a deployed agent definition + def delete(name, version: nil) + query = version ? { version: version } : {} + @api_client.call_api('/agent/{name}', 'DELETE', path_params: { name: name }, query_params: query, + return_http_data_only: true) + end + + private + + def get(path, execution_id) + @api_client.call_api(path, 'GET', path_params: { executionId: execution_id }, return_type: HASH, + return_http_data_only: true) + end + + def post(path, body) + @api_client.call_api(path, 'POST', body: body, return_type: HASH, return_http_data_only: true) + end + + def post_action(execution_id, action, body = nil) + @api_client.call_api("/agent/{executionId}/#{action}", 'POST', path_params: { executionId: execution_id }, + body: body, return_http_data_only: true) + end + end + end + end +end diff --git a/lib/conductor/http/api/task_resource_api.rb b/lib/conductor/http/api/task_resource_api.rb index b0adc11..c096825 100644 --- a/lib/conductor/http/api/task_resource_api.rb +++ b/lib/conductor/http/api/task_resource_api.rb @@ -73,6 +73,21 @@ def update_task(body) ) end + # Update task status using the v2 endpoint (POST /tasks/update-v2) + # Supports lease extension via TaskResult#extend_lease and returns the next + # task for the same worker when the server has one queued. + # @param [TaskResult] body Task result + # @return [Task, nil] Next task if the server returned one, nil on 204 + def update_task_v2(body) + @api_client.call_api( + '/tasks/update-v2', + 'POST', + body: body, + return_type: 'Task', + return_http_data_only: true + ) + end + # Get task details # @param [String] task_id Task ID # @return [Task] Task object diff --git a/lib/conductor/http/models/task.rb b/lib/conductor/http/models/task.rb index 27667df..0912efe 100644 --- a/lib/conductor/http/models/task.rb +++ b/lib/conductor/http/models/task.rb @@ -50,7 +50,8 @@ class Task < BaseModel first_start_time: 'Integer', loop_over_task: 'Boolean', task_definition: 'TaskDef', - queue_wait_time: 'Integer' + queue_wait_time: 'Integer', + runtime_metadata: 'Hash' }.freeze ATTRIBUTE_MAP = { @@ -96,7 +97,8 @@ class Task < BaseModel first_start_time: :firstStartTime, loop_over_task: :loopOverTask, task_definition: :taskDefinition, - queue_wait_time: :queueWaitTime + queue_wait_time: :queueWaitTime, + runtime_metadata: :runtimeMetadata }.freeze attr_accessor :task_type, :status, :input_data, :reference_task_name, @@ -112,6 +114,11 @@ class Task < BaseModel :iteration, :sub_workflow_id, :subworkflow_changed, :parent_task_id, :first_start_time, :loop_over_task, :task_definition, :queue_wait_time + # Wire-only map of secret name => resolved value. The server fills it at poll + # time from TaskDef#runtime_metadata (names) and never persists it. + # @return [Hash] + attr_accessor :runtime_metadata + # Initialize a new Task # @param [Hash] attributes Model attributes in the form of hash def initialize(attributes = {}) @@ -125,6 +132,7 @@ def initialize(attributes = {}) # Set default values for collections @input_data ||= {} @output_data ||= {} + @runtime_metadata ||= {} end # Check if task is in terminal state diff --git a/lib/conductor/http/models/task_def.rb b/lib/conductor/http/models/task_def.rb index 18912ac..b26f2bd 100644 --- a/lib/conductor/http/models/task_def.rb +++ b/lib/conductor/http/models/task_def.rb @@ -37,7 +37,9 @@ class TaskDef < BaseModel execution_name_space: 'String', owner_email: 'String', poll_timeout_seconds: 'Integer', - backoff_scale_factor: 'Integer' + backoff_scale_factor: 'Integer', + enforce_schema: 'Boolean', + runtime_metadata: 'Array' }.freeze ATTRIBUTE_MAP = { @@ -58,7 +60,9 @@ class TaskDef < BaseModel execution_name_space: :executionNameSpace, owner_email: :ownerEmail, poll_timeout_seconds: :pollTimeoutSeconds, - backoff_scale_factor: :backoffScaleFactor + backoff_scale_factor: :backoffScaleFactor, + enforce_schema: :enforceSchema, + runtime_metadata: :runtimeMetadata }.freeze attr_accessor :name, :description, :retry_count, :timeout_seconds, @@ -67,7 +71,12 @@ class TaskDef < BaseModel :concurrent_exec_limit, :rate_limit_per_frequency, :rate_limit_frequency_in_seconds, :isolation_group_id, :execution_name_space, :owner_email, :poll_timeout_seconds, - :backoff_scale_factor + :backoff_scale_factor, :enforce_schema + + # Names of secrets the server must resolve and attach to every task of this + # type (delivered as Task#runtime_metadata). Requires conductor-oss >= 3.32.0-rc.8. + # @return [Array] + attr_accessor :runtime_metadata def initialize(params = {}) @name = params[:name] @@ -88,6 +97,8 @@ def initialize(params = {}) @owner_email = params[:owner_email] @poll_timeout_seconds = params[:poll_timeout_seconds] @backoff_scale_factor = params[:backoff_scale_factor] || 1 + @enforce_schema = params.fetch(:enforce_schema, false) + @runtime_metadata = params[:runtime_metadata] || [] end end end diff --git a/lib/conductor/orkes/orkes_clients.rb b/lib/conductor/orkes/orkes_clients.rb index 6e36a99..a365a5d 100644 --- a/lib/conductor/orkes/orkes_clients.rb +++ b/lib/conductor/orkes/orkes_clients.rb @@ -61,6 +61,10 @@ def get_schema_client Client::SchemaClient.new(@api_client) end + def get_agent_client + Client::AgentClient.new(@api_client) + end + def get_workflow_executor Workflow::WorkflowExecutor.new(@configuration) end diff --git a/lib/conductor/worker/task_runner.rb b/lib/conductor/worker/task_runner.rb index 32fe0c5..c86ce90 100644 --- a/lib/conductor/worker/task_runner.rb +++ b/lib/conductor/worker/task_runner.rb @@ -68,6 +68,9 @@ def initialize(worker, configuration:, event_dispatcher: nil, logger: nil) @poll_count = Concurrent::AtomicFixnum.new(0) @shutdown = Concurrent::AtomicBoolean.new(false) @mutex = Mutex.new + # Prefer POST /tasks/update-v2 (lease extension, next-task return); fall back + # to POST /tasks once if the server does not serve it (404/405). + @use_update_v2 = Concurrent::AtomicBoolean.new(true) end # Main polling loop (runs until shutdown) @@ -458,7 +461,7 @@ def update_task_with_retry(task_result) start_time = Time.now begin - @task_client.update_task(task_result) + send_task_update(task_result) duration_ms = (Time.now - start_time) * 1000 publish_task_update_completed(task_result, duration_ms) @@ -476,6 +479,23 @@ def update_task_with_retry(task_result) end end + # Send the task result to the server, preferring the v2 endpoint + # @param task_result [TaskResult] + def send_task_update(task_result) + return @task_client.update_task(task_result) unless @use_update_v2.true? + + task_result.extend_lease = false if task_result.extend_lease.nil? + begin + @task_client.update_task_v2(task_result) + rescue ApiError => e + raise unless [404, 405].include?(e.status) + + @logger.info('Server does not support /tasks/update-v2, falling back to /tasks') + @use_update_v2.make_false + @task_client.update_task(task_result) + end + end + def publish_task_update_completed(task_result, duration_ms) @event_dispatcher.publish(Events::TaskUpdateCompleted.new( task_type: @worker.task_definition_name, diff --git a/lib/conductor/worker/worker.rb b/lib/conductor/worker/worker.rb index f86a1f4..28bfd38 100644 --- a/lib/conductor/worker/worker.rb +++ b/lib/conductor/worker/worker.rb @@ -22,7 +22,7 @@ class Worker attr_accessor :poll_interval, :thread_count, :domain, :worker_id, :poll_timeout, :register_task_def, :overwrite_task_def, :strict_schema, :paused, :isolation, :executor, - :task_def_template + :task_def_template, :lease_extend_enabled # Default configuration values DEFAULTS = { @@ -36,7 +36,8 @@ class Worker strict_schema: false, paused: false, isolation: :thread, - executor: :thread_pool + executor: :thread_pool, + lease_extend_enabled: false }.freeze # Initialize a worker diff --git a/lib/conductor/worker/worker_config.rb b/lib/conductor/worker/worker_config.rb index 7092a67..62e31bc 100644 --- a/lib/conductor/worker/worker_config.rb +++ b/lib/conductor/worker/worker_config.rb @@ -22,7 +22,8 @@ class WorkerConfig strict_schema: { type: :boolean, default: false }, paused: { type: :boolean, default: false }, isolation: { type: :symbol, default: :thread }, # :thread or :ractor - executor: { type: :symbol, default: :thread_pool } # :thread_pool or :fiber + executor: { type: :symbol, default: :thread_pool }, # :thread_pool or :fiber + lease_extend_enabled: { type: :boolean, default: false } }.freeze class << self diff --git a/spec/conductor/client/agent_client_spec.rb b/spec/conductor/client/agent_client_spec.rb new file mode 100644 index 0000000..314002a --- /dev/null +++ b/spec/conductor/client/agent_client_spec.rb @@ -0,0 +1,69 @@ +# frozen_string_literal: true + +require 'spec_helper' + +RSpec.describe Conductor::Client::AgentClient do + let(:api_client) { instance_double(Conductor::Http::ApiClient) } + let(:agent_api) { instance_double(Conductor::Http::Api::AgentResourceApi) } + let(:client) { described_class.new(api_client) } + + before do + allow(Conductor::Http::Api::AgentResourceApi).to receive(:new).with(api_client).and_return(agent_api) + end + + it 'delegates start/deploy/compile' do + expect(agent_api).to receive(:start).with({ 'prompt' => 'x' }).and_return({ 'executionId' => 'E' }) + expect(agent_api).to receive(:deploy).with({ 'agentConfig' => {} }) + expect(agent_api).to receive(:compile).with({ 'agentConfig' => {} }) + + expect(client.start_agent('prompt' => 'x')).to eq('executionId' => 'E') + client.deploy_agent('agentConfig' => {}) + client.compile_agent('agentConfig' => {}) + end + + it 'delegates status, execution and list' do + expect(agent_api).to receive(:status).with('E') + expect(agent_api).to receive(:execution).with('E') + expect(agent_api).to receive(:executions).with({ size: 1 }) + client.get_status('E') + client.get_execution('E') + client.list_executions(size: 1) + end + + it 'builds approval bodies like the Python client' do + expect(agent_api).to receive(:respond).with('E', { 'approved' => true }) + expect(agent_api).to receive(:respond).with('E', { 'approved' => false, 'reason' => 'Needs a manager' }) + expect(agent_api).to receive(:respond).with('E', { 'message' => 'hello' }) + expect(agent_api).to receive(:respond).with('E', { 'output' => 42 }) + + client.approve('E') + client.reject('E', 'Needs a manager') + client.send_message('E', 'hello') + client.respond('E', 42) + end + + it 'delegates control operations' do + expect(agent_api).to receive(:stop).with('E') + expect(agent_api).to receive(:signal).with('E', 'go') + expect(agent_api).to receive(:pause).with('E') + expect(agent_api).to receive(:resume).with('E') + expect(agent_api).to receive(:cancel).with('E', reason: 'r') + client.stop('E') + client.signal('E', 'go') + client.pause('E') + client.resume('E') + client.cancel('E', reason: 'r') + end + + it 'maps 404 to AgentNotFoundError and other errors to AgentApiError with the server message' do + allow(agent_api).to receive(:status).and_raise(Conductor::ApiError.new('nf', status: 404)) + expect { client.get_status('E') }.to raise_error(Conductor::AgentNotFoundError) + + body = { error: 'agentConfig.model is required', status: 400 }.to_json + allow(agent_api).to receive(:start).and_raise(Conductor::ApiError.new('bad', status: 400, body: body)) + expect { client.start_agent({}) }.to raise_error(Conductor::AgentApiError) { |e| + expect(e.status).to eq(400) + expect(e.error).to eq('agentConfig.model is required') + } + end +end diff --git a/spec/conductor/configuration/token_cache_spec.rb b/spec/conductor/configuration/token_cache_spec.rb new file mode 100644 index 0000000..42ad3e6 --- /dev/null +++ b/spec/conductor/configuration/token_cache_spec.rb @@ -0,0 +1,29 @@ +# frozen_string_literal: true + +require 'spec_helper' + +RSpec.describe Conductor::Configuration, '#auth_token' do + it 'starts with no token and a zero update time' do + config = described_class.new(server_api_url: 'http://a/api') + expect(config.auth_token).to be_nil + expect(config.token_update_time).to eq(0) + end + + it 'caches the token per instance' do + a = described_class.new(server_api_url: 'http://a/api') + b = described_class.new(server_api_url: 'http://b/api') + + a.update_token('token-a') + + expect(a.auth_token).to eq('token-a') + expect(a.token_update_time).to be > 0 + expect(b.auth_token).to be_nil + expect(b.token_update_time).to eq(0) + end + + it 'keeps the deprecated class-level accessors working with a warning' do + expect { described_class.auth_token = 'legacy' }.to output(/deprecated/).to_stderr + expect(described_class.auth_token).to eq('legacy') + described_class.auth_token = nil + end +end diff --git a/spec/conductor/http/api/agent_resource_api_spec.rb b/spec/conductor/http/api/agent_resource_api_spec.rb new file mode 100644 index 0000000..fb7e000 --- /dev/null +++ b/spec/conductor/http/api/agent_resource_api_spec.rb @@ -0,0 +1,96 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'webmock/rspec' + +RSpec.describe Conductor::Http::Api::AgentResourceApi do + let(:base) { 'http://localhost:8080/api' } + let(:configuration) { Conductor::Configuration.new(server_api_url: base) } + let(:api_client) { Conductor::Http::ApiClient.new(configuration: configuration) } + let(:api) { described_class.new(api_client) } + + before { WebMock.enable! } + after { WebMock.reset! } + + it 'POSTs the start request and returns the parsed body' do + stub = stub_request(:post, "#{base}/agent/start") + .with(body: hash_including('prompt' => 'hi', 'agentConfig' => hash_including('name' => 'a'))) + .to_return(status: 200, headers: { 'Content-Type' => 'application/json' }, + body: { executionId: 'EXEC_1', agentName: 'a', requiredWorkers: ['get_weather'] }.to_json) + + result = api.start('agentConfig' => { 'name' => 'a' }, 'prompt' => 'hi') + + expect(stub).to have_been_requested + expect(result).to eq('executionId' => 'EXEC_1', 'agentName' => 'a', 'requiredWorkers' => ['get_weather']) + end + + it 'POSTs deploy and compile' do + stub_request(:post, "#{base}/agent/deploy").to_return(body: { agentName: 'a' }.to_json, + headers: { 'Content-Type' => 'application/json' }) + stub_request(:post, "#{base}/agent/compile").to_return(body: { workflowDef: {}, requiredWorkers: [] }.to_json, + headers: { 'Content-Type' => 'application/json' }) + expect(api.deploy('agentConfig' => {})).to eq('agentName' => 'a') + expect(api.compile('agentConfig' => {})).to include('workflowDef') + end + + it 'GETs status and execution' do + stub_request(:get, "#{base}/agent/EXEC_1/status") + .to_return(body: { executionId: 'EXEC_1', isComplete: true, isWaiting: false }.to_json, + headers: { 'Content-Type' => 'application/json' }) + stub_request(:get, "#{base}/agent/execution/EXEC_1") + .to_return(body: { tokenUsage: { totalTokens: 5 } }.to_json, headers: { 'Content-Type' => 'application/json' }) + + expect(api.status('EXEC_1')).to include('isComplete' => true) + expect(api.execution('EXEC_1')).to eq('tokenUsage' => { 'totalTokens' => 5 }) + end + + it 'GETs executions with query params' do + stub = stub_request(:get, "#{base}/agent/executions").with(query: { 'size' => '5', 'agentName' => 'a' }) + .to_return(body: { totalHits: 0, results: [] }.to_json, + headers: { 'Content-Type' => 'application/json' }) + expect(api.executions(size: 5, agentName: 'a')).to eq('totalHits' => 0, 'results' => []) + expect(stub).to have_been_requested + end + + it 'POSTs respond, stop and signal with the exact bodies' do + respond = stub_request(:post, "#{base}/agent/EXEC_1/respond").with(body: { approved: false, reason: 'no' }.to_json) + stop = stub_request(:post, "#{base}/agent/EXEC_1/stop") + signal = stub_request(:post, "#{base}/agent/EXEC_1/signal").with(body: { message: 'hurry' }.to_json) + + api.respond('EXEC_1', { 'approved' => false, 'reason' => 'no' }) + api.stop('EXEC_1') + api.signal('EXEC_1', 'hurry') + + expect(respond).to have_been_requested + expect(stop).to have_been_requested + expect(signal).to have_been_requested + end + + it 'PUTs pause/resume and DELETEs cancel with a reason' do + pause = stub_request(:put, "#{base}/agent/EXEC_1/pause") + resume = stub_request(:put, "#{base}/agent/EXEC_1/resume") + cancel = stub_request(:delete, "#{base}/agent/EXEC_1/cancel").with(query: { 'reason' => 'bye' }) + + api.pause('EXEC_1') + api.resume('EXEC_1') + api.cancel('EXEC_1', reason: 'bye') + + expect(pause).to have_been_requested + expect(resume).to have_been_requested + expect(cancel).to have_been_requested + end + + it 'lists, gets and deletes deployed agents' do + stub_request(:get, "#{base}/agent/list").to_return(body: [{ name: 'a' }].to_json, + headers: { 'Content-Type' => 'application/json' }) + stub_request(:get, "#{base}/agent/a").with(query: { 'version' => '2' }) + .to_return(body: { name: 'a' }.to_json, + headers: { 'Content-Type' => 'application/json' }) + del = stub_request(:delete, "#{base}/agent/a") + + expect(api.list).to eq([{ 'name' => 'a' }]) + expect(api.get_agent('a', version: 2)).to eq('name' => 'a') + api.delete('a') + expect(del).to have_been_requested + end +end diff --git a/spec/conductor/models/runtime_metadata_spec.rb b/spec/conductor/models/runtime_metadata_spec.rb new file mode 100644 index 0000000..4ec0067 --- /dev/null +++ b/spec/conductor/models/runtime_metadata_spec.rb @@ -0,0 +1,53 @@ +# frozen_string_literal: true + +require 'spec_helper' + +RSpec.describe Conductor::Http::Models::Task, '#runtime_metadata' do + it 'defaults to an empty hash' do + expect(described_class.new.runtime_metadata).to eq({}) + end + + it 'deserializes the wire-only secret map from a poll response' do + task = described_class.from_hash( + 'taskType' => 'create_issue', + 'taskId' => 'TASK_1', + 'inputData' => { 'title' => 'bug' }, + 'runtimeMetadata' => { 'GH_TOKEN' => 'ghp_secret' } + ) + expect(task.runtime_metadata).to eq('GH_TOKEN' => 'ghp_secret') + end +end + +RSpec.describe Conductor::Http::Models::TaskDef, '#runtime_metadata' do + it 'defaults to an empty list and enforce_schema false' do + task_def = described_class.new(name: 't') + expect(task_def.runtime_metadata).to eq([]) + expect(task_def.enforce_schema).to be false + end + + it 'serializes secret names as runtimeMetadata' do + task_def = described_class.new(name: 'create_issue', runtime_metadata: ['GH_TOKEN']) + hash = task_def.to_h + expect(hash['runtimeMetadata']).to eq(['GH_TOKEN']) + expect(hash['enforceSchema']).to be false + end + + it 'keeps agent worker defaults when used as a template (timeout 0 is not overridden)' do + template = described_class.new(name: 'x', timeout_seconds: 0, response_timeout_seconds: 10, + retry_count: 2, retry_delay_seconds: 2, + retry_logic: 'LINEAR_BACKOFF', timeout_policy: 'RETRY', + runtime_metadata: ['GH_TOKEN']) + worker = Conductor::Worker::Worker.new('get_weather', register_task_def: true, + task_def_template: template) { {} } + registrar = Conductor::Worker::TaskDefinitionRegistrar.new(Conductor::Configuration.new, logger: Logger.new(nil)) + task_def = registrar.send(:build_task_definition, worker) + + expect(task_def.name).to eq('get_weather') + expect(task_def.timeout_seconds).to eq(0) + expect(task_def.response_timeout_seconds).to eq(10) + expect(task_def.retry_count).to eq(2) + expect(task_def.retry_logic).to eq('LINEAR_BACKOFF') + expect(task_def.timeout_policy).to eq('RETRY') + expect(task_def.runtime_metadata).to eq(['GH_TOKEN']) + end +end diff --git a/spec/conductor/orkes/orkes_clients_spec.rb b/spec/conductor/orkes/orkes_clients_spec.rb index 31ebeda..69c546d 100644 --- a/spec/conductor/orkes/orkes_clients_spec.rb +++ b/spec/conductor/orkes/orkes_clients_spec.rb @@ -81,6 +81,12 @@ end end + describe '#get_agent_client' do + it 'returns an AgentClient' do + expect(clients.get_agent_client).to be_a(Conductor::Client::AgentClient) + end + end + describe '#get_schema_client' do it 'returns a SchemaClient' do result = clients.get_schema_client diff --git a/spec/conductor/worker/task_runner_spec.rb b/spec/conductor/worker/task_runner_spec.rb index 669d32d..e98bdff 100644 --- a/spec/conductor/worker/task_runner_spec.rb +++ b/spec/conductor/worker/task_runner_spec.rb @@ -37,6 +37,7 @@ allow(Conductor::Client::TaskClient).to receive(:new).and_return(task_client) allow(task_client).to receive(:batch_poll_tasks).and_return([]) allow(task_client).to receive(:update_task) + allow(task_client).to receive(:update_task_v2) end describe '#initialize' do @@ -260,7 +261,7 @@ before do allow(task_client).to receive(:batch_poll_tasks).and_return([task_data]) - allow(task_client).to receive(:update_task).and_raise(StandardError.new('Update failed')) + allow(task_client).to receive(:update_task_v2).and_raise(StandardError.new('Update failed')) event_dispatcher.register(Conductor::Worker::Events::TaskUpdateFailure, ->(event) { received_events << [:update_failure, event] }) @@ -365,7 +366,7 @@ before do allow(task_client).to receive(:batch_poll_tasks).and_return([task_data]) - allow(task_client).to receive(:update_task).and_raise(StandardError.new('Update failed')) + allow(task_client).to receive(:update_task_v2).and_raise(StandardError.new('Update failed')) event_dispatcher.register(Conductor::Worker::Events::TaskUpdateFailure, ->(event) { received_events << [:update_failure, event] }) diff --git a/spec/conductor/worker/task_update_v2_spec.rb b/spec/conductor/worker/task_update_v2_spec.rb new file mode 100644 index 0000000..c399759 --- /dev/null +++ b/spec/conductor/worker/task_update_v2_spec.rb @@ -0,0 +1,41 @@ +# frozen_string_literal: true + +require 'spec_helper' + +RSpec.describe Conductor::Worker::TaskRunner, '#send_task_update' do + let(:configuration) { Conductor::Configuration.new(server_api_url: 'http://localhost:8080/api') } + let(:task_client) { instance_double(Conductor::Client::TaskClient) } + let(:worker) { Conductor::Worker::Worker.new('t') { {} } } + let(:runner) do + allow(Conductor::Client::TaskClient).to receive(:new).and_return(task_client) + described_class.new(worker, configuration: configuration, logger: Logger.new(nil)) + end + let(:result) { Conductor::Http::Models::TaskResult.complete } + + it 'posts to update-v2 with extendLease false by default' do + expect(task_client).to receive(:update_task_v2) do |task_result| + expect(task_result.extend_lease).to be false + nil + end + runner.send(:send_task_update, result) + end + + it 'falls back to /tasks once when the server does not serve update-v2' do + expect(task_client).to receive(:update_task_v2).once.and_raise(Conductor::ApiError.new('nope', status: 404)) + expect(task_client).to receive(:update_task).twice + runner.send(:send_task_update, result) + runner.send(:send_task_update, result) + end + + it 're-raises other API errors' do + allow(task_client).to receive(:update_task_v2).and_raise(Conductor::ApiError.new('boom', status: 500)) + expect { runner.send(:send_task_update, result) }.to raise_error(Conductor::ApiError) + end +end + +RSpec.describe Conductor::Worker::Worker, '#lease_extend_enabled' do + it 'defaults to false and can be enabled' do + expect(described_class.new('t') { {} }.lease_extend_enabled).to be false + expect(described_class.new('t', lease_extend_enabled: true) { {} }.lease_extend_enabled).to be true + end +end From d58dc287ec631cfb3e324e3f04cc12e8ced24099 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Tue, 8 Sep 2026 15:37:48 -0700 Subject: [PATCH 03/20] Phase 1: Conductor::Agents definition layer and agentConfig serializer - ToolDef/ToolType with Python defaults and factories (http, api, mcp, human, agent, media, RAG, wait_for_message) - Tools DSL: `tool def` builds a ToolDef from keyword defaults via the method AST (city: String required, units: 'metric' optional, [String] arrays, %w[] enums), scans literal secret()/secrets_env() names as credentials, describe / requires_approval / tool_credentials, per-module lookup (Weather[:current]), global registry for Agent#add_tool(:name), RubyLLM::Tool adapter, AST-less fallback - Guardrail / RegexGuardrail / LlmGuardrail, Termination::* with & and |, Handoff::*, CallbackHandler, ConversationMemory, PromptTemplate - Agent with team/guardrail/termination/approval sugar (add_tool, add_agent, hands_off_to, redact, stop_when, stop_after, on_approval, >>) - ConfigSerializer: same wire shape as Python (camelCase, nil-compacted, agent credentials top-level, tool credentials under config), team parents inherit the first member's model, member hands_off_to hoists into a swarm team - Contract tests: vendored agent-schema.json + 19 golden configs from python-sdk; every golden serializes identically and validates; examples/agents/golden_agents.rb and dump_agent_configs.rb regenerate them Co-Authored-By: Claude Fable 5.1 --- .rubocop.yml | 6 + examples/agents/dump_agent_configs.rb | 30 ++ examples/agents/golden_agents.rb | 382 ++++++++++++++++++ lib/conductor/agents.rb | 41 ++ lib/conductor/agents/agent.rb | 317 +++++++++++++++ lib/conductor/agents/callback_handler.rb | 64 +++ lib/conductor/agents/config_serializer.rb | 236 +++++++++++ lib/conductor/agents/errors.rb | 26 ++ lib/conductor/agents/guardrail.rb | 142 +++++++ lib/conductor/agents/handoff.rb | 94 +++++ lib/conductor/agents/memory.rb | 75 ++++ lib/conductor/agents/prompt_template.rb | 23 ++ lib/conductor/agents/runtime/secrets.rb | 54 +++ lib/conductor/agents/termination.rb | 192 +++++++++ lib/conductor/agents/tool_def.rb | 272 +++++++++++++ lib/conductor/agents/tools.rb | 171 ++++++++ .../agents/tools/ruby_llm_adapter.rb | 73 ++++ lib/conductor/agents/tools/schema_builder.rb | 190 +++++++++ lib/conductor/agents/tools/secret_scanner.rb | 45 +++ spec/conductor/agents/agent_spec.rb | 164 ++++++++ .../conductor/agents/callback_handler_spec.rb | 47 +++ .../agents/config_serializer_spec.rb | 165 ++++++++ spec/conductor/agents/contract_spec.rb | 40 ++ spec/conductor/agents/guardrail_spec.rb | 59 +++ spec/conductor/agents/handoff_spec.rb | 37 ++ spec/conductor/agents/memory_spec.rb | 43 ++ spec/conductor/agents/termination_spec.rb | 59 +++ spec/conductor/agents/tool_def_spec.rb | 104 +++++ spec/conductor/agents/tools_spec.rb | 147 +++++++ spec/fixtures/agents/agent-schema.json | 87 ++++ .../agents/configs/01_basic_agent.json | 7 + spec/fixtures/agents/configs/02_tools.json | 80 ++++ .../agents/configs/03_structured_output.json | 61 +++ spec/fixtures/agents/configs/05_handoffs.json | 101 +++++ .../configs/06_sequential_pipeline.json | 34 ++ .../agents/configs/07_parallel_agents.json | 34 ++ .../agents/configs/08_router_agent.json | 43 ++ .../agents/configs/10_guardrails.json | 60 +++ .../configs/13_hierarchical_agents.json | 77 ++++ .../configs/17_swarm_orchestration.json | 39 ++ .../19_composable_termination_and.json | 43 ++ .../19_composable_termination_complex.json | 56 +++ .../configs/19_composable_termination_or.json | 22 + .../19_composable_termination_simple.json | 34 ++ .../agents/configs/21_regex_guardrails.json | 56 +++ .../agents/configs/22_llm_guardrails.json | 20 + .../agents/configs/45_agent_tool.json | 79 ++++ .../fixtures/agents/configs/47_callbacks.json | 40 ++ .../agents/configs/52_nested_strategies.json | 44 ++ spec/support/agent_tools.rb | 95 +++++ 50 files changed, 4410 insertions(+) create mode 100644 examples/agents/dump_agent_configs.rb create mode 100644 examples/agents/golden_agents.rb create mode 100644 lib/conductor/agents.rb create mode 100644 lib/conductor/agents/agent.rb create mode 100644 lib/conductor/agents/callback_handler.rb create mode 100644 lib/conductor/agents/config_serializer.rb create mode 100644 lib/conductor/agents/errors.rb create mode 100644 lib/conductor/agents/guardrail.rb create mode 100644 lib/conductor/agents/handoff.rb create mode 100644 lib/conductor/agents/memory.rb create mode 100644 lib/conductor/agents/prompt_template.rb create mode 100644 lib/conductor/agents/runtime/secrets.rb create mode 100644 lib/conductor/agents/termination.rb create mode 100644 lib/conductor/agents/tool_def.rb create mode 100644 lib/conductor/agents/tools.rb create mode 100644 lib/conductor/agents/tools/ruby_llm_adapter.rb create mode 100644 lib/conductor/agents/tools/schema_builder.rb create mode 100644 lib/conductor/agents/tools/secret_scanner.rb create mode 100644 spec/conductor/agents/agent_spec.rb create mode 100644 spec/conductor/agents/callback_handler_spec.rb create mode 100644 spec/conductor/agents/config_serializer_spec.rb create mode 100644 spec/conductor/agents/contract_spec.rb create mode 100644 spec/conductor/agents/guardrail_spec.rb create mode 100644 spec/conductor/agents/handoff_spec.rb create mode 100644 spec/conductor/agents/memory_spec.rb create mode 100644 spec/conductor/agents/termination_spec.rb create mode 100644 spec/conductor/agents/tool_def_spec.rb create mode 100644 spec/conductor/agents/tools_spec.rb create mode 100644 spec/fixtures/agents/agent-schema.json create mode 100644 spec/fixtures/agents/configs/01_basic_agent.json create mode 100644 spec/fixtures/agents/configs/02_tools.json create mode 100644 spec/fixtures/agents/configs/03_structured_output.json create mode 100644 spec/fixtures/agents/configs/05_handoffs.json create mode 100644 spec/fixtures/agents/configs/06_sequential_pipeline.json create mode 100644 spec/fixtures/agents/configs/07_parallel_agents.json create mode 100644 spec/fixtures/agents/configs/08_router_agent.json create mode 100644 spec/fixtures/agents/configs/10_guardrails.json create mode 100644 spec/fixtures/agents/configs/13_hierarchical_agents.json create mode 100644 spec/fixtures/agents/configs/17_swarm_orchestration.json create mode 100644 spec/fixtures/agents/configs/19_composable_termination_and.json create mode 100644 spec/fixtures/agents/configs/19_composable_termination_complex.json create mode 100644 spec/fixtures/agents/configs/19_composable_termination_or.json create mode 100644 spec/fixtures/agents/configs/19_composable_termination_simple.json create mode 100644 spec/fixtures/agents/configs/21_regex_guardrails.json create mode 100644 spec/fixtures/agents/configs/22_llm_guardrails.json create mode 100644 spec/fixtures/agents/configs/45_agent_tool.json create mode 100644 spec/fixtures/agents/configs/47_callbacks.json create mode 100644 spec/fixtures/agents/configs/52_nested_strategies.json create mode 100644 spec/support/agent_tools.rb diff --git a/.rubocop.yml b/.rubocop.yml index a384d8f..2816546 100644 --- a/.rubocop.yml +++ b/.rubocop.yml @@ -199,3 +199,9 @@ RSpec/FilePath: RSpec/VerifiedDoubles: Enabled: false + +# The agents Tools DSL has its own `describe :tool_name, 'text'`; spec support files use it +RSpec/DescribeSymbol: + Exclude: + - 'spec/support/**/*' + - 'examples/**/*' diff --git a/examples/agents/dump_agent_configs.rb b/examples/agents/dump_agent_configs.rb new file mode 100644 index 0000000..7e1129b --- /dev/null +++ b/examples/agents/dump_agent_configs.rb @@ -0,0 +1,30 @@ +#!/usr/bin/env ruby +# frozen_string_literal: true + +# Dump the serialized agentConfig JSON for the golden examples, for cross-SDK comparison +# with python-sdk/examples/agents/dump_agent_configs.py (same file names, sorted keys). +# +# CONDUCTOR_AGENT_LLM_MODEL=anthropic/claude-sonnet-4-6 bundle exec ruby examples/agents/dump_agent_configs.rb [out_dir] +require 'json' +require_relative '../../lib/conductor/agents' +require_relative 'golden_agents' + +out_dir = ARGV[0] || File.join(__dir__, '_configs') +Dir.mkdir(out_dir) unless Dir.exist?(out_dir) + +def sort_keys(value) + case value + when Hash then value.keys.sort.to_h { |k| [k, sort_keys(value[k])] } + when Array then value.map { |v| sort_keys(v) } + else value + end +end + +GoldenAgents::EXAMPLES.each do |name, build| + config = Conductor::Agents::ConfigSerializer.serialize(build.call) + File.write(File.join(out_dir, "#{name}.json"), JSON.pretty_generate(sort_keys(config))) + puts " [OK] #{name}" +rescue StandardError => e + puts " [FAIL] #{name}: #{e.message}" +end +puts "\nConfigs written to #{out_dir}" diff --git a/examples/agents/golden_agents.rb b/examples/agents/golden_agents.rb new file mode 100644 index 0000000..f41d2ae --- /dev/null +++ b/examples/agents/golden_agents.rb @@ -0,0 +1,382 @@ +# frozen_string_literal: true + +# The 19 agents whose serialized agentConfig must match the Python SDK byte for byte +# (spec/fixtures/agents/configs/*.json, vendored from python-sdk examples/agents/_configs). +# +# Used by spec/conductor/agents/contract_spec.rb and by dump_agent_configs.rb. +# Each example keeps its tools in its own module so that tools with the same name but +# different descriptions (get_weather in 02 vs 03) do not collide. +require 'conductor/agents' + +module GoldenAgents + MODEL = ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'anthropic/claude-sonnet-4-6') + A = Conductor::Agents + + module Ex02 + extend A::Tools + tool def get_weather(city: String) + {} + end + describe :get_weather, 'Get current weather for a city.' + tool def calculate(expression: String) + {} + end + describe :calculate, 'Evaluate a math expression.' + tool def send_email(to: String, subject: String, body: String) + {} + end + describe :send_email, 'Send an email.' + requires_approval :send_email + self[:send_email].timeout_seconds = 60 + end + + module Ex03 + extend A::Tools + tool def get_weather(city: String) + {} + end + describe :get_weather, 'Get current weather data for a city.' + end + + module Ex05 + extend A::Tools + tool def check_balance(account_id: String) + {} + end + describe :check_balance, 'Check the balance of a bank account.' + tool def lookup_order(order_id: String) + {} + end + describe :lookup_order, 'Look up the status of an order.' + tool def get_pricing(product: String) + {} + end + describe :get_pricing, 'Get pricing information for a product.' + end + + module Ex10 + extend A::Tools + tool def get_order_status(order_id: String) + {} + end + describe :get_order_status, 'Look up the current status of an order.' + tool def get_customer_info(customer_id: String) + {} + end + describe :get_customer_info, 'Retrieve customer details including payment info on file.' + end + + module Ex19 + extend A::Tools + tool def search(query: String) + '' + end + describe :search, 'Search for information.' + self[:search].output_schema = { 'type' => 'string' } + end + + module Ex21 + extend A::Tools + tool def get_user_profile(user_id: String) + {} + end + describe :get_user_profile, "Retrieve a user's profile from the database." + end + + module Ex45 + extend A::Tools + tool def search_knowledge_base(query: String) + {} + end + describe :search_knowledge_base, 'Search an internal knowledge base for information.' + tool def calculate(expression: String) + {} + end + describe :calculate, 'Evaluate a math expression safely.' + end + + module Ex47 + extend A::Tools + tool def get_facts(topic: String) + {} + end + describe :get_facts, 'Get interesting facts about a topic.' + end + + # 47_callbacks: before_model / after_model hooks + class MonitorHandler < A::CallbackHandler + def on_model_start(**_kwargs) + {} + end + + def on_model_end(**_kwargs) + {} + end + end + + EXAMPLES = { + '01_basic_agent' => lambda { + A::Agent.new(name: 'greeter', model: MODEL) + }, + + '02_tools' => lambda { + A::Agent.new( + name: 'tool_demo_agent', model: MODEL, + tools: [Ex02[:get_weather], Ex02[:calculate], Ex02[:send_email]], + instructions: 'You are a helpful assistant with access to weather, calculator, and email tools.' + ) + }, + + '03_structured_output' => lambda { + weather_report = { + 'title' => 'WeatherReport', + 'type' => 'object', + 'properties' => { + 'city' => { 'title' => 'City', 'type' => 'string' }, + 'temperature' => { 'title' => 'Temperature', 'type' => 'number' }, + 'condition' => { 'title' => 'Condition', 'type' => 'string' }, + 'recommendation' => { 'title' => 'Recommendation', 'type' => 'string' } + }, + 'required' => %w[city temperature condition recommendation] + } + A::Agent.new( + name: 'weather_reporter', model: MODEL, tools: [Ex03[:get_weather]], output_type: weather_report, + instructions: 'You are a weather reporter. Get the weather and provide a recommendation.' + ) + }, + + '05_handoffs' => lambda { + billing = A::Agent.new(name: 'billing', model: MODEL, tools: [Ex05[:check_balance]], + instructions: 'You handle billing questions: balances, payments, invoices.') + technical = A::Agent.new(name: 'technical', model: MODEL, tools: [Ex05[:lookup_order]], + instructions: 'You handle technical questions: order status, shipping, returns.') + sales = A::Agent.new(name: 'sales', model: MODEL, tools: [Ex05[:get_pricing]], + instructions: 'You handle sales questions: pricing, products, promotions.') + A::Agent.new(name: 'support', model: MODEL, agents: [billing, technical, sales], strategy: :handoff, + instructions: 'Route customer requests to the right specialist: billing, technical, or sales.') + }, + + '06_sequential_pipeline' => lambda { + researcher = A::Agent.new( + name: 'researcher', model: MODEL, + instructions: 'You are a researcher. Given a topic, provide key facts and data points. ' \ + 'Be thorough but concise. Output raw research findings.' + ) + writer = A::Agent.new( + name: 'writer', model: MODEL, + instructions: 'You are a writer. Take research findings and write a clear, engaging ' \ + 'article. Use headers and bullet points where appropriate.' + ) + editor = A::Agent.new( + name: 'editor', model: MODEL, + instructions: 'You are an editor. Review the article for clarity, grammar, and tone. ' \ + 'Make improvements and output the final polished version.' + ) + researcher >> writer >> editor + }, + + '07_parallel_agents' => lambda { + market = A::Agent.new( + name: 'market_analyst', model: MODEL, + instructions: 'You are a market analyst. Analyze the given topic from a market perspective: ' \ + 'market size, growth trends, key players, and opportunities.' + ) + risk = A::Agent.new( + name: 'risk_analyst', model: MODEL, + instructions: 'You are a risk analyst. Analyze the given topic for risks: ' \ + 'regulatory risks, technical risks, competitive threats, and mitigation strategies.' + ) + compliance = A::Agent.new( + name: 'compliance', model: MODEL, + instructions: 'You are a compliance specialist. Check the given topic for compliance considerations: ' \ + 'data privacy, regulatory requirements, and industry standards.' + ) + A::Agent.new(name: 'analysis', model: MODEL, agents: [market, risk, compliance], strategy: :parallel) + }, + + '08_router_agent' => lambda { + planner = A::Agent.new(name: 'planner', model: MODEL, + instructions: 'You create implementation plans. Break down tasks into clear numbered steps.') + coder = A::Agent.new(name: 'coder', model: MODEL, + instructions: 'You write code. Output clean, well-documented Python code.') + reviewer = A::Agent.new(name: 'reviewer', model: MODEL, + instructions: 'You review code. Check for bugs, style issues, and suggest improvements.') + A::Agent.new( + name: 'dev_team', model: MODEL, agents: [planner, coder, reviewer], strategy: :router, router: planner, + instructions: 'You are the tech lead. Route requests to the right team member: ' \ + 'planner for design/architecture, coder for implementation, reviewer for code review.' + ) + }, + + '10_guardrails' => lambda { + no_pii = A::Guardrail.new(name: 'no_pii', position: :output, on_fail: :retry) do |content| + if content =~ /\b\d{4}[\s-]?\d{4}[\s-]?\d{4}[\s-]?\d{4}\b/ || content =~ /\b\d{3}-\d{2}-\d{4}\b/ + A::GuardrailResult.new(passed: false, message: 'Your response contains PII. Redact it.') + else + A::GuardrailResult.new(passed: true) + end + end + A::Agent.new( + name: 'support_agent', model: MODEL, tools: [Ex10[:get_order_status], Ex10[:get_customer_info]], + guardrails: [no_pii], + instructions: 'You are a customer support assistant. Use the available tools to ' \ + 'answer questions about orders and customers. Always include all ' \ + 'details from the tool results in your response.' + ) + }, + + '13_hierarchical_agents' => lambda { + backend = A::Agent.new( + name: 'backend_dev', model: MODEL, + instructions: 'You are a backend developer. You design APIs, databases, and server ' \ + 'architecture. Provide technical recommendations with code examples.' + ) + frontend = A::Agent.new( + name: 'frontend_dev', model: MODEL, + instructions: 'You are a frontend developer. You design UI components, user flows, ' \ + 'and client-side architecture. Provide recommendations with code examples.' + ) + content = A::Agent.new( + name: 'content_writer', model: MODEL, + instructions: 'You are a content writer. You create blog posts, landing page copy, ' \ + 'and marketing materials. Write engaging, clear content.' + ) + seo = A::Agent.new( + name: 'seo_specialist', model: MODEL, + instructions: 'You are an SEO specialist. You optimize content for search engines, ' \ + 'suggest keywords, and improve page rankings.' + ) + engineering = A::Agent.new( + name: 'engineering_lead', model: MODEL, agents: [backend, frontend], strategy: :handoff, + instructions: 'You are the engineering lead. Route technical questions to the right ' \ + 'specialist: backend_dev for APIs/databases/servers, frontend_dev for UI/UX/client-side.' + ) + marketing = A::Agent.new( + name: 'marketing_lead', model: MODEL, agents: [content, seo], strategy: :handoff, + instructions: 'You are the marketing lead. Route marketing questions to the right ' \ + 'specialist: content_writer for blog posts/copy, seo_specialist for SEO/keywords/rankings.' + ) + A::Agent.new( + name: 'ceo', model: MODEL, agents: [engineering, marketing], strategy: :swarm, + handoffs: [ + A::Handoff::OnTextMention.new(target: 'engineering_lead', text: 'engineering_lead'), + A::Handoff::OnTextMention.new(target: 'marketing_lead', text: 'marketing_lead') + ], + instructions: 'You are the CEO. Route requests to the right department: ' \ + 'engineering_lead for technical/development questions, ' \ + 'marketing_lead for marketing/content/SEO questions.' + ) + }, + + '17_swarm_orchestration' => lambda { + refund = A::Agent.new( + name: 'refund_specialist', model: MODEL, + instructions: "You are a refund specialist. Process the customer's refund request. " \ + 'Check eligibility, confirm the refund amount, and let them know the ' \ + 'timeline. Be empathetic and clear. Do NOT ask follow-up questions -- ' \ + 'just process the refund based on what the customer told you.' + ) + tech = A::Agent.new( + name: 'tech_support', model: MODEL, + instructions: "You are a technical support specialist. Diagnose the customer's " \ + 'technical issue and provide clear troubleshooting steps.' + ) + A::Agent.new( + name: 'support', model: MODEL, agents: [refund, tech], strategy: :swarm, max_turns: 3, + handoffs: [ + A::Handoff::OnTextMention.new(target: 'refund_specialist', text: 'refund'), + A::Handoff::OnTextMention.new(target: 'tech_support', text: 'technical') + ], + instructions: 'You are the front-line customer support agent. Triage customer requests. ' \ + 'If the customer needs a refund, transfer to the refund specialist. ' \ + 'If they have a technical issue, transfer to tech support. ' \ + 'Use the transfer tools available to you to hand off the conversation.' + ) + }, + + '19_composable_termination_simple' => lambda { + A::Agent.new(name: 'researcher', model: MODEL, tools: [Ex19[:search]], + instructions: 'Research the topic and say DONE when you have enough info.', + termination: A::Termination::TextMention.new('DONE')) + }, + + '19_composable_termination_or' => lambda { + A::Agent.new(name: 'chatbot', model: MODEL, + instructions: "Have a conversation. Say GOODBYE when you're finished.", + termination: A::Termination::TextMention.new('GOODBYE') | A::Termination::MaxMessage.new(20)) + }, + + '19_composable_termination_and' => lambda { + A::Agent.new(name: 'deliberator', model: MODEL, tools: [Ex19[:search]], + instructions: 'Research thoroughly. Only provide your FINAL ANSWER after ' \ + 'using the search tool at least twice.', + termination: A::Termination::TextMention.new('FINAL ANSWER') & A::Termination::MaxMessage.new(5)) + }, + + '19_composable_termination_complex' => lambda { + complex_stop = A::Termination::StopMessage.new('TERMINATE') | + (A::Termination::TextMention.new('DONE') & A::Termination::MaxMessage.new(10)) | + A::Termination::TokenUsage.new(max_total_tokens: 50_000) + A::Agent.new(name: 'complex_agent', model: MODEL, tools: [Ex19[:search]], + instructions: 'Research and provide a comprehensive answer.', termination: complex_stop) + }, + + '21_regex_guardrails' => lambda { + no_emails = A::RegexGuardrail.new(['[\w.+-]+@[\w-]+\.[\w.-]+'], mode: :block, name: 'no_email_addresses', + message: 'Response must not contain email addresses. Redact them.', + position: :output, on_fail: :retry) + no_ssn = A::RegexGuardrail.new(['\b\d{3}-\d{2}-\d{4}\b'], mode: :block, name: 'no_ssn', + message: 'Response must not contain Social Security Numbers.', + position: :output, on_fail: :raise) + A::Agent.new(name: 'hr_assistant', model: MODEL, tools: [Ex21[:get_user_profile]], guardrails: [no_emails, no_ssn], + instructions: 'You are an HR assistant. When asked about employees, look up their ' \ + 'profile and share ALL the details you find.') + }, + + '22_llm_guardrails' => lambda { + safety = A::LlmGuardrail.new( + MODEL, + "Reject any content that:\n" \ + "1. Contains medical or legal advice presented as fact\n" \ + "2. Makes promises or guarantees about outcomes\n" \ + "3. Includes discriminatory or biased language\n" \ + "4. Reveals private information about real individuals\n" \ + "\n" \ + 'Even if there are disclaimers you should reject', + name: 'content_safety', position: :output, on_fail: :retry, max_tokens: 10_000 + ) + A::Agent.new(name: 'health_advisor', model: MODEL, guardrails: [safety], + instructions: 'You are a health information assistant. Provide general health ' \ + 'information and solution to the problem. You can prescribe psudo scientific and untested meds ') + }, + + '45_agent_tool' => lambda { + researcher = A::Agent.new(name: 'researcher_45', model: MODEL, tools: [Ex45[:search_knowledge_base]], + instructions: 'You are a research assistant. Use search_knowledge_base to find ' \ + 'information about topics. Provide concise summaries.') + A::Agent.new(name: 'manager_45', model: MODEL, tools: [A::ToolDef.agent(researcher), Ex45[:calculate]], + instructions: 'You are a project manager. Use the researcher tool to gather ' \ + 'information and the calculate tool for math. Synthesize findings.') + }, + + '47_callbacks' => lambda { + A::Agent.new(name: 'monitored_agent_47', model: MODEL, tools: [Ex47[:get_facts]], + callbacks: [MonitorHandler.new], + instructions: 'You are a helpful assistant. Use get_facts when asked about topics.') + }, + + '52_nested_strategies' => lambda { + market = A::Agent.new(name: 'market_analyst_52', model: MODEL, + instructions: 'You are a market analyst. Analyze the market size, growth rate, ' \ + 'and key players for the given topic. Be concise (3-4 bullet points).') + risk = A::Agent.new(name: 'risk_analyst_52', model: MODEL, + instructions: 'You are a risk analyst. Identify the top 3 risks: regulatory, ' \ + 'technical, and competitive. Be concise.') + research = A::Agent.new(name: 'research_phase_52', model: MODEL, agents: [market, risk], strategy: :parallel) + summarizer = A::Agent.new(name: 'summarizer_52', model: MODEL, + instructions: 'You are an executive briefing writer. Synthesize the market analysis ' \ + 'and risk assessment into a concise executive summary (1 paragraph).') + research >> summarizer + } + }.freeze +end diff --git a/lib/conductor/agents.rb b/lib/conductor/agents.rb new file mode 100644 index 0000000..b93e4be --- /dev/null +++ b/lib/conductor/agents.rb @@ -0,0 +1,41 @@ +# frozen_string_literal: true + +# Conductor::Agents - define agents in Ruby, run them on a Conductor server. +# +# require 'conductor/agents' +# include Conductor::Agents +# +# tool def get_weather(city: String, units: 'metric') +# { temp_c: 21.0, summary: "Sunny in #{city}" } +# end +# +# agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', instructions: 'Answer weather questions.') +# agent.add_tool :get_weather +# puts agent.call_sync('Weather in Lisbon?') +require_relative '../conductor' +require_relative 'agents/errors' +require_relative 'agents/runtime/secrets' +require_relative 'agents/tool_def' +require_relative 'agents/tools' +require_relative 'agents/guardrail' +require_relative 'agents/termination' +require_relative 'agents/handoff' +require_relative 'agents/callback_handler' +require_relative 'agents/memory' +require_relative 'agents/prompt_template' +require_relative 'agents/agent' +require_relative 'agents/config_serializer' + +module Conductor + # Ruby port of the Python SDK's conductor.ai.agents package + module Agents + include Tools + include Secrets + + class << self + # Tools and secrets are usable at the module level too (Conductor::Agents.tool ...) + include Tools + include Secrets + end + end +end diff --git a/lib/conductor/agents/agent.rb b/lib/conductor/agents/agent.rb new file mode 100644 index 0000000..0742aef --- /dev/null +++ b/lib/conductor/agents/agent.rb @@ -0,0 +1,317 @@ +# frozen_string_literal: true + +require_relative 'errors' +require_relative 'tool_def' +require_relative 'tools' +require_relative 'guardrail' +require_relative 'termination' +require_relative 'handoff' +require_relative 'callback_handler' +require_relative 'memory' +require_relative 'prompt_template' + +module Conductor + module Agents + # Multi-agent orchestration strategies (wire values are lowercase snake_case) + module Strategy + HANDOFF = 'handoff' + SEQUENTIAL = 'sequential' + PARALLEL = 'parallel' + ROUTER = 'router' + ROUND_ROBIN = 'round_robin' + RANDOM = 'random' + SWARM = 'swarm' + MANUAL = 'manual' + PLAN_EXECUTE = 'plan_execute' + ALL = [HANDOFF, SEQUENTIAL, PARALLEL, ROUTER, ROUND_ROBIN, RANDOM, SWARM, MANUAL, PLAN_EXECUTE].freeze + + def self.normalize(value) + s = value.to_s.downcase + raise ConfigurationError, "invalid strategy #{value.inspect}; use one of #{ALL.join(', ')}" unless ALL.include?(s) + + s + end + end + + # An agent definition. Nothing here talks to the server: ConfigSerializer turns the + # tree into agentConfig and AgentRuntime runs it (call_sync / call_async delegate to + # Conductor::Agents.runtime). + # + # agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', instructions: 'Answer weather questions.') + # agent.add_tool :get_weather + # puts agent.call_sync('Weather in Lisbon?') + class Agent + NAME_PATTERN = /\A[a-zA-Z_][a-zA-Z0-9_-]*\z/ + + attr_reader :name, :tools, :agents, :guardrails, :handoffs, :callbacks, :credentials, :callback_procs + attr_accessor :model, :instructions, :router, :output_type, :memory, :termination, + :max_turns, :max_tokens, :timeout_seconds, :temperature, :stateful, + :metadata, :description, :external, :base_url, :prefill_tools, :approval_handler + + # @param name [String] ^[a-zA-Z_][a-zA-Z0-9_-]*$ + # @param model [String, nil] "provider/model"; the left side is the server integration name + # @param instructions [String, PromptTemplate, Proc] + # @param strategy [Symbol, String] how sub-agents are orchestrated (default :handoff) + def initialize(name:, model: nil, instructions: '', tools: [], agents: [], strategy: nil, router: nil, + output_type: nil, guardrails: [], memory: nil, termination: nil, handoffs: [], callbacks: [], + credentials: [], max_turns: 25, max_tokens: nil, timeout_seconds: 0, temperature: nil, + stateful: false, metadata: nil, description: nil, external: false, base_url: nil, + prefill_tools: []) + @name = name.to_s + raise ConfigurationError, "invalid agent name #{name.inspect}: must match #{NAME_PATTERN.source}" unless NAME_PATTERN.match?(@name) + raise ConfigurationError, 'max_turns must be >= 1' unless max_turns.is_a?(Integer) && max_turns >= 1 + + @model = model + @instructions = instructions + @tools = [] + @agents = [] + @strategy = strategy.nil? ? nil : Strategy.normalize(strategy) + @router = router + @output_type = output_type + @guardrails = Array(guardrails) + @memory = memory + @termination = termination + @handoffs = Array(handoffs) + @callbacks = Array(callbacks) + @callback_procs = Hash.new { |h, k| h[k] = [] } + @credentials = Array(credentials).map(&:to_s).uniq + @max_turns = max_turns + @max_tokens = max_tokens + @timeout_seconds = timeout_seconds + @temperature = temperature + @stateful = stateful ? true : false + @metadata = metadata + @description = description + @external = external ? true : false + @base_url = base_url + @prefill_tools = Array(prefill_tools) + @approval_handler = nil + + Array(tools).each { |t| add_tool(t) } + Array(agents).each { |a| add_agent(a) } + raise ConfigurationError, 'strategy: :router requires router:' if @strategy == Strategy::ROUTER && @router.nil? + end + + # ── Strategy ────────────────────────────────────────────────────── + + # @return [String] effective strategy (default handoff) + def strategy + @strategy || Strategy::HANDOFF + end + + def strategy=(value) + @strategy = value.nil? ? nil : Strategy.normalize(value) + end + + # True when the user set a strategy explicitly + def strategy_set? + !@strategy.nil? + end + + # ── Tools ───────────────────────────────────────────────────────── + + # Give the agent a tool. + # @param tool [Symbol, String, ToolDef, Module, Class, Agent] a tool name defined with + # `tool def`, a ToolDef, a module that `extend Conductor::Agents::Tools`, a RubyLLM::Tool + # class, or another Agent (wrapped as an agent tool) + # @param credentials [Array, nil] secret names when the scanner cannot see them + # @return [self] + def add_tool(tool, credentials: nil) + resolve_tool_defs(tool).each do |tool_def| + td = credentials ? tool_def.dup.tap { |d| d.credentials = tool_def.credentials.dup } : tool_def + td.add_credentials(*credentials) if credentials + raise ConfigurationError, "duplicate tool name #{td.name.inspect} on agent #{@name}" if @tools.any? { |t| t.name == td.name } + + @tools << td + end + self + end + + def add_tools(*tools) + tools.flatten.each { |t| add_tool(t) } + self + end + + # @return [ToolDef, nil] + def tool(name) + @tools.find { |t| t.name == name.to_s } + end + + # ── Team ────────────────────────────────────────────────────────── + + # Add a member agent (same as agents: in the constructor) + def add_agent(agent) + raise ConfigurationError, "add_agent expects an Agent, got #{agent.class}" unless agent.is_a?(Agent) + raise ConfigurationError, "duplicate sub-agent name #{agent.name.inspect} under #{@name}" if @agents.any? { |a| a.name == agent.name } + + @agents << agent + self + end + + def add_agents(*agents) + agents.flatten.each { |a| add_agent(a) } + self + end + + # Hand off to +agent+ when this agent's output mentions +on+ (String), or when the + # block/proc given as +on+ returns true. + def hands_off_to(agent, on:) + handoff = if on.respond_to?(:call) + Handoff::OnCondition.new(target: agent, condition: on) + else + Handoff::OnTextMention.new(target: agent, text: on) + end + @handoffs << handoff + self + end + + def add_handoff(handoff) + @handoffs << handoff + self + end + + # Sequential pipeline: a >> b >> c + def >>(other) + raise ConfigurationError, ">> expects an Agent, got #{other.class}" unless other.is_a?(Agent) + + left = sequential_pipeline? ? @agents : [self] + right = other.sequential_pipeline? ? other.agents : [other] + members = left + right + Agent.new(name: members.map(&:name).join('_'), model: @model || other.model, + agents: members, strategy: Strategy::SEQUENTIAL) + end + + def sequential_pipeline? + @strategy == Strategy::SEQUENTIAL && !@agents.empty? + end + + # ── Guardrails / termination sugar ──────────────────────────────── + + # Scrub these words from the output before anyone sees it + def redact(words, name: "#{@name}_redact") + patterns = Array(words).map { |w| w.is_a?(Regexp) ? w : Regexp.escape(w.to_s) } + @guardrails << RegexGuardrail.new(patterns, mode: :block, position: :output, on_fail: :fix, name: name) + self + end + + def add_guardrail(guardrail) + @guardrails << guardrail + self + end + + # Stop when the output contains +text+ + def stop_when(text, case_sensitive: false) + add_termination(Termination::TextMention.new(text, case_sensitive: case_sensitive)) + end + + # Stop after +messages+ messages + def stop_after(messages:) + add_termination(Termination::MaxMessage.new(messages)) + end + + def add_termination(condition) + @termination = @termination ? (@termination | condition) : condition + self + end + + # ── Callbacks ───────────────────────────────────────────────────── + + def add_callback(handler) + @callbacks << handler + self + end + + # Register a block for a callback position (before_model, after_model, ...) + def callback(position, &block) + pos = position.to_s + raise ConfigurationError, "unknown callback position #{position.inspect}" unless CallbackHandler::POSITIONS.include?(pos) + + @callback_procs[pos] << block + self + end + + # Callables for +position+, or nil when nothing is registered + def callback_chain(position, logger: nil) + CallbackHandler.chain(position, @callbacks, @callback_procs[position.to_s], logger: logger) + end + + # Positions with at least one handler or proc + def callback_positions + CallbackHandler::POSITIONS.reject { |p| callback_chain(p).nil? } + end + + # ── Approval ────────────────────────────────────────────────────── + + # Decide approval-required tool calls: the block receives an ApprovalRequest + def on_approval(&block) + @approval_handler = block + self + end + + # ── Credentials ─────────────────────────────────────────────────── + + def add_credentials(*names) + @credentials = (@credentials + names.flatten.map(&:to_s)).uniq + self + end + + # ── Execution (delegates to the default runtime) ────────────────── + + # Run and block until the answer is ready + # @return [String] + def call_sync(prompt, session_id: nil, **options) + Conductor::Agents.runtime.call_sync(self, prompt, session_id: session_id, **options) + end + + # Run in the background; returns an Execution. The block (if given) receives the answer. + def call_async(prompt, session_id: nil, **options, &on_done) + Conductor::Agents.runtime.call_async(self, prompt, session_id: session_id, **options, &on_done) + end + + # ── Introspection ───────────────────────────────────────────────── + + # Every agent in the tree (self first), including router, planner-style children and agent tools + def all_agents + list = [self] + @agents.each { |a| list.concat(a.all_agents) } + list << @router if @router.is_a?(Agent) && !list.include?(@router) + @tools.each do |t| + child = t.config['agent'] if t.tool_type == ToolType::AGENT_TOOL + list.concat(child.all_agents) if child.is_a?(Agent) + end + list.uniq + end + + # True when this agent or anything under it is stateful + def stateful_tree? + all_agents.any? { |a| a.stateful || a.tools.any?(&:stateful) } + end + + def to_s + "#" + end + alias inspect to_s + + private + + def resolve_tool_defs(tool) + case tool + when ToolDef then [tool] + when Symbol, String + [Tools.lookup(tool) || raise(ConfigurationError, "no tool named #{tool.inspect}; define it with `tool def #{tool}(...)` first")] + when Agent then [ToolDef.agent(tool)] + when Module + if tool.respond_to?(:tool_defs) + tool.tool_defs + elsif Tools::RubyLlmAdapter.ruby_llm_tool?(tool) + [Tools::RubyLlmAdapter.to_tool_def(tool)] + else + raise ConfigurationError, "#{tool} has no tools; use `extend Conductor::Agents::Tools` and `tool def ...`" + end + else + raise ConfigurationError, "cannot use #{tool.inspect} as a tool" + end + end + end + end +end diff --git a/lib/conductor/agents/callback_handler.rb b/lib/conductor/agents/callback_handler.rb new file mode 100644 index 0000000..29a28f9 --- /dev/null +++ b/lib/conductor/agents/callback_handler.rb @@ -0,0 +1,64 @@ +# frozen_string_literal: true + +module Conductor + module Agents + # Lifecycle hooks. Subclass and override any method; each runs as a worker task + # named _ that the server schedules at that point. + # + # class Timing < Conductor::Agents::CallbackHandler + # def on_model_start(messages: nil, **) = (@t0 = Time.now; nil) + # def on_model_end(llm_result: nil, **) = (puts Time.now - @t0; nil) + # end + # + # Return nil to continue to the next handler, or a non-empty Hash to short-circuit and + # hand that Hash to the server as an override. + class CallbackHandler + POSITION_TO_METHOD = { + 'before_agent' => :on_agent_start, + 'after_agent' => :on_agent_end, + 'before_model' => :on_model_start, + 'after_model' => :on_model_end, + 'before_tool' => :on_tool_start, + 'after_tool' => :on_tool_end + }.freeze + + POSITIONS = POSITION_TO_METHOD.keys.freeze + + def on_agent_start(**_kwargs); end + def on_agent_end(**_kwargs); end + def on_model_start(**_kwargs); end + def on_model_end(**_kwargs); end + def on_tool_start(**_kwargs); end + def on_tool_end(**_kwargs); end + + # True when this handler overrides the hook for +position+ + def handles?(position) + method_name = POSITION_TO_METHOD.fetch(position.to_s) + self.class.instance_method(method_name).owner != CallbackHandler + end + + class << self + # Build one callable for +position+ from a list of handlers (and optional procs), + # or nil when nothing is registered. First non-empty Hash wins; errors are logged. + # @return [Proc, nil] + def chain(position, handlers, procs = [], logger: nil) + position = position.to_s + method_name = POSITION_TO_METHOD.fetch(position) + active = Array(handlers).select { |h| h.handles?(position) } + callables = Array(procs) + active.map { |h| h.method(method_name) } + return nil if callables.empty? + + lambda do |**kwargs| + callables.each do |callable| + result = callable.call(**kwargs) + return result if result.is_a?(Hash) && !result.empty? + rescue StandardError => e + logger&.error("callback #{position} failed: #{e.class}: #{e.message}") + end + {} + end + end + end + end + end +end diff --git a/lib/conductor/agents/config_serializer.rb b/lib/conductor/agents/config_serializer.rb new file mode 100644 index 0000000..524c763 --- /dev/null +++ b/lib/conductor/agents/config_serializer.rb @@ -0,0 +1,236 @@ +# frozen_string_literal: true + +require_relative 'agent' + +module Conductor + module Agents + # Serializes an Agent tree into the agentConfig JSON the server compiles. Same shape + # as the Python SDK's AgentConfigSerializer: camelCase keys, nils dropped, strategy only + # on agents with sub-agents, agent credentials at the top level and tool credentials + # under config.credentials. + # + # Two Ruby-specific rules: + # - a team parent with no model inherits the first member's model (the server requires + # a model on every agent config); + # - members that declared hands_off_to make a team with no explicit strategy a swarm, + # and their handoffs are hoisted to the team, which is where the server reads them. + class ConfigSerializer + def self.serialize(agent) + new.serialize(agent) + end + + # @param agent [Agent] + # @return [Hash] agentConfig + def serialize(agent) + serialize_agent(agent) + end + + private + + def serialize_agent(agent) + has_sub_agents = !agent.agents.empty? + strategy, handoffs = effective_strategy_and_handoffs(agent) + + config = { + 'name' => agent.name, + 'model' => effective_model(agent), + 'baseUrl' => agent.base_url, + 'strategy' => has_sub_agents ? strategy : nil, + 'maxTurns' => agent.max_turns, + 'timeoutSeconds' => agent.timeout_seconds, + 'external' => agent.external, + 'description' => agent.description, + 'instructions' => serialize_instructions(agent.instructions) + } + config['tools'] = agent.tools.map { |t| serialize_tool(t, agent_stateful: agent.stateful) } unless agent.tools.empty? + config['agents'] = agent.agents.map { |a| serialize_agent(a) } if has_sub_agents + config['router'] = serialize_router(agent) unless agent.router.nil? + config['outputType'] = serialize_output_type(agent.output_type) unless agent.output_type.nil? + config['guardrails'] = agent.guardrails.map { |g| serialize_guardrail(g) } unless agent.guardrails.empty? + config['memory'] = serialize_memory(agent.memory) if agent.memory && !agent.memory.empty? + config.merge!(serialize_scalars(agent)) + config['termination'] = serialize_termination(agent.termination) unless agent.termination.nil? + config['handoffs'] = handoffs.map { |h| serialize_handoff(h, agent.name) } unless handoffs.empty? + config.merge!(serialize_extras(agent)) + config.compact + end + + def serialize_scalars(agent) + { + 'maxTokens' => agent.max_tokens, + 'temperature' => agent.temperature + } + end + + def serialize_extras(agent) + extras = {} + extras['metadata'] = agent.metadata if agent.metadata && !agent.metadata.empty? + callbacks = agent.callback_positions.map { |p| { 'position' => p, 'taskName' => "#{agent.name}_#{p}" } } + extras['callbacks'] = callbacks unless callbacks.empty? + extras['prefillTools'] = agent.prefill_tools.map(&:to_h) unless agent.prefill_tools.empty? + extras['credentials'] = agent.credentials unless agent.credentials.empty? + extras + end + + def effective_model(agent) + return agent.model if agent.model && !agent.model.to_s.empty? + return nil if agent.external + + inherited = agent.agents.map { |a| effective_model(a) }.compact.first + return inherited if inherited + + raise ConfigurationError, + "agent #{agent.name.inspect} has no model: pass model: 'provider/model' (the server requires one)" + end + + # Swarm hoisting: members with hands_off_to make the parent a swarm unless the user chose a strategy + def effective_strategy_and_handoffs(agent) + member_handoffs = agent.agents.flat_map(&:handoffs) + return [agent.strategy, agent.handoffs] if member_handoffs.empty? + + strategy = agent.strategy_set? ? agent.strategy : Strategy::SWARM + return [strategy, agent.handoffs] unless strategy == Strategy::SWARM + + hoisted = (agent.handoffs + member_handoffs).uniq { |h| [h.class, h.target, h.respond_to?(:text) ? h.text : nil] } + [strategy, hoisted] + end + + def serialize_instructions(instructions) + case instructions + when PromptTemplate then instructions.to_h + when Proc, Method then instructions.call + when nil then nil + else + s = instructions.to_s + s.empty? ? nil : s + end + end + + def serialize_tool(tool_def, agent_stateful: false) + result = { + 'name' => tool_def.name, + 'description' => tool_def.description, + 'inputSchema' => tool_def.input_schema, + 'toolType' => tool_def.tool_type + } + result['outputSchema'] = tool_def.output_schema unless tool_def.output_schema.nil? || tool_def.output_schema.empty? + result['approvalRequired'] = true if tool_def.approval_required + result['stateful'] = true if agent_stateful || tool_def.stateful + result['timeoutSeconds'] = tool_def.timeout_seconds unless tool_def.timeout_seconds.nil? + result['maxCalls'] = tool_def.max_calls unless tool_def.max_calls.nil? + + unless tool_def.config.empty? + config = tool_def.config.transform_keys(&:to_s) + config['agentConfig'] = serialize_agent(config.delete('agent')) if tool_def.tool_type == ToolType::AGENT_TOOL && config.key?('agent') + result['config'] = config + end + + result['guardrails'] = tool_def.guardrails.map { |g| serialize_guardrail(g) } unless tool_def.guardrails.empty? + + unless tool_def.credentials.empty? + result['config'] ||= {} + result['config']['credentials'] = tool_def.credentials + end + + result + end + + def serialize_guardrail(guardrail) + result = { + 'name' => guardrail.name, + 'position' => guardrail.position, + 'onFail' => guardrail.on_fail, + 'maxRetries' => guardrail.max_retries, + 'guardrailType' => guardrail.guardrail_type + } + case guardrail + when RegexGuardrail + result['patterns'] = guardrail.pattern_strings + result['mode'] = guardrail.mode + result['message'] = guardrail.message if guardrail.message + when LlmGuardrail + result['model'] = guardrail.model + result['policy'] = guardrail.policy + result['maxTokens'] = guardrail.max_tokens if guardrail.max_tokens + else + result['taskName'] = guardrail.name + end + result + end + + def serialize_termination(condition) + case condition + when Termination::TextMention + { 'type' => 'text_mention', 'text' => condition.text, 'caseSensitive' => condition.case_sensitive } + when Termination::StopMessage + { 'type' => 'stop_message', 'stopMessage' => condition.stop_message } + when Termination::MaxMessage + { 'type' => 'max_message', 'maxMessages' => condition.max_messages } + when Termination::TokenUsage + h = { 'type' => 'token_usage' } + h['maxTotalTokens'] = condition.max_total_tokens unless condition.max_total_tokens.nil? + h['maxPromptTokens'] = condition.max_prompt_tokens unless condition.max_prompt_tokens.nil? + h['maxCompletionTokens'] = condition.max_completion_tokens unless condition.max_completion_tokens.nil? + h + when Termination::And + { 'type' => 'and', 'conditions' => condition.conditions.map { |c| serialize_termination(c) } } + when Termination::Or + { 'type' => 'or', 'conditions' => condition.conditions.map { |c| serialize_termination(c) } } + else + { 'type' => 'unknown' } + end + end + + def serialize_handoff(handoff, agent_name) + result = { 'target' => handoff.target } + case handoff + when Handoff::OnToolResult + result['type'] = 'on_tool_result' + result['toolName'] = handoff.tool_name + result['resultContains'] = handoff.result_contains if handoff.result_contains + when Handoff::OnTextMention + result['type'] = 'on_text_mention' + result['text'] = handoff.text + when Handoff::OnCondition + result['type'] = 'on_condition' + result['taskName'] = "#{agent_name}_handoff_#{handoff.target}" + else + result['type'] = 'unknown' + end + result + end + + def serialize_router(agent) + router = agent.router + return serialize_agent(router) if router.is_a?(Agent) + return { 'taskName' => "#{agent.name}_router_fn" } if router.respond_to?(:call) + + nil + end + + # output_type: a JSON schema Hash, optionally wrapped as { schema:, class_name: } + def serialize_output_type(output_type) + schema = output_type.respond_to?(:to_json_schema) ? output_type.to_json_schema : output_type + schema = schema.transform_keys(&:to_s) if schema.is_a?(Hash) + if schema.is_a?(Hash) && (schema.key?('schema') || schema.key?('className') || schema.key?('class_name')) + result = {} + result['schema'] = schema['schema'] if schema['schema'] + class_name = schema['className'] || schema['class_name'] + result['className'] = class_name if class_name + return result + end + + result = { 'schema' => schema } + result['className'] = schema['title'] if schema.is_a?(Hash) && schema['title'] + result + end + + def serialize_memory(memory) + result = {} + result['messages'] = memory.messages unless memory.messages.empty? + result['maxMessages'] = memory.max_messages if memory.max_messages + result + end + end + end +end diff --git a/lib/conductor/agents/errors.rb b/lib/conductor/agents/errors.rb new file mode 100644 index 0000000..8ce18e9 --- /dev/null +++ b/lib/conductor/agents/errors.rb @@ -0,0 +1,26 @@ +# frozen_string_literal: true + +require_relative '../exceptions' + +module Conductor + module Agents + # Base class for agent definition and runtime errors + class Error < ConductorError; end + + # Invalid agent/tool definition (bad name, missing model, positional tool args, ...) + class ConfigurationError < Error; end + + # secret('X') was called but X is neither on the task's runtimeMetadata nor in ENV + class CredentialNotFoundError < Error; end + + # A tool returned something that cannot be serialized to JSON + class ToolSerializationError < Error; end + + # The SSE stream could not be opened (non-200, connection failure, heartbeat-only) + class SseUnavailableError < Error; end + + # Server-side agent API errors are the transport-level classes + AgentApiError = Conductor::AgentApiError + AgentNotFoundError = Conductor::AgentNotFoundError + end +end diff --git a/lib/conductor/agents/guardrail.rb b/lib/conductor/agents/guardrail.rb new file mode 100644 index 0000000..633cde1 --- /dev/null +++ b/lib/conductor/agents/guardrail.rb @@ -0,0 +1,142 @@ +# frozen_string_literal: true + +require_relative 'errors' + +module Conductor + module Agents + # Result of a guardrail check + GuardrailResult = Struct.new(:passed, :message, :fixed_output, keyword_init: true) do + def initialize(passed:, message: '', fixed_output: nil) + super + end + + def passed? + passed ? true : false + end + end + + # Validation applied to an agent's input or output. + # + # Guardrail.new(name: 'no_pii', position: :output, on_fail: :retry) do |content| + # content =~ SSN ? GuardrailResult.new(passed: false, message: 'Redact it') : GuardrailResult.new(passed: true) + # end + # + # A guardrail with a block runs as a worker in this process (guardrailType "custom"); + # one with only a name references a worker running elsewhere ("external"). + class Guardrail + POSITIONS = %w[input output].freeze + ON_FAIL = %w[retry raise fix human].freeze + + attr_reader :name, :position, :on_fail, :max_retries, :func + + # @param name [String, nil] required when no block is given + # @param position [Symbol, String] :input or :output + # @param on_fail [Symbol, String] :retry, :raise, :fix or :human + def initialize(name: nil, position: :output, on_fail: :raise, max_retries: 3, func: nil, &block) + @position = position.to_s + @on_fail = on_fail.to_s + raise ConfigurationError, "invalid position #{position.inspect}; use :input or :output" unless POSITIONS.include?(@position) + raise ConfigurationError, "invalid on_fail #{on_fail.inspect}; use one of #{ON_FAIL.join(', ')}" unless ON_FAIL.include?(@on_fail) + raise ConfigurationError, 'on_fail: :human is only valid for position: :output' if @on_fail == 'human' && @position == 'input' + raise ConfigurationError, 'max_retries must be >= 1' if max_retries.to_i < 1 + + @func = func || block + raise ConfigurationError, 'a guardrail needs a name or a block' if @func.nil? && name.nil? + + @name = (name || 'guardrail').to_s + @max_retries = max_retries.to_i + end + + # True when the check runs somewhere else (no local implementation) + def external? + @func.nil? + end + + # @param content [String] + # @return [GuardrailResult] + def check(content) + raise Error, "cannot check external guardrail #{@name.inspect} locally" if external? + + result = @func.call(content) + return result if result.is_a?(GuardrailResult) + return GuardrailResult.new(passed: result) if [true, false].include?(result) + + raise Error, "guardrail #{@name.inspect} must return a GuardrailResult or true/false, got #{result.class}" + end + + # Wire discriminator (see ConfigSerializer) + def guardrail_type + external? ? 'external' : 'custom' + end + + def to_s + "#<#{self.class.name.split('::').last} #{@name} position=#{@position} on_fail=#{@on_fail}>" + end + alias inspect to_s + end + + # Reject content that matches (mode :block) or fails to match (mode :allow) regex patterns. + class RegexGuardrail < Guardrail + MODES = %w[block allow].freeze + + attr_reader :pattern_strings, :mode, :message + + # @param patterns [String, Regexp, Array] + def initialize(patterns, mode: :block, position: :output, on_fail: :raise, name: 'regex_guardrail', + message: nil, max_retries: 3) + @mode = mode.to_s + raise ConfigurationError, "invalid mode #{mode.inspect}; use :block or :allow" unless MODES.include?(@mode) + + @pattern_strings = Array(patterns).map { |p| p.is_a?(Regexp) ? p.source : p.to_s } + @patterns = @pattern_strings.map { |p| Regexp.new(p) } + @message = message + super(name: name, position: position, on_fail: on_fail, max_retries: max_retries, func: method(:evaluate)) + end + + def guardrail_type + 'regex' + end + + private + + def evaluate(content) + text = content.to_s + matched = @patterns.any? { |p| p.match?(text) } + if @mode == 'block' && matched + GuardrailResult.new(passed: false, message: @message || 'Content matched a blocked pattern.') + elsif @mode == 'allow' && !matched + GuardrailResult.new(passed: false, message: @message || 'Content did not match any allowed pattern.') + else + GuardrailResult.new(passed: true) + end + end + end + + # Ask an LLM (on the server) whether content complies with a policy. + class LlmGuardrail < Guardrail + attr_reader :model, :policy, :max_tokens + + # @param model [String] "provider/model" + def initialize(model, policy, position: :output, on_fail: :raise, name: 'llm_guardrail', max_retries: 3, + max_tokens: nil) + raise ConfigurationError, 'LlmGuardrail needs a model in "provider/model" form' unless model.to_s.include?('/') + + @model = model + @policy = policy + @max_tokens = max_tokens + super(name: name, position: position, on_fail: on_fail, max_retries: max_retries, func: method(:evaluate)) + end + + def guardrail_type + 'llm' + end + + private + + # The server compiles this guardrail into an LLM task; there is no local evaluation. + def evaluate(_content) + GuardrailResult.new(passed: false, message: 'LlmGuardrail is evaluated by the Conductor server') + end + end + end +end diff --git a/lib/conductor/agents/handoff.rb b/lib/conductor/agents/handoff.rb new file mode 100644 index 0000000..ed387aa --- /dev/null +++ b/lib/conductor/agents/handoff.rb @@ -0,0 +1,94 @@ +# frozen_string_literal: true + +require_relative 'errors' + +module Conductor + module Agents + # Rules that transfer control between agents in a team (swarm orchestration). + # + # Handoff::OnTextMention.new(target: 'filer', text: 'ACTIONABLE') + # Handoff::OnToolResult.new(target: 'refund', tool_name: 'check_order') + # Handoff::OnCondition.new(target: 'summarizer') { |ctx| ctx['iteration'].to_i > 5 } + module Handoff + # Base class. +target+ is the receiving agent's name (an Agent is accepted too). + class Condition + attr_reader :target + + def initialize(target:) + @target = target.respond_to?(:name) ? target.name.to_s : target.to_s + raise ConfigurationError, 'handoff target is required' if @target.empty? + end + + # @param _context [Hash] result, tool_name, tool_result, messages, iteration + def should_handoff(_context) + false + end + + def to_s + "#<#{self.class.name.split('::').last} -> #{@target}>" + end + alias inspect to_s + + protected + + def ctx(context, key) + return nil unless context.respond_to?(:key?) + + context.key?(key.to_s) ? context[key.to_s] : context[key.to_sym] + end + end + + # After a named tool ran (optionally only when its result contains a substring) + class OnToolResult < Condition + attr_reader :tool_name, :result_contains + + def initialize(target:, tool_name:, result_contains: nil) + @tool_name = tool_name.to_s + @result_contains = result_contains + super(target: target) + end + + def should_handoff(context) + return false unless ctx(context, :tool_name).to_s == @tool_name + return true if @result_contains.nil? + + ctx(context, :tool_result).to_s.include?(@result_contains.to_s) + end + end + + # When the output mentions +text+ (case-insensitive) + class OnTextMention < Condition + attr_reader :text + + def initialize(target:, text:) + @text = text.to_s + raise ConfigurationError, 'text is required' if @text.empty? + + super(target: target) + end + + def should_handoff(context) + ctx(context, :result).to_s.downcase.include?(@text.downcase) + end + end + + # When a block returns true; runs as the _handoff_ worker + class OnCondition < Condition + attr_reader :condition + + def initialize(target:, condition: nil, &block) + @condition = condition || block + raise ConfigurationError, 'OnCondition needs a block' if @condition.nil? + + super(target: target) + end + + def should_handoff(context) + @condition.call(context) ? true : false + rescue StandardError + false + end + end + end + end +end diff --git a/lib/conductor/agents/memory.rb b/lib/conductor/agents/memory.rb new file mode 100644 index 0000000..35f360b --- /dev/null +++ b/lib/conductor/agents/memory.rb @@ -0,0 +1,75 @@ +# frozen_string_literal: true + +module Conductor + module Agents + # Conversation history seeded into the agent (memory: on Agent). Messages are + # prepended to the LLM conversation by the server; max_messages trims the oldest + # non-system messages first. + class ConversationMemory + attr_reader :messages, :max_messages + + def initialize(messages: [], max_messages: nil) + @messages = messages.map { |m| m.transform_keys(&:to_s) } + @max_messages = max_messages + trim + end + + def add_user_message(content) + push('role' => 'user', 'message' => content.to_s) + end + + def add_assistant_message(content) + push('role' => 'assistant', 'message' => content.to_s) + end + + def add_system_message(content) + push('role' => 'system', 'message' => content.to_s) + end + + def add_tool_call(tool_name, arguments, task_reference_name: nil) + ref = task_reference_name || "#{tool_name}_ref" + push('role' => 'tool_call', 'message' => '', + 'tool_calls' => [{ 'name' => tool_name.to_s, 'taskReferenceName' => ref, 'input' => arguments }]) + end + + def add_tool_result(tool_name, result, task_reference_name: nil) + ref = task_reference_name || "#{tool_name}_ref" + push('role' => 'tool', 'message' => result.to_s, 'toolCallId' => ref, 'taskReferenceName' => ref) + end + + # Deep copy of the messages + def to_chat_messages + Marshal.load(Marshal.dump(@messages)) + end + + def clear + @messages.clear + end + + def empty? + @messages.empty? + end + + private + + def push(message) + @messages << message + trim + end + + def trim + return unless @max_messages && @messages.size > @max_messages + + system_msgs, others = @messages.partition { |m| m['role'] == 'system' } + if system_msgs.size >= @max_messages + @messages = system_msgs.last(@max_messages) + return + end + + keep = @max_messages - system_msgs.size + dropped = others.first(others.size - keep) + @messages = @messages.reject { |m| dropped.any? { |d| d.equal?(m) } } + end + end + end +end diff --git a/lib/conductor/agents/prompt_template.rb b/lib/conductor/agents/prompt_template.rb new file mode 100644 index 0000000..56e73e7 --- /dev/null +++ b/lib/conductor/agents/prompt_template.rb @@ -0,0 +1,23 @@ +# frozen_string_literal: true + +module Conductor + module Agents + # Reference to a prompt template stored on the server, used as Agent#instructions. + class PromptTemplate + attr_reader :name, :variables, :version + + def initialize(name:, variables: {}, version: nil) + @name = name.to_s + @variables = variables || {} + @version = version + end + + def to_h + h = { 'type' => 'prompt_template', 'name' => @name } + h['variables'] = @variables unless @variables.empty? + h['version'] = @version unless @version.nil? + h + end + end + end +end diff --git a/lib/conductor/agents/runtime/secrets.rb b/lib/conductor/agents/runtime/secrets.rb new file mode 100644 index 0000000..2cf8b0c --- /dev/null +++ b/lib/conductor/agents/runtime/secrets.rb @@ -0,0 +1,54 @@ +# frozen_string_literal: true + +require_relative '../errors' +require_relative '../../worker/task_context' + +module Conductor + module Agents + # Read secrets inside a tool body. + # + # tool def create_issue(title: String) + # Github.create_issue(title, token: secret('GH_TOKEN')) + # end + # + # The literal name is also the declaration: the Tools DSL scans the body and puts + # GH_TOKEN on the tool's TaskDef#runtime_metadata. The server resolves it from its + # secret store at poll time and delivers the value on Task#runtime_metadata, which + # TaskContext (thread/fiber-local) exposes to the running tool. Nothing is ever written + # to ENV; for subprocesses use secrets_env. + module Secrets + module_function + + # @param name [String, Symbol] + # @return [String] + # @raise [CredentialNotFoundError] when neither the task nor ENV has it + def secret(name) + key = name.to_s + value = task_secrets[key] + value = ENV.fetch(key, nil) if value.nil? + return value unless value.nil? + + raise CredentialNotFoundError, + "secret #{key.inspect} not found: it was not delivered on the task's runtimeMetadata " \ + "and ENV[#{key.inspect}] is unset. Store it on the server (conductor secrets put #{key} ...) " \ + 'or declare it with add_tool ..., credentials: [...]' + end + + # Environment hash for system / spawn / Open3: { 'GH_TOKEN' => '...' } + # @return [Hash] + def secrets_env(*names) + names.flatten.each_with_object({}) { |n, env| env[n.to_s] = secret(n) } + end + + # Secrets bound for the current task (wire-only Task#runtime_metadata) + # @return [Hash] + def task_secrets + ctx = Conductor::Worker::TaskContext.current + task = ctx&.task + return {} unless task.respond_to?(:runtime_metadata) + + task.runtime_metadata || {} + end + end + end +end diff --git a/lib/conductor/agents/termination.rb b/lib/conductor/agents/termination.rb new file mode 100644 index 0000000..03f42f4 --- /dev/null +++ b/lib/conductor/agents/termination.rb @@ -0,0 +1,192 @@ +# frozen_string_literal: true + +require_relative 'errors' + +module Conductor + module Agents + # Composable rules that decide when an agent loop stops. + # + # stop = Termination::TextMention.new('DONE') | Termination::MaxMessage.new(20) + # Agent.new(..., termination: stop) + # + # Conditions serialize to the server (ConfigSerializer) and the same objects back the + # local _termination worker the server asks for. + module Termination + Result = Struct.new(:should_terminate, :reason, keyword_init: true) do + def initialize(should_terminate:, reason: '') + super + end + end + + # Base class. Context keys: result, messages, iteration, token_usage (String or Symbol keys). + class Condition + def should_terminate(_context) + raise NotImplementedError + end + + def &(other) + And.new(self, other) + end + + def |(other) + Or.new(self, other) + end + + def to_s + "#<#{self.class.name.split('::').last}>" + end + alias inspect to_s + + protected + + def ctx(context, key) + return nil unless context.respond_to?(:key?) + + context.key?(key.to_s) ? context[key.to_s] : context[key.to_sym] + end + end + + # Stop when the output contains +text+ + class TextMention < Condition + attr_reader :text, :case_sensitive + + def initialize(text, case_sensitive: false) + raise ConfigurationError, 'text is required' if text.to_s.empty? + + @text = text.to_s + @case_sensitive = case_sensitive ? true : false + super() + end + + def should_terminate(context) + result = ctx(context, :result).to_s + needle = @text + unless @case_sensitive + result = result.downcase + needle = needle.downcase + end + return Result.new(should_terminate: true, reason: "Text '#{@text}' found in output") if result.include?(needle) + + Result.new(should_terminate: false) + end + end + + # Stop when the whole output (stripped) equals +stop_message+ + class StopMessage < Condition + attr_reader :stop_message + + def initialize(stop_message = 'TERMINATE') + @stop_message = stop_message.to_s + super() + end + + def should_terminate(context) + return Result.new(should_terminate: true, reason: "Stop message '#{@stop_message}' received") if ctx(context, :result).to_s.strip == @stop_message + + Result.new(should_terminate: false) + end + end + + # Stop after +max_messages+ messages (falls back to the loop iteration count) + class MaxMessage < Condition + attr_reader :max_messages + + def initialize(max_messages) + raise ConfigurationError, 'max_messages must be >= 1' unless max_messages.is_a?(Integer) && max_messages >= 1 + + @max_messages = max_messages + super() + end + + def should_terminate(context) + messages = ctx(context, :messages) + count = messages.is_a?(Array) ? messages.size : 0 + count = ctx(context, :iteration).to_i if count.zero? + return Result.new(should_terminate: true, reason: "Message count (#{count}) >= limit (#{@max_messages})") if count >= @max_messages + + Result.new(should_terminate: false) + end + end + + # Stop when token usage crosses a budget + class TokenUsage < Condition + attr_reader :max_total_tokens, :max_prompt_tokens, :max_completion_tokens + + def initialize(max_total_tokens: nil, max_prompt_tokens: nil, max_completion_tokens: nil) + raise ConfigurationError, 'at least one token limit must be specified' if [max_total_tokens, max_prompt_tokens, max_completion_tokens].all?(&:nil?) + + @max_total_tokens = max_total_tokens + @max_prompt_tokens = max_prompt_tokens + @max_completion_tokens = max_completion_tokens + super() + end + + def should_terminate(context) + usage = ctx(context, :token_usage) + return Result.new(should_terminate: false) unless usage.respond_to?(:key?) + + checks = [ + [@max_total_tokens, usage_value(usage, 'total_tokens', 'totalTokens'), 'Total'], + [@max_prompt_tokens, usage_value(usage, 'prompt_tokens', 'promptTokens'), 'Prompt'], + [@max_completion_tokens, usage_value(usage, 'completion_tokens', 'completionTokens'), 'Completion'] + ] + checks.each do |limit, value, label| + next if limit.nil? || value < limit + + return Result.new(should_terminate: true, reason: "#{label} tokens (#{value}) >= limit (#{limit})") + end + Result.new(should_terminate: false) + end + + private + + def usage_value(usage, *keys) + keys.each do |k| + v = usage[k] || usage[k.to_sym] + return v.to_i unless v.nil? + end + 0 + end + end + + # All children must trigger + class And < Condition + attr_reader :conditions + + def initialize(*conditions) + @conditions = conditions.flat_map { |c| c.is_a?(And) ? c.conditions : [c] } + super() + end + + def should_terminate(context) + reasons = [] + @conditions.each do |c| + r = c.should_terminate(context) + return Result.new(should_terminate: false) unless r.should_terminate + + reasons << r.reason unless r.reason.to_s.empty? + end + Result.new(should_terminate: true, reason: reasons.join(' AND ')) + end + end + + # Any child triggers + class Or < Condition + attr_reader :conditions + + def initialize(*conditions) + @conditions = conditions.flat_map { |c| c.is_a?(Or) ? c.conditions : [c] } + super() + end + + def should_terminate(context) + @conditions.each do |c| + r = c.should_terminate(context) + return r if r.should_terminate + end + Result.new(should_terminate: false) + end + end + end + end +end diff --git a/lib/conductor/agents/tool_def.rb b/lib/conductor/agents/tool_def.rb new file mode 100644 index 0000000..183f29e --- /dev/null +++ b/lib/conductor/agents/tool_def.rb @@ -0,0 +1,272 @@ +# frozen_string_literal: true + +require 'json' +require_relative 'errors' + +module Conductor + module Agents + # Wire values for ToolConfig#toolType. Only +worker+ and +cli+ tools run in this + # process; every other type is executed by the Conductor server. + module ToolType + WORKER = 'worker' + HTTP = 'http' + API = 'api' + MCP = 'mcp' + HUMAN = 'human' + AGENT_TOOL = 'agent_tool' + GENERATE_IMAGE = 'generate_image' + GENERATE_AUDIO = 'generate_audio' + GENERATE_VIDEO = 'generate_video' + GENERATE_PDF = 'generate_pdf' + RAG_INDEX = 'rag_index' + RAG_SEARCH = 'rag_search' + PULL_WORKFLOW_MESSAGES = 'pull_workflow_messages' + CLI = 'cli' + + MEDIA = [GENERATE_IMAGE, GENERATE_AUDIO, GENERATE_VIDEO, GENERATE_PDF].freeze + RAG = [RAG_INDEX, RAG_SEARCH].freeze + LOCAL = [WORKER, CLI].freeze + ALL = [WORKER, HTTP, API, MCP, HUMAN, AGENT_TOOL, *MEDIA, *RAG, PULL_WORKFLOW_MESSAGES, CLI].freeze + + def self.valid?(type) + ALL.include?(type.to_s) + end + end + + # A tool call with pre-filled arguments (Agent#prefill_tools) + PrefillToolCall = Struct.new(:tool_name, :arguments, :tool_def, keyword_init: true) do + def to_h + { 'toolName' => tool_name, 'arguments' => arguments || {} } + end + end + + # Definition of one tool. Same fields and defaults as the Python SDK's ToolDef. + # + # Worker tools are created by the Tools DSL (+tool def ...+); server-side tools by + # the factories below (ToolDef.http, .mcp, .human, .agent, ...). + class ToolDef + RETRY_POLICIES = %w[fixed linear_backoff exponential_backoff].freeze + RETRY_LOGIC = { + 'fixed' => 'FIXED', + 'linear_backoff' => 'LINEAR_BACKOFF', + 'exponential_backoff' => 'EXPONENTIAL_BACKOFF' + }.freeze + CREDENTIAL_PLACEHOLDER = /\$\{(\w+)\}/ + + attr_accessor :name, :description, :input_schema, :output_schema, :func, + :approval_required, :timeout_seconds, :tool_type, :config, + :guardrails, :credentials, :stateful, :max_calls, + :retry_count, :retry_delay_seconds, :retry_policy + + # @param name [String] tool name; for worker tools this is also the Conductor task name + # @param func [Proc, Method, nil] local implementation; nil for server-side tools + def initialize(name:, description: '', input_schema: nil, output_schema: nil, func: nil, + approval_required: false, timeout_seconds: nil, tool_type: ToolType::WORKER, + config: nil, guardrails: nil, credentials: nil, stateful: false, max_calls: nil, + retry_count: 2, retry_delay_seconds: 2, retry_policy: 'linear_backoff') + raise ConfigurationError, 'tool name is required' if name.nil? || name.to_s.empty? + raise ConfigurationError, "unknown tool_type #{tool_type.inspect}" unless ToolType.valid?(tool_type) + unless RETRY_POLICIES.include?(retry_policy.to_s) || RETRY_LOGIC.value?(retry_policy.to_s) + raise ConfigurationError, "retry_policy must be one of #{RETRY_POLICIES.join(', ')}" + end + + @name = name.to_s + @description = description.to_s + @input_schema = input_schema || {} + @output_schema = output_schema || {} + @func = func + @approval_required = approval_required ? true : false + @timeout_seconds = timeout_seconds + @tool_type = tool_type.to_s + @config = config || {} + @guardrails = Array(guardrails) + @credentials = Array(credentials).map(&:to_s).uniq + @stateful = stateful ? true : false + @max_calls = max_calls + @retry_count = retry_count + @retry_delay_seconds = retry_delay_seconds + @retry_policy = retry_policy.to_s + end + + # True when this tool needs a worker polling in this process + def local? + !@func.nil? && ToolType::LOCAL.include?(@tool_type) + end + + def server_side? + !local? + end + + # Conductor retryLogic value for this tool's retry policy + def retry_logic + RETRY_LOGIC.fetch(@retry_policy) { @retry_policy.upcase } + end + + # Add secret names this tool needs (deduplicated) + # @return [self] + def add_credentials(*names) + @credentials = (@credentials + names.flatten.map(&:to_s)).uniq + self + end + + # Build a pre-filled call for Agent#prefill_tools + def call(**args) + PrefillToolCall.new(tool_name: @name, arguments: args.transform_keys(&:to_s), tool_def: self) + end + + def to_s + "#" + end + alias inspect to_s + + class << self + # Tool backed by an HTTP endpoint; the server makes the call. + # Headers may reference secrets as ${NAME}; every placeholder must be listed in +credentials+. + def http(name, url, description: '', method: 'GET', headers: nil, input_schema: nil, + accept: ['application/json'], content_type: 'application/json', credentials: nil) + creds = Array(credentials).map(&:to_s) + validate_placeholders!(headers, creds) + new( + name: name, description: description, + input_schema: input_schema || { 'type' => 'object', 'properties' => {} }, + tool_type: ToolType::HTTP, + config: { 'url' => url, 'method' => method.to_s.upcase, 'headers' => headers || {}, + 'accept' => accept, 'contentType' => content_type }, + credentials: creds + ) + end + + # Tools discovered from an OpenAPI endpoint; the server does the discovery. + def api(url, name: 'api_tools', description: nil, headers: nil, tool_names: nil, max_tools: 64, credentials: nil) + creds = Array(credentials).map(&:to_s) + validate_placeholders!(headers, creds) + config = { 'url' => url } + config['headers'] = headers if headers + config['tool_names'] = Array(tool_names) if tool_names + config['max_tools'] = max_tools + new(name: name, description: description || "API tools from #{url}", tool_type: ToolType::API, + config: config, credentials: creds) + end + + # Tools served by an MCP server; discovery (LIST_MCP_TOOLS) and calls happen on the server. + def mcp(server_url, name: 'mcp_tools', description: nil, headers: nil, tool_names: nil, + max_tools: 64, credentials: nil) + creds = Array(credentials).map(&:to_s) + validate_placeholders!(headers, creds) + config = { 'server_url' => server_url } + config['headers'] = headers if headers + config['tool_names'] = Array(tool_names) if tool_names + config['max_tools'] = max_tools + new(name: name, description: description || "MCP tools from #{server_url}", tool_type: ToolType::MCP, + config: config, credentials: creds) + end + + # Tool that pauses for a human answer (Conductor HUMAN task) + def human(name, description:, input_schema: nil) + new( + name: name, description: description, tool_type: ToolType::HUMAN, + input_schema: input_schema || { + 'type' => 'object', + 'properties' => { 'question' => { 'type' => 'string', + 'description' => 'The question or request for the human operator.' } }, + 'required' => ['question'] + } + ) + end + + # Another agent exposed as a tool (runs as a sub-workflow) + def agent(agent, name: nil, description: nil, retry_count: nil, retry_delay_seconds: nil, optional: nil) + agent_name = agent.respond_to?(:name) ? agent.name : agent.to_s + config = { 'agent' => agent } + config['retryCount'] = retry_count unless retry_count.nil? + config['retryDelaySeconds'] = retry_delay_seconds unless retry_delay_seconds.nil? + config['optional'] = optional unless optional.nil? + new( + name: name || agent_name, + description: description || "Invoke the #{agent_name} agent", + input_schema: { + 'type' => 'object', + 'properties' => { 'request' => { 'type' => 'string', + 'description' => 'The request or question to send to this agent.' } }, + 'required' => ['request'] + }, + tool_type: ToolType::AGENT_TOOL, + config: config + ) + end + + # Media generation tools (server-side) + def image(name, description:, llm_provider:, model:, input_schema: nil, **defaults) + media(ToolType::GENERATE_IMAGE, 'GENERATE_IMAGE', name, description, llm_provider, model, input_schema, defaults) + end + + def audio(name, description:, llm_provider:, model:, input_schema: nil, **defaults) + media(ToolType::GENERATE_AUDIO, 'GENERATE_AUDIO', name, description, llm_provider, model, input_schema, defaults) + end + + def video(name, description:, llm_provider:, model:, input_schema: nil, **defaults) + media(ToolType::GENERATE_VIDEO, 'GENERATE_VIDEO', name, description, llm_provider, model, input_schema, defaults) + end + + def pdf(name = 'generate_pdf', description: 'Generate a PDF document.', input_schema: nil, **defaults) + new(name: name, description: description, tool_type: ToolType::GENERATE_PDF, + input_schema: input_schema || { 'type' => 'object', 'properties' => {} }, + config: { 'taskType' => 'GENERATE_PDF' }.merge(stringify(defaults))) + end + + # RAG tools (server-side) + def index(name, description:, vector_db:, index:, embedding_model_provider:, embedding_model:, + namespace: 'default_ns', chunk_size: nil, chunk_overlap: nil, dimensions: nil, input_schema: nil) + config = { 'taskType' => 'LLM_INDEX_TEXT', 'vectorDB' => vector_db, 'namespace' => namespace, 'index' => index, + 'embeddingModelProvider' => embedding_model_provider, 'embeddingModel' => embedding_model } + config['chunkSize'] = chunk_size if chunk_size + config['chunkOverlap'] = chunk_overlap if chunk_overlap + config['dimensions'] = dimensions if dimensions + new(name: name, description: description, tool_type: ToolType::RAG_INDEX, + input_schema: input_schema || { 'type' => 'object', 'properties' => {} }, config: config) + end + + def search(name, description:, vector_db:, index:, embedding_model_provider:, embedding_model:, + namespace: 'default_ns', max_results: 5, dimensions: nil, input_schema: nil) + config = { 'taskType' => 'LLM_SEARCH_INDEX', 'vectorDB' => vector_db, 'namespace' => namespace, 'index' => index, + 'embeddingModelProvider' => embedding_model_provider, 'embeddingModel' => embedding_model, + 'maxResults' => max_results } + config['dimensions'] = dimensions if dimensions + new(name: name, description: description, tool_type: ToolType::RAG_SEARCH, + input_schema: input_schema || { 'type' => 'object', 'properties' => {} }, config: config) + end + + # Wait for messages posted to the execution (PULL_WORKFLOW_MESSAGES) + def wait_for_message(name, description:, batch_size: 1, blocking: true) + config = { 'batchSize' => batch_size } + config['blocking'] = false unless blocking + new(name: name, description: description, tool_type: ToolType::PULL_WORKFLOW_MESSAGES, + input_schema: { 'type' => 'object', 'properties' => {} }, config: config) + end + + private + + def media(tool_type, task_type, name, description, llm_provider, model, input_schema, defaults) + new(name: name, description: description, tool_type: tool_type, + input_schema: input_schema || { 'type' => 'object', 'properties' => {} }, + config: { 'taskType' => task_type, 'llmProvider' => llm_provider, 'model' => model }.merge(stringify(defaults))) + end + + def stringify(hash) + hash.transform_keys(&:to_s) + end + + def validate_placeholders!(headers, credentials) + return unless headers + + placeholders = headers.to_s.scan(CREDENTIAL_PLACEHOLDER).flatten.uniq + missing = placeholders - credentials + return if missing.empty? + + raise ConfigurationError, + "Header placeholder(s) #{missing.inspect} not declared in credentials: #{credentials.inspect}" + end + end + end + end +end diff --git a/lib/conductor/agents/tools.rb b/lib/conductor/agents/tools.rb new file mode 100644 index 0000000..f87ac97 --- /dev/null +++ b/lib/conductor/agents/tools.rb @@ -0,0 +1,171 @@ +# frozen_string_literal: true + +require_relative 'errors' +require_relative 'runtime/secrets' +require_relative 'tool_def' +require_relative 'tools/schema_builder' +require_relative 'tools/secret_scanner' +require_relative 'tools/ruby_llm_adapter' + +module Conductor + module Agents + # The tool DSL. + # + # include Conductor::Agents # top level, or + # module Weather; extend Conductor::Agents::Tools; ... end + # + # tool def get_weather(city: String, units: 'metric') + # { temp_c: 21.0 } + # end + # describe :get_weather, 'Get the current weather for a city.' + # requires_approval :get_weather + # + # +tool+ receives the Symbol that +def+ returns, builds a ToolDef from the method + # (schema from keyword defaults, secrets from literal secret() calls) and registers it + # both on the receiver (Weather[:get_weather], Weather.tool_defs) and in the global + # registry that Agent#add_tool(:get_weather) consults. + module Tools + TOOL_OPTIONS = %i[description output_schema approval_required timeout_seconds credentials + stateful max_calls retry_count retry_delay_seconds retry_policy].freeze + + # Weather[:current] on a module that `extend Conductor::Agents::Tools` + module Lookup + # @return [ToolDef] + def [](name) + fetch_tool(name) + end + end + + class << self + # A module that extends Tools also gets [] and the secret helpers + def extended(base) + base.extend(Lookup) + base.extend(Secrets) + end + + # Global name => ToolDef registry shared by every scope that defines tools + def registry + @registry ||= {} + end + + def registry_mutex + @registry_mutex ||= Mutex.new + end + + def register(tool_def) + registry_mutex.synchronize { registry[tool_def.name] = tool_def } + tool_def + end + + # @return [ToolDef, nil] + def lookup(name) + registry_mutex.synchronize { registry[name.to_s] } + end + + # Forget every registered tool (tests) + def clear! + registry_mutex.synchronize { registry.clear } + end + + # Build a ToolDef from a bound Method + # @param method [Method] + # @param name [String, nil] tool name override + def build(method, name: nil, **options) + unknown = options.keys - TOOL_OPTIONS + raise ConfigurationError, "unknown tool option(s): #{unknown.inspect}" unless unknown.empty? + + tool_name = (name || method.name).to_s + input_schema = SchemaBuilder.input_schema(method) + credentials = SecretScanner.scan(method) + + ToolDef.new( + name: tool_name, + description: options.fetch(:description) { humanize(method.name) }, + input_schema: input_schema, + output_schema: options.fetch(:output_schema) { SchemaBuilder.default_output_schema }, + func: method, + approval_required: options.fetch(:approval_required, false), + timeout_seconds: options[:timeout_seconds], + credentials: credentials + Array(options[:credentials]), + stateful: options.fetch(:stateful, false), + max_calls: options[:max_calls], + retry_count: options.fetch(:retry_count, 2), + retry_delay_seconds: options.fetch(:retry_delay_seconds, 2), + retry_policy: options.fetch(:retry_policy, 'linear_backoff') + ) + end + + # "get_weather" => "Get weather" + def humanize(name) + words = name.to_s.tr('_', ' ').strip + return '' if words.empty? + + words[0].upcase + words[1..] + end + end + + # Mark a method as a tool + # @param name [Symbol, String, Method] method name (what +def+ returns) or a Method + # @param options [Hash] ToolDef overrides: description:, output_schema:, approval_required:, + # timeout_seconds:, credentials:, stateful:, max_calls:, retry_count:, retry_delay_seconds:, retry_policy: + # @return [ToolDef] + def tool(name, **options) + method = name.is_a?(Method) ? name : resolve_tool_method(name.to_sym) + tool_def = Tools.build(method, name: name.is_a?(Method) ? name.name : name, **options) + tool_registry[tool_def.name] = tool_def + Tools.register(tool_def) + end + + # Override the description the LLM sees + def describe(name, text) + fetch_tool(name).description = text.to_s + end + + # Require a human approval before the tool runs + def requires_approval(name, enabled: true) + fetch_tool(name).approval_required = enabled + end + + # Declare secret names the scanner could not see (dynamic names) + def tool_credentials(name, *secret_names) + fetch_tool(name).add_credentials(*secret_names) + end + + # Every tool defined in this scope, in definition order + # @return [Array] + def tool_defs + tool_registry.values + end + + private + + def tool_registry + @conductor_tool_registry ||= {} # rubocop:disable Naming/MemoizedInstanceVariableName + end + + def fetch_tool(name) + tool_registry[name.to_s] || Tools.lookup(name) || + raise(ConfigurationError, "no tool named #{name.inspect}; define it with `tool def #{name}(...)` first") + end + + # Find the method behind +tool def name+ for the current receiver: + # - top level / objects: the method is on self + # - module with `extend Tools`: `def` made an instance method; module_function it + # - class bodies (e.g. inside RSpec.describe): bind the instance method to a bare instance + def resolve_tool_method(name) + return method(name) if respond_to?(name, true) + + if is_a?(Module) && (method_defined?(name) || private_method_defined?(name)) + if instance_of?(Module) + module_function(name) + return method(name) + end + + return instance_method(name).bind(allocate) + end + + raise ConfigurationError, "tool #{name.inspect}: no such method on #{inspect}" + end + end + end +end diff --git a/lib/conductor/agents/tools/ruby_llm_adapter.rb b/lib/conductor/agents/tools/ruby_llm_adapter.rb new file mode 100644 index 0000000..243948b --- /dev/null +++ b/lib/conductor/agents/tools/ruby_llm_adapter.rb @@ -0,0 +1,73 @@ +# frozen_string_literal: true + +module Conductor + module Agents + module Tools + # Adapts a RubyLLM::Tool class to a ToolDef so it can be given to an Agent as-is. + # + # RubyLLM exposes +name+, +description+ and +parameters+ (name => Parameter with + # +type+, +description+, +required+) and executes via +#execute(**args)+. RubyLLM is an + # optional dependency: this adapter only engages when RubyLLM::Tool is defined. + module RubyLlmAdapter + module_function + + # @return [Boolean] true when +klass+ is a RubyLLM::Tool subclass + def ruby_llm_tool?(klass) + return false unless defined?(::RubyLLM::Tool) + + klass.is_a?(Class) && klass < ::RubyLLM::Tool + end + + # @param klass [Class] a RubyLLM::Tool subclass + # @return [ToolDef] + def to_tool_def(klass) + instance = klass.new + name = read(klass, instance, :name) || snake_case(klass.name.to_s.split('::').last) + description = read(klass, instance, :description) || '' + params = read(klass, instance, :parameters) || {} + + properties = {} + required = [] + params.each do |param_name, param| + schema = { 'type' => (fetch(param, :type) || 'string').to_s } + desc = fetch(param, :description) + schema['description'] = desc if desc + properties[param_name.to_s] = schema + required << param_name.to_s if fetch(param, :required) != false + end + + input_schema = { 'type' => 'object', 'properties' => properties } + input_schema['required'] = required unless required.empty? + + ToolDef.new( + name: name.to_s, + description: description.to_s, + input_schema: input_schema, + output_schema: SchemaBuilder.default_output_schema, + func: ->(**args) { instance.execute(**args) } + ) + end + + def read(klass, instance, attr) + if klass.respond_to?(attr) + klass.public_send(attr) + elsif instance.respond_to?(attr) + instance.public_send(attr) + end + end + + def fetch(param, key) + if param.respond_to?(key) + param.public_send(key) + elsif param.respond_to?(:[]) + param[key] || param[key.to_s] + end + end + + def snake_case(name) + name.gsub(/([A-Z]+)([A-Z][a-z])/, '\1_\2').gsub(/([a-z\d])([A-Z])/, '\1_\2').downcase + end + end + end + end +end diff --git a/lib/conductor/agents/tools/schema_builder.rb b/lib/conductor/agents/tools/schema_builder.rb new file mode 100644 index 0000000..60c2f1c --- /dev/null +++ b/lib/conductor/agents/tools/schema_builder.rb @@ -0,0 +1,190 @@ +# frozen_string_literal: true + +module Conductor + module Agents + module Tools + # Builds a JSON schema for a tool method from its keyword arguments. + # + # Ruby exposes keyword names and whether they are required (+Method#parameters+) but + # never the default expressions, and the Tools DSL uses the default as the type: + # + # tool def get_weather(city: String, units: 'metric', limit: 10, tags: [String], mode: %w[a b]) + # + # so the defaults are read from the method's AST (RubyVM::AbstractSyntaxTree.of). That + # works on MRI whenever the method's source file is on disk (and for eval'd code on + # Ruby >= 3.2 with RubyVM.keep_script_lines = true). When the AST is unavailable the + # builder falls back to Python's behaviour: every keyword becomes an untyped property + # ({}), required when the keyword has no default. + module SchemaBuilder # rubocop:disable Metrics/ModuleLength + CLASS_TYPES = { + 'String' => 'string', 'Symbol' => 'string', + 'Integer' => 'integer', + 'Float' => 'number', 'Numeric' => 'number', 'BigDecimal' => 'number', + 'TrueClass' => 'boolean', 'FalseClass' => 'boolean', + 'Hash' => 'object', 'Array' => 'array', + 'Time' => 'string', 'Date' => 'string', 'DateTime' => 'string' + }.freeze + + POSITIONAL = %i[req opt rest].freeze + + module_function + + # @param method [Method, UnboundMethod] + # @return [Hash] JSON schema for the tool input + def input_schema(method) + params = method.parameters + positional = params.select { |kind, _| POSITIONAL.include?(kind) } + unless positional.empty? + raise ConfigurationError, + "tool #{method.name}: use keyword arguments only (found positional #{positional.map(&:last).inspect})" + end + + defaults = keyword_defaults(method) + properties = {} + required = [] + + params.each do |kind, name| + case kind + when :keyreq + properties[name.to_s] = defaults.key?(name) ? schema_for(defaults[name]) : {} + required << name.to_s + when :key + if defaults.key?(name) + schema, is_required = schema_for_default(defaults[name]) + properties[name.to_s] = schema + required << name.to_s if is_required + else + properties[name.to_s] = {} + end + end + end + + schema = { 'type' => 'object', 'properties' => properties } + schema['required'] = required unless required.empty? + schema + end + + # Default output schema for worker tools: Dispatch always returns a JSON object + def default_output_schema + { 'type' => 'object', 'additionalProperties' => {} } + end + + # Map keyword name => default AST node, or {} when no AST is available + def keyword_defaults(method) + ast = ast_of(method) + return {} unless ast + + args = find_node(ast, :ARGS) + return {} unless args + + kw = args.children[7] + result = {} + each_node(kw) do |node| + next unless node.type == :KW_ARG + + lasgn = node.children[0] + next unless lasgn.respond_to?(:type) && lasgn.type == :LASGN + + name, default = lasgn.children + # required keywords (city:) carry a Symbol placeholder instead of a default node + result[name] = default if default.is_a?(RubyVM::AbstractSyntaxTree::Node) + end + result + end + + def ast_of(method) + return nil unless defined?(RubyVM::AbstractSyntaxTree) + + RubyVM::AbstractSyntaxTree.of(method) + rescue StandardError + nil + end + + # Schema for a default that stands for a *type* (class constant or array of one) + def schema_for(node) + schema_for_default(node).first + end + + # @return [Array(Hash, Boolean)] schema and whether the parameter is required + def schema_for_default(node) + case node.type + when :CONST + type = CLASS_TYPES[node.children[0].to_s] + [type ? { 'type' => type } : {}, true] + when :COLON2 + type = CLASS_TYPES[node.children[1].to_s] + [type ? { 'type' => type } : {}, true] + when :STR + [{ 'type' => 'string', 'default' => node.children[0] }, false] + when :DSTR, :XSTR, :DXSTR + [{ 'type' => 'string' }, false] + when :LIT, :INTEGER, :FLOAT, :RATIONAL, :IMAGINARY + literal_schema(node.children[0]) + when :TRUE + [{ 'type' => 'boolean', 'default' => true }, false] + when :FALSE + [{ 'type' => 'boolean', 'default' => false }, false] + when :NIL + [{}, false] + when :LIST, :ZLIST + list_schema(node) + when :HASH + [{ 'type' => 'object', 'default' => {} }, false] + else + [{}, false] + end + end + + def literal_schema(value) + case value + when Integer then [{ 'type' => 'integer', 'default' => value }, false] + when Float then [{ 'type' => 'number', 'default' => value }, false] + when Symbol then [{ 'type' => 'string', 'default' => value.to_s }, false] + when Regexp then [{ 'type' => 'string', 'pattern' => value.source }, false] + else [{}, false] + end + end + + def list_schema(node) + elements = node.type == :ZLIST ? [] : node.children.compact + return [{ 'type' => 'array', 'default' => [] }, false] if elements.empty? + + if elements.size == 1 && %i[CONST COLON2].include?(elements[0].type) + item, = schema_for_default(elements[0]) + return [{ 'type' => 'array', 'items' => item }, true] + end + + if elements.all? { |e| e.type == :STR } + values = elements.map { |e| e.children[0] } + return [{ 'type' => 'string', 'enum' => values, 'default' => values.first }, false] + end + + if elements.all? { |e| %i[LIT INTEGER].include?(e.type) && e.children[0].is_a?(Integer) } + values = elements.map { |e| e.children[0] } + return [{ 'type' => 'integer', 'enum' => values, 'default' => values.first }, false] + end + + [{ 'type' => 'array' }, false] + end + + def find_node(node, type) + found = nil + each_node(node) do |n| + if n.type == type + found = n + break + end + end + found + end + + def each_node(node, &block) + return unless node.is_a?(RubyVM::AbstractSyntaxTree::Node) + + block.call(node) + node.children.each { |child| each_node(child, &block) } + end + end + end + end +end diff --git a/lib/conductor/agents/tools/secret_scanner.rb b/lib/conductor/agents/tools/secret_scanner.rb new file mode 100644 index 0000000..279f320 --- /dev/null +++ b/lib/conductor/agents/tools/secret_scanner.rb @@ -0,0 +1,45 @@ +# frozen_string_literal: true + +module Conductor + module Agents + module Tools + # Finds the secret names a tool body reads, so they can be declared on the wire + # (TaskDef#runtime_metadata / tool.config.credentials) without a separate list. + # + # Only string literals are picked up: + # + # secret('GH_TOKEN') # => GH_TOKEN + # secrets_env('A', 'B') # => A, B + # secret(name) # dynamic: declare with add_tool ..., credentials: [...] + module SecretScanner + SECRET_METHODS = %i[secret secrets_env].freeze + + module_function + + # @param method [Method, UnboundMethod] + # @return [Array] literal secret names, in order of appearance + def scan(method) + ast = SchemaBuilder.ast_of(method) + return [] unless ast + + names = [] + SchemaBuilder.each_node(ast) do |node| + method_id, args = call_parts(node) + next unless SECRET_METHODS.include?(method_id) + + SchemaBuilder.each_node(args) { |a| names << a.children[0] if a.type == :STR } + end + names.uniq + end + + def call_parts(node) + case node.type + when :FCALL then [node.children[0], node.children[1]] + when :CALL, :QCALL then [node.children[1], node.children[2]] + else [nil, nil] + end + end + end + end + end +end diff --git a/spec/conductor/agents/agent_spec.rb b/spec/conductor/agents/agent_spec.rb new file mode 100644 index 0000000..07e5ff1 --- /dev/null +++ b/spec/conductor/agents/agent_spec.rb @@ -0,0 +1,164 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'support/agent_tools' + +RSpec.describe Conductor::Agents::Agent do + let(:model) { 'openai/gpt-4o' } + + describe '#initialize' do + it 'validates the name pattern and max_turns' do + expect { described_class.new(name: '1bad', model: model) }.to raise_error(Conductor::Agents::ConfigurationError, /name/) + expect { described_class.new(name: 'has space', model: model) }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: 'ok', model: model, max_turns: 0) }.to raise_error(Conductor::Agents::ConfigurationError, /max_turns/) + expect(described_class.new(name: 'ok-name_1', model: model).name).to eq('ok-name_1') + end + + it 'normalizes the strategy and requires a router for :router' do + expect(described_class.new(name: 'a', model: model).strategy).to eq('handoff') + expect(described_class.new(name: 'a', model: model).strategy_set?).to be false + expect(described_class.new(name: 'a', model: model, strategy: :round_robin).strategy).to eq('round_robin') + expect { described_class.new(name: 'a', model: model, strategy: :zigzag) }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: 'a', model: model, strategy: :router) }.to raise_error(Conductor::Agents::ConfigurationError, /router/) + end + + it 'accepts tools and agents lists' do + child = described_class.new(name: 'c', model: model) + agent = described_class.new(name: 'a', model: model, tools: [:current], agents: [child]) + expect(agent.tools.map(&:name)).to eq(['current']) + expect(agent.agents).to eq([child]) + end + end + + describe '#add_tool' do + let(:agent) { described_class.new(name: 'a', model: model) } + + it 'accepts a symbol from the global registry, a ToolDef, a Tools module and an Agent' do + agent.add_tool :current + agent.add_tool Conductor::Agents::ToolDef.http('fetch', 'http://x') + agent.add_tool SpecTools::Github + agent.add_tool described_class.new(name: 'helper', model: model) + expect(agent.tools.map(&:name)).to eq(%w[current fetch create_issue gh_cli dynamic_secret helper]) + expect(agent.tool('helper').tool_type).to eq('agent_tool') + end + + it 'adds explicit credentials on a copy so the registry tool stays untouched' do + agent.add_tool :create_issue, credentials: ['EXTRA'] + expect(agent.tool('create_issue').credentials).to eq(%w[GH_TOKEN EXTRA]) + expect(SpecTools::Github[:create_issue].credentials).to eq(['GH_TOKEN']) + end + + it 'rejects unknown names, duplicates and junk' do + expect { agent.add_tool :missing }.to raise_error(Conductor::Agents::ConfigurationError, /no tool named/) + agent.add_tool :current + expect { agent.add_tool :current }.to raise_error(Conductor::Agents::ConfigurationError, /duplicate/) + expect { agent.add_tool 42 }.to raise_error(Conductor::Agents::ConfigurationError) + expect { agent.add_tool Comparable }.to raise_error(Conductor::Agents::ConfigurationError, /no tools/) + end + end + + describe 'team sugar' do + it 'add_agent rejects non-agents and duplicate names' do + team = described_class.new(name: 'team') + a = described_class.new(name: 'a', model: model) + team.add_agent(a) + expect { team.add_agent(described_class.new(name: 'a', model: model)) }.to raise_error(Conductor::Agents::ConfigurationError, /duplicate/) + expect { team.add_agent('a') }.to raise_error(Conductor::Agents::ConfigurationError) + team.add_agents(described_class.new(name: 'b', model: model), described_class.new(name: 'c', model: model)) + expect(team.agents.map(&:name)).to eq(%w[a b c]) + end + + it 'hands_off_to builds OnTextMention or OnCondition' do + triage = described_class.new(name: 'triage', model: model) + filer = described_class.new(name: 'filer', model: model) + triage.hands_off_to filer, on: 'ACTIONABLE' + triage.hands_off_to filer, on: ->(ctx) { ctx['iteration'] > 3 } + expect(triage.handoffs[0]).to be_a(Conductor::Agents::Handoff::OnTextMention) + expect(triage.handoffs[0].target).to eq('filer') + expect(triage.handoffs[0].text).to eq('ACTIONABLE') + expect(triage.handoffs[1]).to be_a(Conductor::Agents::Handoff::OnCondition) + end + + it '>> builds a flattened sequential pipeline' do + a = described_class.new(name: 'a', model: model) + b = described_class.new(name: 'b', model: model) + c = described_class.new(name: 'c', model: model) + pipeline = a >> b >> c + expect(pipeline.name).to eq('a_b_c') + expect(pipeline.strategy).to eq('sequential') + expect(pipeline.agents).to eq([a, b, c]) + expect(pipeline.model).to eq(model) + end + end + + describe 'guardrail and termination sugar' do + let(:agent) { described_class.new(name: 'a', model: model) } + + it 'redact adds a fixing regex guardrail' do + agent.redact %w[password api.key] + g = agent.guardrails.first + expect(g).to be_a(Conductor::Agents::RegexGuardrail) + expect(g.on_fail).to eq('fix') + expect(g.position).to eq('output') + expect(g.pattern_strings).to eq(['password', 'api\.key']) + expect(g.name).to eq('a_redact') + end + + it 'stop_when and stop_after combine with OR' do + agent.stop_when 'ISSUE_FILED' + expect(agent.termination).to be_a(Conductor::Agents::Termination::TextMention) + agent.stop_after messages: 12 + expect(agent.termination).to be_a(Conductor::Agents::Termination::Or) + expect(agent.termination.conditions.map(&:class)).to eq([Conductor::Agents::Termination::TextMention, + Conductor::Agents::Termination::MaxMessage]) + end + end + + describe 'callbacks' do + it 'collects positions from handlers and blocks' do + handler = Class.new(Conductor::Agents::CallbackHandler) do + def on_tool_end(**_kwargs) + nil + end + end.new + agent = described_class.new(name: 'a', model: model, callbacks: [handler]) + agent.callback(:before_model) { |**_| nil } + expect(agent.callback_positions).to eq(%w[before_model after_tool]) + expect { agent.callback(:sideways) { nil } }.to raise_error(Conductor::Agents::ConfigurationError) + end + end + + describe '#on_approval' do + it 'stores the handler' do + agent = described_class.new(name: 'a', model: model) + handler = proc { |req| req.approve } + agent.on_approval(&handler) + expect(agent.approval_handler).to eq(handler) + end + end + + describe 'introspection' do + it 'walks the whole tree and detects stateful members' do + leaf = described_class.new(name: 'leaf', model: model, stateful: true) + router = described_class.new(name: 'router', model: model) + tool_agent = described_class.new(name: 'tool_agent', model: model) + root = described_class.new(name: 'root', model: model, agents: [leaf], strategy: :router, router: router) + root.add_tool tool_agent + expect(root.all_agents.map(&:name)).to eq(%w[root leaf router tool_agent]) + expect(root.stateful_tree?).to be true + expect(described_class.new(name: 'x', model: model).stateful_tree?).to be false + end + end + + describe '#call_sync / #call_async' do + it 'delegates to the default runtime' do + runtime = double('runtime') + allow(Conductor::Agents).to receive(:runtime).and_return(runtime) + agent = described_class.new(name: 'a', model: model) + expect(runtime).to receive(:call_sync).with(agent, 'hi', session_id: 's').and_return('answer') + expect(runtime).to receive(:call_async).with(agent, 'hi', session_id: nil) + expect(agent.call_sync('hi', session_id: 's')).to eq('answer') + agent.call_async('hi') + end + end +end diff --git a/spec/conductor/agents/callback_handler_spec.rb b/spec/conductor/agents/callback_handler_spec.rb new file mode 100644 index 0000000..ec688e9 --- /dev/null +++ b/spec/conductor/agents/callback_handler_spec.rb @@ -0,0 +1,47 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::CallbackHandler do + let(:timing) do + Class.new(described_class) do + def on_model_start(**_kwargs) + { 'seen' => true } + end + end + end + + let(:noisy) do + Class.new(described_class) do + def on_model_start(**_kwargs) + raise 'boom' + end + + def on_model_end(**_kwargs) + nil + end + end + end + + it 'knows which positions a handler overrides' do + handler = timing.new + expect(handler.handles?('before_model')).to be true + expect(handler.handles?(:after_model)).to be false + end + + it 'chains handlers with first-non-empty-hash-wins and skips errors' do + chain = described_class.chain('before_model', [noisy.new, timing.new], logger: Logger.new(nil)) + expect(chain.call(messages: [])).to eq('seen' => true) + end + + it 'returns nil when nothing is registered and {} when handlers return nil' do + expect(described_class.chain('after_agent', [timing.new])).to be_nil + expect(described_class.chain('after_model', [noisy.new]).call(llm_result: 'x')).to eq({}) + end + + it 'runs procs before handlers' do + chain = described_class.chain('before_model', [timing.new], [->(**_) { { 'proc' => 1 } }]) + expect(chain.call).to eq('proc' => 1) + end +end diff --git a/spec/conductor/agents/config_serializer_spec.rb b/spec/conductor/agents/config_serializer_spec.rb new file mode 100644 index 0000000..2992e6c --- /dev/null +++ b/spec/conductor/agents/config_serializer_spec.rb @@ -0,0 +1,165 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'support/agent_tools' + +RSpec.describe Conductor::Agents::ConfigSerializer do + a = Conductor::Agents + let(:model) { 'openai/gpt-4o' } + + def serialize(agent) + described_class.serialize(agent) + end + + it 'emits the always-present keys and drops nils and empties' do + config = serialize(a::Agent.new(name: 'greeter', model: model)) + expect(config).to eq('name' => 'greeter', 'model' => model, 'maxTurns' => 25, 'timeoutSeconds' => 0, + 'external' => false) + end + + it 'emits strategy only for agents with sub-agents' do + leaf = a::Agent.new(name: 'leaf', model: model, strategy: :parallel) + expect(serialize(leaf)).not_to have_key('strategy') + team = a::Agent.new(name: 'team', model: model, agents: [leaf], strategy: :parallel) + expect(serialize(team)['strategy']).to eq('parallel') + expect(serialize(team)['agents'].first['name']).to eq('leaf') + end + + it 'inherits a missing team model from the first member and rejects model-less leaves' do + team = a::Agent.new(name: 'bug_desk') + team.add_agent a::Agent.new(name: 'triage', model: 'openai/gpt-4o-mini') + team.add_agent a::Agent.new(name: 'filer', model: 'anthropic/claude-sonnet-4-5') + expect(serialize(team)['model']).to eq('openai/gpt-4o-mini') + + expect { serialize(a::Agent.new(name: 'lonely')) }.to raise_error(a::ConfigurationError, /no model/) + expect(serialize(a::Agent.new(name: 'ext', external: true))).not_to have_key('model') + end + + it 'puts agent credentials at the top level and tool credentials under config' do + agent = a::Agent.new(name: 'filer', model: model, credentials: ['GH_TOKEN']) + agent.add_tool :create_issue, credentials: ['EXTRA'] + config = serialize(agent) + expect(config['credentials']).to eq(['GH_TOKEN']) + tool = config['tools'].first + expect(tool['config']).to eq('credentials' => %w[GH_TOKEN EXTRA]) + expect(tool).not_to have_key('credentials') + expect(tool['approvalRequired']).to be true + end + + it 'serializes worker tools with schema, description and default output schema' do + agent = a::Agent.new(name: 'w', model: model, tools: [:current]) + tool = serialize(agent)['tools'].first + expect(tool).to eq( + 'name' => 'current', 'description' => 'Current', 'toolType' => 'worker', + 'inputSchema' => { 'type' => 'object', + 'properties' => { 'city' => { 'type' => 'string' }, + 'units' => { 'type' => 'string', 'default' => 'metric' } }, + 'required' => ['city'] }, + 'outputSchema' => { 'type' => 'object', 'additionalProperties' => {} } + ) + end + + it 'marks tools stateful when the agent is stateful' do + agent = a::Agent.new(name: 'w', model: model, tools: [:current], stateful: true) + expect(serialize(agent)['tools'].first['stateful']).to be true + end + + it 'replaces the agent in an agent_tool config with agentConfig' do + child = a::Agent.new(name: 'child', model: model) + parent = a::Agent.new(name: 'parent', model: model, tools: [a::ToolDef.agent(child, optional: false)]) + tool = serialize(parent)['tools'].first + expect(tool['toolType']).to eq('agent_tool') + expect(tool['config']['agentConfig']['name']).to eq('child') + expect(tool['config']['optional']).to be false + expect(tool['config']).not_to have_key('agent') + end + + it 'serializes instructions as string, prompt template or callable' do + expect(serialize(a::Agent.new(name: 'x', model: model, instructions: ''))).not_to have_key('instructions') + tpl = a::PromptTemplate.new(name: 'support_prompt', variables: { 'tone' => 'kind' }, version: 2) + expect(serialize(a::Agent.new(name: 'x', model: model, instructions: tpl))['instructions']).to eq( + 'type' => 'prompt_template', 'name' => 'support_prompt', 'variables' => { 'tone' => 'kind' }, 'version' => 2 + ) + expect(serialize(a::Agent.new(name: 'x', model: model, instructions: -> { 'dynamic' }))['instructions']).to eq('dynamic') + end + + it 'serializes guardrails of every kind' do + agent = a::Agent.new( + name: 'g', model: model, + guardrails: [ + a::RegexGuardrail.new('x', name: 'rx', message: 'no x', on_fail: :retry), + a::LlmGuardrail.new(model, 'policy', name: 'llm', max_tokens: 10), + a::Guardrail.new(name: 'custom', on_fail: :fix) { |_| true }, + a::Guardrail.new(name: 'remote', position: :input) + ] + ) + expect(serialize(agent)['guardrails']).to eq([ + { 'name' => 'rx', 'position' => 'output', 'onFail' => 'retry', 'maxRetries' => 3, + 'guardrailType' => 'regex', 'patterns' => ['x'], 'mode' => 'block', 'message' => 'no x' }, + { 'name' => 'llm', 'position' => 'output', 'onFail' => 'raise', 'maxRetries' => 3, + 'guardrailType' => 'llm', 'model' => model, 'policy' => 'policy', 'maxTokens' => 10 }, + { 'name' => 'custom', 'position' => 'output', 'onFail' => 'fix', 'maxRetries' => 3, + 'guardrailType' => 'custom', 'taskName' => 'custom' }, + { 'name' => 'remote', 'position' => 'input', 'onFail' => 'raise', 'maxRetries' => 3, + 'guardrailType' => 'external', 'taskName' => 'remote' } + ]) + end + + it 'serializes handoffs including on_condition task names' do + team = a::Agent.new(name: 'team', model: model, strategy: :swarm, + agents: [a::Agent.new(name: 'b', model: model)], + handoffs: [a::Handoff::OnToolResult.new(target: 'b', tool_name: 't', result_contains: 'x'), + a::Handoff::OnCondition.new(target: 'b') { true }]) + expect(serialize(team)['handoffs']).to eq([ + { 'target' => 'b', 'type' => 'on_tool_result', 'toolName' => 't', 'resultContains' => 'x' }, + { 'target' => 'b', 'type' => 'on_condition', 'taskName' => 'team_handoff_b' } + ]) + end + + it 'hoists member hands_off_to into a swarm team when no strategy was chosen' do + triage = a::Agent.new(name: 'triage', model: model) + filer = a::Agent.new(name: 'filer', model: model) + triage.hands_off_to filer, on: 'ACTIONABLE' + team = a::Agent.new(name: 'bug_desk') + team.add_agents triage, filer + config = serialize(team) + expect(config['strategy']).to eq('swarm') + expect(config['handoffs']).to eq([{ 'target' => 'filer', 'type' => 'on_text_mention', 'text' => 'ACTIONABLE' }]) + + explicit = a::Agent.new(name: 'seq', strategy: :sequential) + explicit.add_agents triage, filer + expect(serialize(explicit)['strategy']).to eq('sequential') + expect(serialize(explicit)).not_to have_key('handoffs') + end + + it 'serializes memory, callbacks, prefill tools, metadata and numeric options' do + memory = a::ConversationMemory.new(max_messages: 5) + memory.add_user_message('hi') + agent = a::Agent.new(name: 'm', model: model, memory: memory, max_tokens: 100, temperature: 0.2, + metadata: { 'team' => 'x' }, prefill_tools: [a::ToolDef.new(name: 't').call(a: 1)]) + agent.callback(:after_model) { |**_| nil } + config = serialize(agent) + expect(config['memory']).to eq('messages' => [{ 'role' => 'user', 'message' => 'hi' }], 'maxMessages' => 5) + expect(config['callbacks']).to eq([{ 'position' => 'after_model', 'taskName' => 'm_after_model' }]) + expect(config['prefillTools']).to eq([{ 'toolName' => 't', 'arguments' => { 'a' => 1 } }]) + expect(config['metadata']).to eq('team' => 'x') + expect(config['maxTokens']).to eq(100) + expect(config['temperature']).to eq(0.2) + end + + it 'serializes a router agent or router task reference' do + child = a::Agent.new(name: 'c', model: model) + by_agent = a::Agent.new(name: 'r', model: model, agents: [child], strategy: :router, router: child) + expect(serialize(by_agent)['router']['name']).to eq('c') + by_proc = a::Agent.new(name: 'r2', model: model, agents: [child], strategy: :router, router: ->(_ctx) { 'c' }) + expect(serialize(by_proc)['router']).to eq('taskName' => 'r2_router_fn') + end + + it 'serializes output_type from a schema hash' do + schema = { 'title' => 'Report', 'type' => 'object', 'properties' => {} } + expect(serialize(a::Agent.new(name: 'o', model: model, output_type: schema))['outputType']).to eq( + 'schema' => schema, 'className' => 'Report' + ) + expect(serialize(a::Agent.new(name: 'o', model: model, output_type: { class_name: 'X' }))['outputType']).to eq('className' => 'X') + end +end diff --git a/spec/conductor/agents/contract_spec.rb b/spec/conductor/agents/contract_spec.rb new file mode 100644 index 0000000..56926f6 --- /dev/null +++ b/spec/conductor/agents/contract_spec.rb @@ -0,0 +1,40 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'json' +require 'json_schemer' +require 'conductor/agents' +require_relative '../../../examples/agents/golden_agents' + +# Contract tests: the Ruby serializer must produce exactly what the Python SDK produces +# (golden configs vendored from python-sdk/examples/agents/_configs) and every config must +# validate against the published agent schema. +RSpec.describe 'agentConfig contract' do + fixtures = File.expand_path('../../fixtures/agents', __dir__) + schema = JSONSchemer.schema(JSON.parse(File.read(File.join(fixtures, 'agent-schema.json')))) + golden_files = Dir[File.join(fixtures, 'configs', '*.json')] + + it 'has a golden fixture for every example and vice versa' do + expect(golden_files.map { |f| File.basename(f, '.json') }).to match_array(GoldenAgents::EXAMPLES.keys) + end + + golden_files.each do |file| + name = File.basename(file, '.json') + + it "serializes #{name} exactly like the Python SDK" do + expected = JSON.parse(File.read(file)) + actual = JSON.parse(JSON.generate(Conductor::Agents::ConfigSerializer.serialize(GoldenAgents::EXAMPLES.fetch(name).call))) + expect(actual).to eq(expected) + end + + it "#{name} validates against agent-schema.json" do + config = JSON.parse(JSON.generate(Conductor::Agents::ConfigSerializer.serialize(GoldenAgents::EXAMPLES.fetch(name).call))) + errors = schema.validate(config).map { |e| e['error'] } + expect(errors).to eq([]) + end + end + + it 'rejects unknown root keys (the schema is closed)' do + expect(schema.valid?({ 'name' => 'x', 'bogus' => 1 })).to be false + end +end diff --git a/spec/conductor/agents/guardrail_spec.rb b/spec/conductor/agents/guardrail_spec.rb new file mode 100644 index 0000000..05bae94 --- /dev/null +++ b/spec/conductor/agents/guardrail_spec.rb @@ -0,0 +1,59 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::Guardrail do + it 'validates position, on_fail, human-on-input and max_retries' do + expect { described_class.new(name: 'g', position: :middle) }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: 'g', on_fail: :explode) }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: 'g', position: :input, on_fail: :human) }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: 'g', max_retries: 0) }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new }.to raise_error(Conductor::Agents::ConfigurationError) + end + + it 'is external without a block and custom with one' do + external = described_class.new(name: 'remote') + expect(external.external?).to be true + expect(external.guardrail_type).to eq('external') + expect { external.check('x') }.to raise_error(Conductor::Agents::Error) + + custom = described_class.new(name: 'no_pii', on_fail: :retry) { |c| !c.include?('ssn') } + expect(custom.guardrail_type).to eq('custom') + expect(custom.check('fine').passed?).to be true + expect(custom.check('ssn 1').passed?).to be false + expect(custom.on_fail).to eq('retry') + expect(custom.position).to eq('output') + end +end + +RSpec.describe Conductor::Agents::RegexGuardrail do + it 'blocks matches in block mode with the custom or default message' do + g = described_class.new(['\d{3}-\d{2}-\d{4}'], name: 'no_ssn', message: 'No SSNs') + expect(g.check('my ssn is 123-45-6789').message).to eq('No SSNs') + expect(g.check('nothing here').passed?).to be true + expect(described_class.new('x').check('x').message).to eq('Content matched a blocked pattern.') + end + + it 'requires a match in allow mode and accepts Regexp patterns' do + g = described_class.new(/^\s*[{\[]/, mode: :allow) + expect(g.check('{"a":1}').passed?).to be true + expect(g.check('nope').message).to eq('Content did not match any allowed pattern.') + expect(g.pattern_strings).to eq(['^\s*[{\[]']) + expect(g.guardrail_type).to eq('regex') + end + + it 'rejects invalid modes' do + expect { described_class.new('x', mode: :maybe) }.to raise_error(Conductor::Agents::ConfigurationError) + end +end + +RSpec.describe Conductor::Agents::LlmGuardrail do + it 'stores model, policy and max_tokens and evaluates on the server' do + g = described_class.new('openai/gpt-4o-mini', 'No medical advice', name: 'safety', max_tokens: 100) + expect(g.guardrail_type).to eq('llm') + expect(g.model).to eq('openai/gpt-4o-mini') + expect(g.check('anything').passed?).to be false + expect { described_class.new('gpt-4o', 'p') }.to raise_error(Conductor::Agents::ConfigurationError) + end +end diff --git a/spec/conductor/agents/handoff_spec.rb b/spec/conductor/agents/handoff_spec.rb new file mode 100644 index 0000000..1b896d5 --- /dev/null +++ b/spec/conductor/agents/handoff_spec.rb @@ -0,0 +1,37 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::Handoff do + h = described_class + + it 'accepts an Agent or a name as target' do + agent = Conductor::Agents::Agent.new(name: 'filer', model: 'openai/gpt-4o') + expect(h::OnTextMention.new(target: agent, text: 'x').target).to eq('filer') + expect(h::OnTextMention.new(target: 'filer', text: 'x').target).to eq('filer') + expect { h::OnTextMention.new(target: '', text: 'x') }.to raise_error(Conductor::Agents::ConfigurationError) + end + + it 'OnTextMention matches case-insensitively' do + cond = h::OnTextMention.new(target: 'filer', text: 'ACTIONABLE') + expect(cond.should_handoff('result' => 'This is actionable')).to be true + expect(cond.should_handoff(result: 'nope')).to be false + end + + it 'OnToolResult matches the tool and optional substring' do + cond = h::OnToolResult.new(target: 'refund', tool_name: 'check_order', result_contains: 'broken') + expect(cond.should_handoff('tool_name' => 'check_order', 'tool_result' => 'item broken')).to be true + expect(cond.should_handoff('tool_name' => 'check_order', 'tool_result' => 'fine')).to be false + expect(cond.should_handoff('tool_name' => 'other')).to be false + expect(h::OnToolResult.new(target: 'r', tool_name: 'x').should_handoff('tool_name' => 'x')).to be true + end + + it 'OnCondition calls the block and swallows errors' do + cond = h::OnCondition.new(target: 'summarizer') { |ctx| ctx['iteration'] > 5 } + expect(cond.should_handoff('iteration' => 6)).to be true + expect(cond.should_handoff('iteration' => 1)).to be false + expect(cond.should_handoff({})).to be false + expect { h::OnCondition.new(target: 's') }.to raise_error(Conductor::Agents::ConfigurationError) + end +end diff --git a/spec/conductor/agents/memory_spec.rb b/spec/conductor/agents/memory_spec.rb new file mode 100644 index 0000000..6694ef7 --- /dev/null +++ b/spec/conductor/agents/memory_spec.rb @@ -0,0 +1,43 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::ConversationMemory do + it 'records messages in the Conductor chat format' do + m = described_class.new + m.add_system_message('sys') + m.add_user_message('hi') + m.add_assistant_message('hello') + m.add_tool_call('get_weather', { 'city' => 'Lisbon' }) + m.add_tool_result('get_weather', { temp: 21 }) + expect(m.messages).to eq([ + { 'role' => 'system', 'message' => 'sys' }, + { 'role' => 'user', 'message' => 'hi' }, + { 'role' => 'assistant', 'message' => 'hello' }, + { 'role' => 'tool_call', 'message' => '', + 'tool_calls' => [{ 'name' => 'get_weather', 'taskReferenceName' => 'get_weather_ref', + 'input' => { 'city' => 'Lisbon' } }] }, + { 'role' => 'tool', 'message' => '{:temp=>21}', 'toolCallId' => 'get_weather_ref', + 'taskReferenceName' => 'get_weather_ref' } + ]) + end + + it 'trims the oldest non-system messages first' do + m = described_class.new(max_messages: 3) + m.add_system_message('sys') + m.add_user_message('one') + m.add_assistant_message('two') + m.add_user_message('three') + expect(m.messages.map { |x| x['message'] }).to eq(%w[sys two three]) + end + + it 'deep copies in to_chat_messages and clears' do + m = described_class.new(messages: [{ role: 'user', message: 'x' }]) + copy = m.to_chat_messages + copy[0]['message'] = 'changed' + expect(m.messages[0]['message']).to eq('x') + m.clear + expect(m).to be_empty + end +end diff --git a/spec/conductor/agents/termination_spec.rb b/spec/conductor/agents/termination_spec.rb new file mode 100644 index 0000000..b7d8b90 --- /dev/null +++ b/spec/conductor/agents/termination_spec.rb @@ -0,0 +1,59 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::Termination do + t = described_class + + it 'TextMention matches case-insensitively by default' do + cond = t::TextMention.new('DONE') + expect(cond.should_terminate('result' => 'all done').should_terminate).to be true + expect(t::TextMention.new('DONE', case_sensitive: true).should_terminate(result: 'all done').should_terminate).to be false + expect(cond.should_terminate('result' => 'working').should_terminate).to be false + expect { t::TextMention.new('') }.to raise_error(Conductor::Agents::ConfigurationError) + end + + it 'StopMessage needs an exact stripped match' do + cond = t::StopMessage.new + expect(cond.should_terminate('result' => " TERMINATE \n").should_terminate).to be true + expect(cond.should_terminate('result' => 'TERMINATE now').should_terminate).to be false + end + + it 'MaxMessage counts messages or falls back to iteration' do + cond = t::MaxMessage.new(2) + expect(cond.should_terminate('messages' => [1, 2]).should_terminate).to be true + expect(cond.should_terminate('messages' => [], 'iteration' => 1).should_terminate).to be false + expect(cond.should_terminate('iteration' => 5).reason).to include('5') + expect { t::MaxMessage.new(0) }.to raise_error(Conductor::Agents::ConfigurationError) + end + + it 'TokenUsage checks each configured limit' do + cond = t::TokenUsage.new(max_total_tokens: 100) + expect(cond.should_terminate('token_usage' => { 'total_tokens' => 100 }).should_terminate).to be true + expect(cond.should_terminate('token_usage' => { 'totalTokens' => 10 }).should_terminate).to be false + expect(cond.should_terminate({}).should_terminate).to be false + expect { t::TokenUsage.new }.to raise_error(Conductor::Agents::ConfigurationError) + end + + it 'combines with & and | and flattens same-type children' do + a = t::TextMention.new('A') + b = t::MaxMessage.new(3) + c = t::StopMessage.new('X') + both = a & b & c + expect(both).to be_a(t::And) + expect(both.conditions.size).to eq(3) + expect(both.should_terminate('result' => 'A X', 'iteration' => 3).should_terminate).to be false + expect(both.should_terminate('result' => 'A', 'iteration' => 3).should_terminate).to be false + + either = a | b | c + expect(either).to be_a(t::Or) + expect(either.conditions.size).to eq(3) + expect(either.should_terminate('result' => 'X').reason).to include('X') + expect(either.should_terminate('result' => 'nothing').should_terminate).to be false + + nested = c | (a & b) + expect(nested.conditions.size).to eq(2) + expect(nested.conditions.last).to be_a(t::And) + end +end diff --git a/spec/conductor/agents/tool_def_spec.rb b/spec/conductor/agents/tool_def_spec.rb new file mode 100644 index 0000000..d314017 --- /dev/null +++ b/spec/conductor/agents/tool_def_spec.rb @@ -0,0 +1,104 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::ToolDef do + describe '#initialize' do + it 'applies the Python defaults' do + td = described_class.new(name: 't') + expect(td.description).to eq('') + expect(td.input_schema).to eq({}) + expect(td.tool_type).to eq('worker') + expect(td.approval_required).to be false + expect(td.retry_count).to eq(2) + expect(td.retry_delay_seconds).to eq(2) + expect(td.retry_policy).to eq('linear_backoff') + expect(td.retry_logic).to eq('LINEAR_BACKOFF') + expect(td.credentials).to eq([]) + expect(td.server_side?).to be true + end + + it 'rejects unknown tool types and retry policies' do + expect { described_class.new(name: 't', tool_type: 'magic') }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: 't', retry_policy: 'never') }.to raise_error(Conductor::Agents::ConfigurationError) + expect { described_class.new(name: '') }.to raise_error(Conductor::Agents::ConfigurationError) + end + + it 'is local only for worker/cli tools with a func' do + expect(described_class.new(name: 't', func: -> {}).local?).to be true + expect(described_class.new(name: 't', func: -> {}, tool_type: 'http').local?).to be false + expect(described_class.new(name: 't').local?).to be false + end + end + + describe '#call' do + it 'builds a prefilled tool call' do + call = described_class.new(name: 'get_weather').call(city: 'Lisbon') + expect(call.to_h).to eq('toolName' => 'get_weather', 'arguments' => { 'city' => 'Lisbon' }) + end + end + + describe '.http' do + it 'serializes the Python config keys' do + td = described_class.http('fetch', 'https://x.test/a', description: 'Fetch', method: 'post', + headers: { 'Authorization' => 'Bearer ${API_KEY}' }, credentials: ['API_KEY']) + expect(td.tool_type).to eq('http') + expect(td.config).to eq('url' => 'https://x.test/a', 'method' => 'POST', + 'headers' => { 'Authorization' => 'Bearer ${API_KEY}' }, + 'accept' => ['application/json'], 'contentType' => 'application/json') + expect(td.credentials).to eq(['API_KEY']) + expect(td.input_schema).to eq('type' => 'object', 'properties' => {}) + end + + it 'rejects undeclared ${NAME} placeholders' do + expect do + described_class.http('fetch', 'https://x', headers: { 'X' => '${SECRET}' }) + end.to raise_error(Conductor::Agents::ConfigurationError, /SECRET/) + end + end + + describe '.mcp' do + it 'defaults the name and carries server_url / max_tools' do + td = described_class.mcp('http://mcp:3001/mcp', tool_names: %w[a b]) + expect(td.name).to eq('mcp_tools') + expect(td.tool_type).to eq('mcp') + expect(td.config).to eq('server_url' => 'http://mcp:3001/mcp', 'tool_names' => %w[a b], 'max_tools' => 64) + end + end + + describe '.human' do + it 'uses a question schema by default' do + td = described_class.human('ask_user', description: 'Ask the user.') + expect(td.tool_type).to eq('human') + expect(td.input_schema['required']).to eq(['question']) + end + end + + describe '.agent' do + it 'wraps an agent with the request schema and retry config' do + agent = Conductor::Agents::Agent.new(name: 'researcher', model: 'openai/gpt-4o') + td = described_class.agent(agent, retry_count: 0, optional: false) + expect(td.name).to eq('researcher') + expect(td.description).to eq('Invoke the researcher agent') + expect(td.tool_type).to eq('agent_tool') + expect(td.config).to eq('agent' => agent, 'retryCount' => 0, 'optional' => false) + expect(td.input_schema['required']).to eq(['request']) + end + end + + describe 'media, rag and message factories' do + it 'set the task type in config' do + expect(described_class.image('img', description: 'd', llm_provider: 'openai', model: 'dall-e-3').config['taskType']).to eq('GENERATE_IMAGE') + expect(described_class.pdf.config).to eq('taskType' => 'GENERATE_PDF') + idx = described_class.index('idx', description: 'd', vector_db: 'pinecone', index: 'docs', + embedding_model_provider: 'openai', embedding_model: 'e', chunk_size: 100) + expect(idx.config).to include('taskType' => 'LLM_INDEX_TEXT', 'chunkSize' => 100, 'namespace' => 'default_ns') + search = described_class.search('s', description: 'd', vector_db: 'pinecone', index: 'docs', + embedding_model_provider: 'openai', embedding_model: 'e') + expect(search.config).to include('taskType' => 'LLM_SEARCH_INDEX', 'maxResults' => 5) + expect(described_class.wait_for_message('w', description: 'd', blocking: false).config).to eq('batchSize' => 1, 'blocking' => false) + expect(described_class.wait_for_message('w', description: 'd').config).to eq('batchSize' => 1) + end + end +end diff --git a/spec/conductor/agents/tools_spec.rb b/spec/conductor/agents/tools_spec.rb new file mode 100644 index 0000000..267860f --- /dev/null +++ b/spec/conductor/agents/tools_spec.rb @@ -0,0 +1,147 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'support/agent_tools' + +RSpec.describe Conductor::Agents::Tools do + let(:weather) { SpecTools::Weather } + + describe 'tool def' do + it 'names the tool after the method and humanizes the description' do + td = weather[:current] + expect(td).to be_a(Conductor::Agents::ToolDef) + expect(td.name).to eq('current') + expect(td.description).to eq('Current') + expect(td.local?).to be true + end + + it 'registers the tool globally so Agent#add_tool(:name) finds it' do + expect(described_class.lookup(:current)).to equal(weather[:current]) + end + + it 'types required parameters from class defaults and optional ones from literals' do + expect(weather[:current].input_schema).to eq( + 'type' => 'object', + 'properties' => { + 'city' => { 'type' => 'string' }, + 'units' => { 'type' => 'string', 'default' => 'metric' } + }, + 'required' => ['city'] + ) + end + + it 'handles integers, booleans, floats, typed arrays, enums, nil and hash defaults' do + schema = weather[:forecast].input_schema + expect(schema['required']).to eq(%w[city tags]) + expect(schema['properties']).to eq( + 'city' => { 'type' => 'string' }, + 'days' => { 'type' => 'integer', 'default' => 3 }, + 'detailed' => { 'type' => 'boolean', 'default' => false }, + 'tags' => { 'type' => 'array', 'items' => { 'type' => 'string' } }, + 'mode' => { 'type' => 'string', 'enum' => %w[brief full], 'default' => 'brief' }, + 'ratio' => { 'type' => 'number', 'default' => 0.5 }, + 'extra' => {}, + 'opts' => { 'type' => 'object', 'default' => {} } + ) + end + + it 'defaults the output schema to an object' do + expect(weather[:current].output_schema).to eq('type' => 'object', 'additionalProperties' => {}) + end + + it 'keeps the method callable as a normal method' do + expect(weather.current(city: 'Lisbon')).to include(temp_c: 21.0) + expect(weather[:current].func.call(city: 'Porto', units: 'imperial')[:summary]).to eq('Sunny in Porto (imperial)') + end + + it 'rejects positional parameters' do + expect { SpecTools::Refunds.tool(:positional) }.to raise_error(Conductor::Agents::ConfigurationError, /keyword/) + end + + it 'raises for unknown methods and unknown options' do + expect { weather.tool(:nope) }.to raise_error(Conductor::Agents::ConfigurationError, /no such method/) + expect { weather.tool(:current, colour: 'red') }.to raise_error(Conductor::Agents::ConfigurationError, /unknown tool option/) + end + + it 'treats keywords without defaults as required untyped properties' do + expect(SpecTools::Plain[:lookup].input_schema).to eq( + 'type' => 'object', + 'properties' => { 'city' => {}, 'units' => { 'type' => 'string', 'default' => 'metric' } }, + 'required' => ['city'] + ) + end + + it 'falls back to untyped properties when the AST is unavailable' do + allow(Conductor::Agents::Tools::SchemaBuilder).to receive(:ast_of).and_return(nil) + schema = Conductor::Agents::Tools::SchemaBuilder.input_schema(SpecTools::Plain.method(:lookup)) + expect(schema).to eq('type' => 'object', 'properties' => { 'city' => {}, 'units' => {} }, 'required' => ['city']) + end + end + + describe 'describe / requires_approval / tool_credentials' do + it 'overrides the description' do + expect(weather[:forecast].description).to eq('Multi-day forecast.') + end + + it 'marks approval' do + expect(SpecTools::Refunds[:issue_refund].approval_required).to be true + expect(weather[:current].approval_required).to be false + end + + it 'adds explicit credentials' do + SpecTools::Github.tool_credentials(:dynamic_secret, 'DYN_KEY') + expect(SpecTools::Github[:dynamic_secret].credentials).to eq(['DYN_KEY']) + end + end + + describe 'secret scanning' do + it 'declares literal secret() names' do + expect(SpecTools::Github[:create_issue].credentials).to eq(['GH_TOKEN']) + end + + it 'declares every literal in secrets_env()' do + expect(SpecTools::Github[:gh_cli].credentials).to eq(%w[GH_TOKEN GH_HOST]) + end + + it 'ignores dynamic names' do + expect(SpecTools::Github[:dynamic_secret].credentials).not_to include('name') + end + end + + describe '#tool_defs' do + it 'lists the tools of a module in definition order' do + expect(weather.tool_defs.map(&:name)).to eq(%w[current forecast]) + end + end + + describe 'RubyLLM adapter' do + before do + stub_const('RubyLLM', SpecTools::FakeRubyLLM) + stub_const('WeatherLookup', Class.new(SpecTools::FakeRubyLLM::Tool) do + desc 'Gets current weather for a location' + param :latitude, type: :number, desc: 'Latitude' + param :longitude, type: :number, desc: 'Longitude' + param :units, type: :string, required: false + + def execute(latitude:, longitude:, units: 'metric') + { lat: latitude, lon: longitude, units: units } + end + end) + end + + it 'converts the class to a ToolDef and executes an instance' do + td = Conductor::Agents::Tools::RubyLlmAdapter.to_tool_def(WeatherLookup) + expect(td.name).to eq('weather_lookup') + expect(td.description).to eq('Gets current weather for a location') + expect(td.input_schema['required']).to eq(%w[latitude longitude]) + expect(td.input_schema['properties']['latitude']).to eq('type' => 'number', 'description' => 'Latitude') + expect(td.func.call(latitude: 1.0, longitude: 2.0)).to eq(lat: 1.0, lon: 2.0, units: 'metric') + end + + it 'is picked up by Agent#add_tool' do + agent = Conductor::Agents::Agent.new(name: 'a', model: 'openai/gpt-4o') + agent.add_tool WeatherLookup + expect(agent.tool('weather_lookup')).not_to be_nil + end + end +end diff --git a/spec/fixtures/agents/agent-schema.json b/spec/fixtures/agents/agent-schema.json new file mode 100644 index 0000000..12656b4 --- /dev/null +++ b/spec/fixtures/agents/agent-schema.json @@ -0,0 +1,87 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/conductor-oss/python-sdk/blob/main/docs/agents/reference/agent-schema.json", + "title": "Conductor Python AgentConfig", + "description": "The public wire shape emitted by AgentConfigSerializer under agentConfig.", + "$ref": "#/$defs/agentConfig", + "$defs": { + "jsonValue": { + "anyOf": [ + { "type": "string" }, { "type": "number" }, { "type": "integer" }, + { "type": "boolean" }, { "type": "null" }, + { "type": "array", "items": { "$ref": "#/$defs/jsonValue" } }, + { "type": "object", "additionalProperties": { "$ref": "#/$defs/jsonValue" } } + ] + }, + "taskReference": { + "type": "object", + "required": ["taskName"], + "properties": { "taskName": { "type": "string", "minLength": 1 } }, + "additionalProperties": false + }, + "tool": { + "type": "object", + "required": ["name"], + "properties": { + "name": { "type": "string", "minLength": 1 }, + "description": { "type": "string" }, + "toolType": { "type": "string" }, + "taskName": { "type": "string" }, + "inputSchema": { "type": "object" }, + "outputSchema": { "type": "object" }, + "credentials": { "type": "array", "items": { "type": "string" } }, + "approvalRequired": { "type": "boolean" } + }, + "additionalProperties": true + }, + "agentConfig": { + "type": "object", + "required": ["name"], + "properties": { + "name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_-]*$" }, + "model": { "type": ["string", "null"] }, + "baseUrl": { "type": ["string", "null"] }, + "strategy": { "type": ["string", "null"] }, + "maxTurns": { "type": ["integer", "null"], "minimum": 0 }, + "timeoutSeconds": { "type": ["integer", "null"], "minimum": 0 }, + "external": { "type": ["boolean", "null"] }, + "instructions": { "anyOf": [{ "type": "string" }, { "type": "object" }, { "type": "null" }] }, + "tools": { "type": "array", "items": { "$ref": "#/$defs/tool" } }, + "agents": { "type": "array", "items": { "$ref": "#/$defs/agentConfig" } }, + "router": { "anyOf": [{ "$ref": "#/$defs/taskReference" }, { "$ref": "#/$defs/agentConfig" }] }, + "outputType": { "type": "object" }, + "guardrails": { "type": "array", "items": { "type": "object" } }, + "memory": { "type": "object" }, + "maxTokens": { "type": "integer", "minimum": 0 }, + "contextWindowBudget": { "type": "integer", "minimum": 0 }, + "temperature": { "type": "number" }, + "reasoningEffort": { "type": "string" }, + "stopWhen": { "$ref": "#/$defs/taskReference" }, + "termination": { "type": "object" }, + "handoffs": { "type": "array", "items": { "type": "object" } }, + "allowedTransitions": { "type": "array", "items": { "type": "string" } }, + "introduction": { "type": "string" }, + "metadata": { "type": "object" }, + "enablePlanning": { "type": "boolean" }, + "planner": { "$ref": "#/$defs/agentConfig" }, + "fallback": { "$ref": "#/$defs/agentConfig" }, + "callbacks": { "type": "array", "items": { "type": "object" } }, + "includeContents": { "type": "boolean" }, + "thinkingConfig": { "type": "object" }, + "requiredTools": { "type": "array", "items": { "type": "string" } }, + "prefillTools": { "type": "array", "items": { "type": "object" } }, + "fallbackMaxTurns": { "type": "integer", "minimum": 0 }, + "planSource": { "type": "string" }, + "plannerContext": { "type": "array", "items": { "type": "object" } }, + "synthesize": { "type": "boolean" }, + "maskedFields": { "type": "array", "items": { "type": "string" } }, + "gate": { "type": "object" }, + "codeExecution": { "type": "object" }, + "cliConfig": { "type": "object" }, + "credentials": { "type": "array", "items": { "type": "string" } }, + "_framework": { "type": "string" } + }, + "additionalProperties": false + } + } +} diff --git a/spec/fixtures/agents/configs/01_basic_agent.json b/spec/fixtures/agents/configs/01_basic_agent.json new file mode 100644 index 0000000..62be834 --- /dev/null +++ b/spec/fixtures/agents/configs/01_basic_agent.json @@ -0,0 +1,7 @@ +{ + "external": false, + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "greeter", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/02_tools.json b/spec/fixtures/agents/configs/02_tools.json new file mode 100644 index 0000000..8648227 --- /dev/null +++ b/spec/fixtures/agents/configs/02_tools.json @@ -0,0 +1,80 @@ +{ + "external": false, + "instructions": "You are a helpful assistant with access to weather, calculator, and email tools.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "tool_demo_agent", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Get current weather for a city.", + "inputSchema": { + "properties": { + "city": { + "type": "string" + } + }, + "required": [ + "city" + ], + "type": "object" + }, + "name": "get_weather", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Evaluate a math expression.", + "inputSchema": { + "properties": { + "expression": { + "type": "string" + } + }, + "required": [ + "expression" + ], + "type": "object" + }, + "name": "calculate", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "approvalRequired": true, + "description": "Send an email.", + "inputSchema": { + "properties": { + "body": { + "type": "string" + }, + "subject": { + "type": "string" + }, + "to": { + "type": "string" + } + }, + "required": [ + "to", + "subject", + "body" + ], + "type": "object" + }, + "name": "send_email", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "timeoutSeconds": 60, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/03_structured_output.json b/spec/fixtures/agents/configs/03_structured_output.json new file mode 100644 index 0000000..b79ea4a --- /dev/null +++ b/spec/fixtures/agents/configs/03_structured_output.json @@ -0,0 +1,61 @@ +{ + "external": false, + "instructions": "You are a weather reporter. Get the weather and provide a recommendation.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "weather_reporter", + "outputType": { + "className": "WeatherReport", + "schema": { + "properties": { + "city": { + "title": "City", + "type": "string" + }, + "condition": { + "title": "Condition", + "type": "string" + }, + "recommendation": { + "title": "Recommendation", + "type": "string" + }, + "temperature": { + "title": "Temperature", + "type": "number" + } + }, + "required": [ + "city", + "temperature", + "condition", + "recommendation" + ], + "title": "WeatherReport", + "type": "object" + } + }, + "timeoutSeconds": 0, + "tools": [ + { + "description": "Get current weather data for a city.", + "inputSchema": { + "properties": { + "city": { + "type": "string" + } + }, + "required": [ + "city" + ], + "type": "object" + }, + "name": "get_weather", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/05_handoffs.json b/spec/fixtures/agents/configs/05_handoffs.json new file mode 100644 index 0000000..1bd3afa --- /dev/null +++ b/spec/fixtures/agents/configs/05_handoffs.json @@ -0,0 +1,101 @@ +{ + "agents": [ + { + "external": false, + "instructions": "You handle billing questions: balances, payments, invoices.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "billing", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Check the balance of a bank account.", + "inputSchema": { + "properties": { + "account_id": { + "type": "string" + } + }, + "required": [ + "account_id" + ], + "type": "object" + }, + "name": "check_balance", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + }, + { + "external": false, + "instructions": "You handle technical questions: order status, shipping, returns.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "technical", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Look up the status of an order.", + "inputSchema": { + "properties": { + "order_id": { + "type": "string" + } + }, + "required": [ + "order_id" + ], + "type": "object" + }, + "name": "lookup_order", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + }, + { + "external": false, + "instructions": "You handle sales questions: pricing, products, promotions.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "sales", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Get pricing information for a product.", + "inputSchema": { + "properties": { + "product": { + "type": "string" + } + }, + "required": [ + "product" + ], + "type": "object" + }, + "name": "get_pricing", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } + ], + "external": false, + "instructions": "Route customer requests to the right specialist: billing, technical, or sales.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "support", + "strategy": "handoff", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/06_sequential_pipeline.json b/spec/fixtures/agents/configs/06_sequential_pipeline.json new file mode 100644 index 0000000..d224d42 --- /dev/null +++ b/spec/fixtures/agents/configs/06_sequential_pipeline.json @@ -0,0 +1,34 @@ +{ + "agents": [ + { + "external": false, + "instructions": "You are a researcher. Given a topic, provide key facts and data points. Be thorough but concise. Output raw research findings.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "researcher", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a writer. Take research findings and write a clear, engaging article. Use headers and bullet points where appropriate.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "writer", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are an editor. Review the article for clarity, grammar, and tone. Make improvements and output the final polished version.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "editor", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "researcher_writer_editor", + "strategy": "sequential", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/07_parallel_agents.json b/spec/fixtures/agents/configs/07_parallel_agents.json new file mode 100644 index 0000000..83f6025 --- /dev/null +++ b/spec/fixtures/agents/configs/07_parallel_agents.json @@ -0,0 +1,34 @@ +{ + "agents": [ + { + "external": false, + "instructions": "You are a market analyst. Analyze the given topic from a market perspective: market size, growth trends, key players, and opportunities.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "market_analyst", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a risk analyst. Analyze the given topic for risks: regulatory risks, technical risks, competitive threats, and mitigation strategies.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "risk_analyst", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a compliance specialist. Check the given topic for compliance considerations: data privacy, regulatory requirements, and industry standards.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "compliance", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "analysis", + "strategy": "parallel", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/08_router_agent.json b/spec/fixtures/agents/configs/08_router_agent.json new file mode 100644 index 0000000..fcac598 --- /dev/null +++ b/spec/fixtures/agents/configs/08_router_agent.json @@ -0,0 +1,43 @@ +{ + "agents": [ + { + "external": false, + "instructions": "You create implementation plans. Break down tasks into clear numbered steps.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "planner", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You write code. Output clean, well-documented Python code.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "coder", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You review code. Check for bugs, style issues, and suggest improvements.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "reviewer", + "timeoutSeconds": 0 + } + ], + "external": false, + "instructions": "You are the tech lead. Route requests to the right team member: planner for design/architecture, coder for implementation, reviewer for code review.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "dev_team", + "router": { + "external": false, + "instructions": "You create implementation plans. Break down tasks into clear numbered steps.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "planner", + "timeoutSeconds": 0 + }, + "strategy": "router", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/10_guardrails.json b/spec/fixtures/agents/configs/10_guardrails.json new file mode 100644 index 0000000..443dd21 --- /dev/null +++ b/spec/fixtures/agents/configs/10_guardrails.json @@ -0,0 +1,60 @@ +{ + "external": false, + "guardrails": [ + { + "guardrailType": "custom", + "maxRetries": 3, + "name": "no_pii", + "onFail": "retry", + "position": "output", + "taskName": "no_pii" + } + ], + "instructions": "You are a customer support assistant. Use the available tools to answer questions about orders and customers. Always include all details from the tool results in your response.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "support_agent", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Look up the current status of an order.", + "inputSchema": { + "properties": { + "order_id": { + "type": "string" + } + }, + "required": [ + "order_id" + ], + "type": "object" + }, + "name": "get_order_status", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Retrieve customer details including payment info on file.", + "inputSchema": { + "properties": { + "customer_id": { + "type": "string" + } + }, + "required": [ + "customer_id" + ], + "type": "object" + }, + "name": "get_customer_info", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/13_hierarchical_agents.json b/spec/fixtures/agents/configs/13_hierarchical_agents.json new file mode 100644 index 0000000..b06385a --- /dev/null +++ b/spec/fixtures/agents/configs/13_hierarchical_agents.json @@ -0,0 +1,77 @@ +{ + "agents": [ + { + "agents": [ + { + "external": false, + "instructions": "You are a backend developer. You design APIs, databases, and server architecture. Provide technical recommendations with code examples.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "backend_dev", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a frontend developer. You design UI components, user flows, and client-side architecture. Provide recommendations with code examples.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "frontend_dev", + "timeoutSeconds": 0 + } + ], + "external": false, + "instructions": "You are the engineering lead. Route technical questions to the right specialist: backend_dev for APIs/databases/servers, frontend_dev for UI/UX/client-side.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "engineering_lead", + "strategy": "handoff", + "timeoutSeconds": 0 + }, + { + "agents": [ + { + "external": false, + "instructions": "You are a content writer. You create blog posts, landing page copy, and marketing materials. Write engaging, clear content.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "content_writer", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are an SEO specialist. You optimize content for search engines, suggest keywords, and improve page rankings.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "seo_specialist", + "timeoutSeconds": 0 + } + ], + "external": false, + "instructions": "You are the marketing lead. Route marketing questions to the right specialist: content_writer for blog posts/copy, seo_specialist for SEO/keywords/rankings.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "marketing_lead", + "strategy": "handoff", + "timeoutSeconds": 0 + } + ], + "external": false, + "handoffs": [ + { + "target": "engineering_lead", + "text": "engineering_lead", + "type": "on_text_mention" + }, + { + "target": "marketing_lead", + "text": "marketing_lead", + "type": "on_text_mention" + } + ], + "instructions": "You are the CEO. Route requests to the right department: engineering_lead for technical/development questions, marketing_lead for marketing/content/SEO questions.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "ceo", + "strategy": "swarm", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/17_swarm_orchestration.json b/spec/fixtures/agents/configs/17_swarm_orchestration.json new file mode 100644 index 0000000..f4bc0af --- /dev/null +++ b/spec/fixtures/agents/configs/17_swarm_orchestration.json @@ -0,0 +1,39 @@ +{ + "agents": [ + { + "external": false, + "instructions": "You are a refund specialist. Process the customer's refund request. Check eligibility, confirm the refund amount, and let them know the timeline. Be empathetic and clear. Do NOT ask follow-up questions -- just process the refund based on what the customer told you.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "refund_specialist", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a technical support specialist. Diagnose the customer's technical issue and provide clear troubleshooting steps.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "tech_support", + "timeoutSeconds": 0 + } + ], + "external": false, + "handoffs": [ + { + "target": "refund_specialist", + "text": "refund", + "type": "on_text_mention" + }, + { + "target": "tech_support", + "text": "technical", + "type": "on_text_mention" + } + ], + "instructions": "You are the front-line customer support agent. Triage customer requests. If the customer needs a refund, transfer to the refund specialist. If they have a technical issue, transfer to tech support. Use the transfer tools available to you to hand off the conversation.", + "maxTurns": 3, + "model": "anthropic/claude-sonnet-4-6", + "name": "support", + "strategy": "swarm", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/19_composable_termination_and.json b/spec/fixtures/agents/configs/19_composable_termination_and.json new file mode 100644 index 0000000..932ab3a --- /dev/null +++ b/spec/fixtures/agents/configs/19_composable_termination_and.json @@ -0,0 +1,43 @@ +{ + "external": false, + "instructions": "Research thoroughly. Only provide your FINAL ANSWER after using the search tool at least twice.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "deliberator", + "termination": { + "conditions": [ + { + "caseSensitive": false, + "text": "FINAL ANSWER", + "type": "text_mention" + }, + { + "maxMessages": 5, + "type": "max_message" + } + ], + "type": "and" + }, + "timeoutSeconds": 0, + "tools": [ + { + "description": "Search for information.", + "inputSchema": { + "properties": { + "query": { + "type": "string" + } + }, + "required": [ + "query" + ], + "type": "object" + }, + "name": "search", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/19_composable_termination_complex.json b/spec/fixtures/agents/configs/19_composable_termination_complex.json new file mode 100644 index 0000000..99cf2e6 --- /dev/null +++ b/spec/fixtures/agents/configs/19_composable_termination_complex.json @@ -0,0 +1,56 @@ +{ + "external": false, + "instructions": "Research and provide a comprehensive answer.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "complex_agent", + "termination": { + "conditions": [ + { + "stopMessage": "TERMINATE", + "type": "stop_message" + }, + { + "conditions": [ + { + "caseSensitive": false, + "text": "DONE", + "type": "text_mention" + }, + { + "maxMessages": 10, + "type": "max_message" + } + ], + "type": "and" + }, + { + "maxTotalTokens": 50000, + "type": "token_usage" + } + ], + "type": "or" + }, + "timeoutSeconds": 0, + "tools": [ + { + "description": "Search for information.", + "inputSchema": { + "properties": { + "query": { + "type": "string" + } + }, + "required": [ + "query" + ], + "type": "object" + }, + "name": "search", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/19_composable_termination_or.json b/spec/fixtures/agents/configs/19_composable_termination_or.json new file mode 100644 index 0000000..6ff6ca6 --- /dev/null +++ b/spec/fixtures/agents/configs/19_composable_termination_or.json @@ -0,0 +1,22 @@ +{ + "external": false, + "instructions": "Have a conversation. Say GOODBYE when you're finished.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "chatbot", + "termination": { + "conditions": [ + { + "caseSensitive": false, + "text": "GOODBYE", + "type": "text_mention" + }, + { + "maxMessages": 20, + "type": "max_message" + } + ], + "type": "or" + }, + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/19_composable_termination_simple.json b/spec/fixtures/agents/configs/19_composable_termination_simple.json new file mode 100644 index 0000000..596e674 --- /dev/null +++ b/spec/fixtures/agents/configs/19_composable_termination_simple.json @@ -0,0 +1,34 @@ +{ + "external": false, + "instructions": "Research the topic and say DONE when you have enough info.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "researcher", + "termination": { + "caseSensitive": false, + "text": "DONE", + "type": "text_mention" + }, + "timeoutSeconds": 0, + "tools": [ + { + "description": "Search for information.", + "inputSchema": { + "properties": { + "query": { + "type": "string" + } + }, + "required": [ + "query" + ], + "type": "object" + }, + "name": "search", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/21_regex_guardrails.json b/spec/fixtures/agents/configs/21_regex_guardrails.json new file mode 100644 index 0000000..3b4a0f7 --- /dev/null +++ b/spec/fixtures/agents/configs/21_regex_guardrails.json @@ -0,0 +1,56 @@ +{ + "external": false, + "guardrails": [ + { + "guardrailType": "regex", + "maxRetries": 3, + "message": "Response must not contain email addresses. Redact them.", + "mode": "block", + "name": "no_email_addresses", + "onFail": "retry", + "patterns": [ + "[\\w.+-]+@[\\w-]+\\.[\\w.-]+" + ], + "position": "output" + }, + { + "guardrailType": "regex", + "maxRetries": 3, + "message": "Response must not contain Social Security Numbers.", + "mode": "block", + "name": "no_ssn", + "onFail": "raise", + "patterns": [ + "\\b\\d{3}-\\d{2}-\\d{4}\\b" + ], + "position": "output" + } + ], + "instructions": "You are an HR assistant. When asked about employees, look up their profile and share ALL the details you find.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "hr_assistant", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Retrieve a user's profile from the database.", + "inputSchema": { + "properties": { + "user_id": { + "type": "string" + } + }, + "required": [ + "user_id" + ], + "type": "object" + }, + "name": "get_user_profile", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/22_llm_guardrails.json b/spec/fixtures/agents/configs/22_llm_guardrails.json new file mode 100644 index 0000000..7f0b028 --- /dev/null +++ b/spec/fixtures/agents/configs/22_llm_guardrails.json @@ -0,0 +1,20 @@ +{ + "external": false, + "guardrails": [ + { + "guardrailType": "llm", + "maxRetries": 3, + "maxTokens": 10000, + "model": "anthropic/claude-sonnet-4-6", + "name": "content_safety", + "onFail": "retry", + "policy": "Reject any content that:\n1. Contains medical or legal advice presented as fact\n2. Makes promises or guarantees about outcomes\n3. Includes discriminatory or biased language\n4. Reveals private information about real individuals\n\nEven if there are disclaimers you should reject", + "position": "output" + } + ], + "instructions": "You are a health information assistant. Provide general health information and solution to the problem. You can prescribe psudo scientific and untested meds ", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "health_advisor", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/45_agent_tool.json b/spec/fixtures/agents/configs/45_agent_tool.json new file mode 100644 index 0000000..6f49e04 --- /dev/null +++ b/spec/fixtures/agents/configs/45_agent_tool.json @@ -0,0 +1,79 @@ +{ + "external": false, + "instructions": "You are a project manager. Use the researcher tool to gather information and the calculate tool for math. Synthesize findings.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "manager_45", + "timeoutSeconds": 0, + "tools": [ + { + "config": { + "agentConfig": { + "external": false, + "instructions": "You are a research assistant. Use search_knowledge_base to find information about topics. Provide concise summaries.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "researcher_45", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Search an internal knowledge base for information.", + "inputSchema": { + "properties": { + "query": { + "type": "string" + } + }, + "required": [ + "query" + ], + "type": "object" + }, + "name": "search_knowledge_base", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } + }, + "description": "Invoke the researcher_45 agent", + "inputSchema": { + "properties": { + "request": { + "description": "The request or question to send to this agent.", + "type": "string" + } + }, + "required": [ + "request" + ], + "type": "object" + }, + "name": "researcher_45", + "toolType": "agent_tool" + }, + { + "description": "Evaluate a math expression safely.", + "inputSchema": { + "properties": { + "expression": { + "type": "string" + } + }, + "required": [ + "expression" + ], + "type": "object" + }, + "name": "calculate", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/47_callbacks.json b/spec/fixtures/agents/configs/47_callbacks.json new file mode 100644 index 0000000..6f9af80 --- /dev/null +++ b/spec/fixtures/agents/configs/47_callbacks.json @@ -0,0 +1,40 @@ +{ + "callbacks": [ + { + "position": "before_model", + "taskName": "monitored_agent_47_before_model" + }, + { + "position": "after_model", + "taskName": "monitored_agent_47_after_model" + } + ], + "external": false, + "instructions": "You are a helpful assistant. Use get_facts when asked about topics.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "monitored_agent_47", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Get interesting facts about a topic.", + "inputSchema": { + "properties": { + "topic": { + "type": "string" + } + }, + "required": [ + "topic" + ], + "type": "object" + }, + "name": "get_facts", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] +} \ No newline at end of file diff --git a/spec/fixtures/agents/configs/52_nested_strategies.json b/spec/fixtures/agents/configs/52_nested_strategies.json new file mode 100644 index 0000000..eee7b9f --- /dev/null +++ b/spec/fixtures/agents/configs/52_nested_strategies.json @@ -0,0 +1,44 @@ +{ + "agents": [ + { + "agents": [ + { + "external": false, + "instructions": "You are a market analyst. Analyze the market size, growth rate, and key players for the given topic. Be concise (3-4 bullet points).", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "market_analyst_52", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a risk analyst. Identify the top 3 risks: regulatory, technical, and competitive. Be concise.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "risk_analyst_52", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "research_phase_52", + "strategy": "parallel", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are an executive briefing writer. Synthesize the market analysis and risk assessment into a concise executive summary (1 paragraph).", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "summarizer_52", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "research_phase_52_summarizer_52", + "strategy": "sequential", + "timeoutSeconds": 0 +} \ No newline at end of file diff --git a/spec/support/agent_tools.rb b/spec/support/agent_tools.rb new file mode 100644 index 0000000..9341ca0 --- /dev/null +++ b/spec/support/agent_tools.rb @@ -0,0 +1,95 @@ +# frozen_string_literal: true + +# Tool definitions used by the agents specs. They live in a real file so the +# Tools DSL can read keyword defaults and secret() literals from the AST. +require 'conductor/agents' + +module SpecTools + module Weather + extend Conductor::Agents::Tools + + tool def current(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city} (#{units})" } + end + + tool def forecast(city: String, days: 3, detailed: false, tags: [String], mode: %w[brief full], ratio: 0.5, + extra: nil, opts: {}) + { city: city, days: days, detailed: detailed, tags: tags, mode: mode, ratio: ratio, extra: extra, opts: opts } + end + describe :forecast, 'Multi-day forecast.' + end + + module Github + extend Conductor::Agents::Tools + + tool def create_issue(title: String, body: '') + { title: title, body: body, token: secret('GH_TOKEN') } + end + + tool def gh_cli(title: String) + secrets_env('GH_TOKEN', 'GH_HOST') + end + + tool def dynamic_secret(name: String) + secret(name) + end + requires_approval :create_issue + end + + module Plain + extend Conductor::Agents::Tools + + tool def lookup(city:, units: 'metric') + [city, units] + end + end + + module Refunds + extend Conductor::Agents::Tools + + tool def issue_refund(order_id: String, amount: Float) + { refunded: amount, order_id: order_id } + end + requires_approval :issue_refund + + def self.positional(city, units: 'metric') + [city, units] + end + end + + # A stand-in for RubyLLM::Tool so the adapter can be exercised without the gem + module FakeRubyLLM + class Tool + class Param + attr_reader :type, :description, :required + + def initialize(type:, description:, required:) + @type = type + @description = description + @required = required + end + end + + class << self + def desc(text = nil) + @description = text if text + @description + end + + attr_reader :description + + def param(name, type: :string, desc: nil, required: true) + (@parameters ||= {})[name] = Param.new(type: type, description: desc, required: required) + end + + def parameters + @parameters || {} + end + + def name + super.split('::').last.gsub(/([a-z\d])([A-Z])/, '\1_\2').downcase + end + end + end + end +end From aefd104ca1eeb4e4744187841b16bed7469350b8 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Tue, 8 Sep 2026 15:46:50 -0700 Subject: [PATCH 04/20] Phase 2: agent runtime, SSE streaming, tool dispatch, secrets, system workers - AgentRuntime: call_sync / call_async (POST /agent/start, register the workers the server lists in requiredWorkers, stream on a background thread), deploy, compile, serve, shutdown; Conductor::Agents.runtime default instance and .configure - SseClient over Net::HTTP: incremental parser (comments, ids, multi-line data), reconnect with Last-Event-ID, heartbeat-only detection, SseUnavailableError -> StatusPoller fallback over GET /agent/{id}/status - Execution (done?, waiting?, result with timeout, partial_text, finish_reason, tool_calls, token_usage summed over sub-executions, pause/resume/cancel/stop/signal, Execution.find) and ApprovalRequest (batch tool calls, argument accessors, approve / reject / send_message posting the server's respond bodies); on_approval wiring with execution.pending when no handler is registered - Dispatch: strips server-injected keys (method, _agent_state, _agent_tool_name, _allowed_commands), validates required args, coerces by JSON schema, wraps scalar results, terminal failures for missing args/credentials/unserializable output - Secrets.secret / secrets_env read Task#runtime_metadata via TaskContext, fall back to ENV, never write ENV - ToolRegistry builds Worker.define-style workers with Python's TaskDef defaults (retry 2, LINEAR_BACKOFF, timeout 0, response timeout 10, RETRY, runtimeMetadata) and SystemWorkers for _termination, custom guardrails, callbacks and on_condition handoffs; unknown required workers are logged - AgentConfig.from_env for CONDUCTOR_AGENT_* settings Co-Authored-By: Claude Fable 5.1 --- lib/conductor/agents.rb | 31 ++ lib/conductor/agents/runtime/agent_config.rb | 52 ++++ lib/conductor/agents/runtime/agent_runtime.rb | 259 +++++++++++++++++ .../agents/runtime/approval_request.rb | 105 +++++++ lib/conductor/agents/runtime/dispatch.rb | 137 +++++++++ lib/conductor/agents/runtime/execution.rb | 270 ++++++++++++++++++ lib/conductor/agents/runtime/sse_client.rb | 175 ++++++++++++ lib/conductor/agents/runtime/status_poller.rb | 49 ++++ .../agents/runtime/system_workers.rb | 105 +++++++ lib/conductor/agents/runtime/tool_registry.rb | 145 ++++++++++ spec/conductor/agents/agent_config_spec.rb | 30 ++ spec/conductor/agents/agent_runtime_spec.rb | 216 ++++++++++++++ .../conductor/agents/approval_request_spec.rb | 46 +++ spec/conductor/agents/dispatch_spec.rb | 124 ++++++++ spec/conductor/agents/execution_spec.rb | 79 +++++ spec/conductor/agents/sse_client_spec.rb | 89 ++++++ spec/conductor/agents/status_poller_spec.rb | 32 +++ spec/conductor/agents/system_workers_spec.rb | 57 ++++ spec/conductor/agents/tool_registry_spec.rb | 76 +++++ 19 files changed, 2077 insertions(+) create mode 100644 lib/conductor/agents/runtime/agent_config.rb create mode 100644 lib/conductor/agents/runtime/agent_runtime.rb create mode 100644 lib/conductor/agents/runtime/approval_request.rb create mode 100644 lib/conductor/agents/runtime/dispatch.rb create mode 100644 lib/conductor/agents/runtime/execution.rb create mode 100644 lib/conductor/agents/runtime/sse_client.rb create mode 100644 lib/conductor/agents/runtime/status_poller.rb create mode 100644 lib/conductor/agents/runtime/system_workers.rb create mode 100644 lib/conductor/agents/runtime/tool_registry.rb create mode 100644 spec/conductor/agents/agent_config_spec.rb create mode 100644 spec/conductor/agents/agent_runtime_spec.rb create mode 100644 spec/conductor/agents/approval_request_spec.rb create mode 100644 spec/conductor/agents/dispatch_spec.rb create mode 100644 spec/conductor/agents/execution_spec.rb create mode 100644 spec/conductor/agents/sse_client_spec.rb create mode 100644 spec/conductor/agents/status_poller_spec.rb create mode 100644 spec/conductor/agents/system_workers_spec.rb create mode 100644 spec/conductor/agents/tool_registry_spec.rb diff --git a/lib/conductor/agents.rb b/lib/conductor/agents.rb index b93e4be..4562b6d 100644 --- a/lib/conductor/agents.rb +++ b/lib/conductor/agents.rb @@ -25,6 +25,15 @@ require_relative 'agents/prompt_template' require_relative 'agents/agent' require_relative 'agents/config_serializer' +require_relative 'agents/runtime/agent_config' +require_relative 'agents/runtime/dispatch' +require_relative 'agents/runtime/system_workers' +require_relative 'agents/runtime/tool_registry' +require_relative 'agents/runtime/execution' +require_relative 'agents/runtime/approval_request' +require_relative 'agents/runtime/sse_client' +require_relative 'agents/runtime/status_poller' +require_relative 'agents/runtime/agent_runtime' module Conductor # Ruby port of the Python SDK's conductor.ai.agents package @@ -36,6 +45,28 @@ class << self # Tools and secrets are usable at the module level too (Conductor::Agents.tool ...) include Tools include Secrets + + # The default runtime used by Agent#call_sync / #call_async (built from the environment) + # @return [AgentRuntime] + def runtime + @runtime ||= AgentRuntime.new + end + + attr_writer :runtime + + # Replace the default runtime + # Conductor::Agents.configure(configuration: Conductor::Configuration.new(server_api_url: '...')) + # @return [AgentRuntime] + def configure(configuration: nil, agent_config: nil, logger: nil) + @runtime&.shutdown + @runtime = AgentRuntime.new(configuration: configuration, agent_config: agent_config, logger: logger) + end + + # Stop the default runtime's workers and streams + def shutdown + @runtime&.shutdown + @runtime = nil + end end end end diff --git a/lib/conductor/agents/runtime/agent_config.rb b/lib/conductor/agents/runtime/agent_config.rb new file mode 100644 index 0000000..6784529 --- /dev/null +++ b/lib/conductor/agents/runtime/agent_config.rb @@ -0,0 +1,52 @@ +# frozen_string_literal: true + +module Conductor + module Agents + # Runtime knobs, read from CONDUCTOR_AGENT_* environment variables (same names and + # defaults as the Python SDK's AgentConfig). + class AgentConfig + TRUE_VALUES = %w[true 1 yes on].freeze + FALSE_VALUES = %w[false 0 no off].freeze + + attr_accessor :worker_poll_interval_ms, :worker_thread_count, :auto_register_integrations, + :streaming_enabled, :status_poll_interval_seconds, :system_worker_thread_count + + def initialize(worker_poll_interval_ms: 100, worker_thread_count: 1, auto_register_integrations: false, + streaming_enabled: true, status_poll_interval_seconds: 0.5, system_worker_thread_count: 10) + @worker_poll_interval_ms = worker_poll_interval_ms + @worker_thread_count = worker_thread_count + @auto_register_integrations = auto_register_integrations + @streaming_enabled = streaming_enabled + @status_poll_interval_seconds = status_poll_interval_seconds + @system_worker_thread_count = system_worker_thread_count + end + + # @param env [Hash] defaults to ENV + def self.from_env(env = ENV) + new( + worker_poll_interval_ms: int(env, 'CONDUCTOR_AGENT_WORKER_POLL_INTERVAL', 100), + worker_thread_count: int(env, 'CONDUCTOR_AGENT_WORKER_THREADS', 1), + auto_register_integrations: bool(env, 'CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER', false), + streaming_enabled: bool(env, 'CONDUCTOR_AGENT_STREAMING_ENABLED', true) + ) + end + + def self.int(env, key, default) + raw = env[key].to_s.strip + raw.empty? ? default : Integer(raw, 10) + rescue ArgumentError + default + end + + def self.bool(env, key, default) + raw = env[key].to_s.strip.downcase + return default if raw.empty? + return true if TRUE_VALUES.include?(raw) + return false if FALSE_VALUES.include?(raw) + + default + end + private_class_method :int, :bool + end + end +end diff --git a/lib/conductor/agents/runtime/agent_runtime.rb b/lib/conductor/agents/runtime/agent_runtime.rb new file mode 100644 index 0000000..b9aec86 --- /dev/null +++ b/lib/conductor/agents/runtime/agent_runtime.rb @@ -0,0 +1,259 @@ +# frozen_string_literal: true + +require 'securerandom' +require 'logger' +require 'set' +require_relative '../errors' +require_relative '../config_serializer' +require_relative 'agent_config' +require_relative 'execution' +require_relative 'approval_request' +require_relative 'sse_client' +require_relative 'status_poller' +require_relative 'tool_registry' +require_relative '../../client/agent_client' +require_relative '../../worker/task_handler' + +module Conductor + module Agents + # Runs agents against a Conductor server: serializes the agentConfig, starts the + # execution, registers the workers the server asks for, and streams the result. + # + # runtime = Conductor::Agents::AgentRuntime.new(configuration: Conductor::Configuration.new) + # runtime.call_sync(agent, 'Weather in Lisbon?') + # + # Conductor::Agents.runtime holds a default instance built from the environment; + # Agent#call_sync / #call_async use it. + class AgentRuntime + attr_reader :configuration, :agent_config, :client, :api_client, :logger + + def initialize(configuration: nil, agent_config: nil, logger: nil, api_client: nil, agent_client: nil) + @configuration = configuration || Configuration.new + @agent_config = agent_config || AgentConfig.from_env + @logger = logger || Logger.new($stdout, level: Logger::INFO, progname: 'conductor-agents') + @api_client = api_client || Http::ApiClient.new(configuration: @configuration) + @client = agent_client || Client::AgentClient.new(@api_client) + @registry = ToolRegistry.new(@agent_config, logger: @logger) + @handlers = [] + @running_workers = Set.new + @stream_threads = [] + @mutex = Mutex.new + end + + # Run and wait for the answer + # @return [String] + def call_sync(agent, prompt, session_id: nil, timeout: nil, **options) + call_async(agent, prompt, session_id: session_id, **options).result(timeout: timeout) + end + + # Start the agent and return immediately with an Execution + # @param session_id [String, nil] conversation id to continue + # @param media [Array, nil], context [Hash, nil], idempotency_key [String, nil], timeout_seconds [Integer, nil] + # @yield [answer, execution] runs on the stream thread when the execution finishes + # @return [Execution] + def call_async(agent, prompt, session_id: nil, media: nil, context: nil, idempotency_key: nil, + timeout_seconds: nil, &on_done) + payload = start_payload(agent, prompt, session_id: session_id, media: media, context: context, + idempotency_key: idempotency_key, timeout_seconds: timeout_seconds) + response = @client.start_agent(payload) + execution_id = response['executionId'] || raise(Error, "server returned no executionId: #{response.inspect}") + + execution = Execution.new(execution_id, client: @client, agent_name: response['agentName'] || agent.name, runtime: self) + start_workers(agent, response['requiredWorkers'], domain: payload['runId']) + attach(execution, agent: agent, &on_done) + execution + end + + # Register agents on the server without running them + # @return [Array] deployed agent names + def deploy(*agents) + agents.flatten.map do |agent| + response = @client.deploy_agent('agentConfig' => ConfigSerializer.serialize(agent)) + response['agentName'] || agent.name + end + end + + # Compile without registering: { "workflowDef", "requiredWorkers" } + def compile(agent) + @client.compile_agent('agentConfig' => ConfigSerializer.serialize(agent)) + end + + # Deploy, start workers for every tool, and (by default) block until INT/TERM + def serve(*agents, blocking: true) + agents = agents.flatten + agents.each do |agent| + response = @client.deploy_agent('agentConfig' => ConfigSerializer.serialize(agent)) + start_workers(agent, response['requiredWorkers'], domain: nil) + end + return self unless blocking + + wait_for_signal + shutdown + self + end + + # Follow an execution on a background thread (used by call_async and Execution#result) + def attach(execution, agent: nil, &on_done) + execution.attached! + thread = Thread.new do + Thread.current.name = "conductor-agent-stream-#{execution.execution_id}" + follow(execution, agent, &on_done) + end + @mutex.synchronize { @stream_threads << thread } + thread + end + + # Stop workers and stream threads + def shutdown(timeout: 5) + handlers, threads = @mutex.synchronize do + h = @handlers.dup + t = @stream_threads.dup + @handlers.clear + @stream_threads.clear + @running_workers.clear + [h, t] + end + handlers.each { |h| h.stop(timeout: timeout) } + threads.each do |t| + t.join(timeout) + t.kill if t.alive? + end + self + end + + # Names of the workers currently polling + def running_workers + @mutex.synchronize { @running_workers.map(&:first) } + end + + # Build the AgentStartRequest body + def start_payload(agent, prompt, session_id: nil, media: nil, context: nil, idempotency_key: nil, timeout_seconds: nil) + payload = { + 'agentConfig' => ConfigSerializer.serialize(agent), + 'prompt' => prompt.to_s, + 'sessionId' => session_id.to_s, + 'media' => Array(media) + } + payload['context'] = context if context && !context.empty? + payload['idempotencyKey'] = idempotency_key if idempotency_key + payload['timeoutSeconds'] = timeout_seconds if timeout_seconds + payload['runId'] = SecureRandom.hex(16) if agent.stateful_tree? + payload + end + + private + + # Start workers for the tools and system tasks the server requires (skipping ones already polling) + def start_workers(agent, required_workers, domain:) + workers = @registry.workers_for(agent, required_workers: required_workers, domain: domain) + fresh = @mutex.synchronize do + workers.reject { |w| @running_workers.include?([w.task_definition_name, w.domain]) } + .each { |w| @running_workers << [w.task_definition_name, w.domain] } + end + return if fresh.empty? + + handler = Worker::TaskHandler.new(workers: fresh, configuration: @configuration, logger: @logger, + scan_for_annotated_workers: false, register_task_definitions: true) + handler.start + @mutex.synchronize { @handlers << handler } + @logger.info("agent workers started: #{fresh.map(&:task_definition_name).join(', ')}") + end + + def follow(execution, agent, &on_done) + events = event_source(execution.execution_id) + events.each do |event| + handle_event(execution, agent, event) + break if execution.done? + end + execution.fail('stream ended before the execution finished') unless execution.done? + rescue StandardError => e + @logger.error("stream for #{execution.execution_id} failed: #{e.class}: #{e.message}") + execution.fail("#{e.class}: #{e.message}") unless execution.done? + ensure + run_callback(on_done, execution.answer, execution) if on_done + end + + def event_source(execution_id) + poller = StatusPoller.new(@client, interval: @agent_config.status_poll_interval_seconds, logger: @logger) + return poller.each_event(execution_id) unless @agent_config.streaming_enabled + + sse = SseClient.new(@api_client, logger: @logger) + Enumerator.new do |y| + sse.each_event(execution_id) { |ev| y << ev } + rescue SseUnavailableError => e + @logger.info("SSE unavailable (#{e.message}); polling status instead") + poller.each_event(execution_id) { |ev| y << ev } + end + end + + def handle_event(execution, agent, event) + data = event['data'] || {} + execution.record_event(event) + case event['event'].to_s + when 'message' + execution.append_text(data['content']) + when 'tool_call' + execution.add_tool_call(data['toolName'], strip_injected(data['args'])) + when 'tool_result' + execution.add_tool_result(data['toolName'], data['result']) + when 'waiting' + handle_waiting(execution, agent, data) + when 'done' + execution.token_usage = fetch_token_usage(execution.execution_id) + execution.finish(status: 'COMPLETED', output: data['output'] || {}) + when 'error' + execution.finish(status: data['status'] || 'FAILED', output: data['output'] || {}, + reason: data['content'] || 'execution failed') + end + end + + def handle_waiting(execution, agent, data) + pending = data['pendingTool'] || {} + request = ApprovalRequest.new(execution.execution_id, pending, client: @client, execution: execution) + execution.mark_waiting(request) + handler = agent&.approval_handler + return if handler.nil? || request.tool_calls.empty? + + run_callback(handler, request) + end + + def run_callback(callable, *args) + callable.call(*args) + rescue StandardError => e + @logger.error("callback raised #{e.class}: #{e.message}") + end + + def strip_injected(args) + return {} unless args.is_a?(Hash) + + args.reject { |k, _| Dispatch::INJECTED_KEYS.include?(k.to_s) } + end + + # Sum tokenUsage over the execution and its sub-agent executions + def fetch_token_usage(execution_id, visited = Set.new) + return TokenUsage.new if visited.include?(execution_id) || visited.size > 50 + + visited << execution_id + run = @client.get_execution(execution_id) + usage = run['tokenUsage'] || {} + total = TokenUsage.new(prompt_tokens: usage['promptTokens'].to_i, completion_tokens: usage['completionTokens'].to_i, + total_tokens: usage['totalTokens'].to_i) + Array(run['tasks']).each do |task| + sub = task['subWorkflowId'] + total += fetch_token_usage(sub, visited) if sub && !sub.to_s.empty? + end + total + rescue StandardError => e + @logger.debug("token usage unavailable for #{execution_id}: #{e.message}") + TokenUsage.new + end + + def wait_for_signal + queue = Queue.new + %w[INT TERM].each { |sig| trap(sig) { queue << sig } } + @logger.info('serving agents; press Ctrl-C to stop') + queue.pop + end + end + end +end diff --git a/lib/conductor/agents/runtime/approval_request.rb b/lib/conductor/agents/runtime/approval_request.rb new file mode 100644 index 0000000..3e512b7 --- /dev/null +++ b/lib/conductor/agents/runtime/approval_request.rb @@ -0,0 +1,105 @@ +# frozen_string_literal: true + +require_relative '../errors' +require_relative 'execution' + +module Conductor + module Agents + # A tool call (or batch of them) waiting for a human decision. + # + # agent.on_approval do |request| + # request.amount < 100 ? request.approve : request.reject('Needs a manager') + # end + # + # The server pauses on one HUMAN task per turn, so a request may carry several tool + # calls; +tool_name+ and the argument accessors (request.amount) read the first one. + class ApprovalRequest + attr_reader :execution_id, :task_ref_name, :tool_calls, :response_schema, :raw + + # @param pending_tool [Hash] the SSE "waiting" event's pendingTool payload + def initialize(execution_id, pending_tool, client:, execution: nil) + @execution_id = execution_id + @client = client + @execution = execution + @raw = pending_tool || {} + @task_ref_name = @raw['taskRefName'] + @response_schema = @raw['response_schema'] + @tool_calls = extract_tool_calls(@raw) + @responded = false + end + + # First tool call's name + def tool_name + @tool_calls.first&.name + end + + # First tool call's arguments (String keys) + def arguments + @tool_calls.first&.arguments || {} + end + + def responded? + @responded + end + + # Let the tool run + def approve + respond { @client.approve(@execution_id) } + end + + # Skip the tool; the run ends COMPLETED with finish_reason :rejected + def reject(reason = '') + respond { @client.reject(@execution_id, reason) } + end + + # Free-text answer (human tools / feedback) + def send_message(message) + respond { @client.send_message(@execution_id, message) } + end + + # request.amount, request.order_id ... read the first tool call's arguments + def method_missing(name, *args, &block) + key = name.to_s + return arguments[key] if args.empty? && arguments.key?(key) + + super + end + + def respond_to_missing?(name, include_private = false) + arguments.key?(name.to_s) || super + end + + def to_s + calls = @tool_calls.map(&:to_s).join(', ') + "#" + end + alias inspect to_s + + private + + def respond + raise Error, 'approval request already answered' if @responded + + yield + @responded = true + @execution&.clear_waiting + self + end + + def extract_tool_calls(raw) + calls = raw['toolCalls'] || raw['tool_calls'] + if calls.is_a?(Array) && !calls.empty? + return calls.map do |c| + c = c.transform_keys(&:to_s) + ToolCall.new(name: c['name'], arguments: (c['args'] || c['arguments'] || c['parameters'] || {}).transform_keys(&:to_s)) + end + end + + name = raw['tool_name'] || raw['toolName'] + return [] if name.nil? + + [ToolCall.new(name: name, arguments: (raw['parameters'] || raw['args'] || {}).transform_keys(&:to_s))] + end + end + end +end diff --git a/lib/conductor/agents/runtime/dispatch.rb b/lib/conductor/agents/runtime/dispatch.rb new file mode 100644 index 0000000..e62c254 --- /dev/null +++ b/lib/conductor/agents/runtime/dispatch.rb @@ -0,0 +1,137 @@ +# frozen_string_literal: true + +require 'json' +require_relative '../errors' +require_relative 'secrets' +require_relative '../../http/models/task_result' + +module Conductor + module Agents + # Executes one tool task: maps the task's inputData onto the tool method's keyword + # arguments, runs it, and shapes the result for the server. + # + # The server sends the LLM's arguments as top-level keys plus a few injected keys + # (method, _agent_state, _agent_tool_name, _allowed_commands) that are stripped here. + # Missing required arguments, missing credentials and unserializable results are + # terminal failures; anything raised by the tool itself is a retryable failure. + module Dispatch + INJECTED_KEYS = %w[method _agent_state _agent_tool_name _allowed_commands __conductor_agent_ctx__].freeze + WORKER_ID = 'agent-sdk' + TRUE_STRINGS = %w[true 1 yes].freeze + FALSE_STRINGS = %w[false 0 no].freeze + + # Raised when the LLM omitted a required argument + class MissingArgumentError < Error; end + + module_function + + # @param task [Http::Models::Task] + # @param tool_def [ToolDef] + # @return [Http::Models::TaskResult] + def run_tool_task(task, tool_def, logger: nil) + result = base_result(task) + input = (task.input_data || {}).transform_keys(&:to_s) + args = input.except(*INJECTED_KEYS) + + check_credentials!(tool_def) + kwargs = coerce_args(args, tool_def) + output = tool_def.func.call(**kwargs) + return output if output.is_a?(Http::Models::TaskResult) + + result.status = Http::Models::TaskResultStatus::COMPLETED + result.output_data = normalize_output(tool_def, output) + result + rescue MissingArgumentError, CredentialNotFoundError, ToolSerializationError => e + logger&.error("tool #{tool_def.name}: #{e.message}") + terminal_failure(result, e) + rescue StandardError => e + logger&.error("tool #{tool_def.name} raised #{e.class}: #{e.message}") + result.status = Http::Models::TaskResultStatus::FAILED + result.reason_for_incompletion = "#{e.class}: #{e.message}" + result + end + + # Map input keys to the tool's keyword arguments, coercing by the JSON schema + # @return [Hash] + def coerce_args(args, tool_def) + schema = tool_def.input_schema || {} + properties = schema['properties'] || {} + required = Array(schema['required']) + accepts_rest = tool_def.func.respond_to?(:parameters) && tool_def.func.parameters.any? { |kind, _| kind == :keyrest } + + missing = required.reject { |name| args.key?(name) } + raise MissingArgumentError, "tool #{tool_def.name}: missing required argument(s) #{missing.join(', ')}" unless missing.empty? + + args.each_with_object({}) do |(name, value), kwargs| + if properties.key?(name) + kwargs[name.to_sym] = coerce_value(value, properties[name]) + elsif accepts_rest || properties.empty? + kwargs[name.to_sym] = value + end + end + end + + # Coerce one value to its JSON schema type (LLMs often send numbers and JSON as strings) + def coerce_value(value, property) + return value unless property.is_a?(Hash) + + type = property['type'] + case type + when 'integer' + value.is_a?(String) ? (Integer(value, 10) rescue value) : value # rubocop:disable Style/RescueModifier + when 'number' + value.is_a?(String) ? (Float(value) rescue value) : value # rubocop:disable Style/RescueModifier + when 'boolean' + return value unless value.is_a?(String) + + lower = value.strip.downcase + return true if TRUE_STRINGS.include?(lower) + return false if FALSE_STRINGS.include?(lower) + + value + when 'array', 'object' + return value unless value.is_a?(String) + + parsed = JSON.parse(value) + expected = type == 'array' ? Array : Hash + parsed.is_a?(expected) ? parsed : value + when 'string' + value.is_a?(Hash) || value.is_a?(Array) ? JSON.generate(value) : value + else + value + end + rescue JSON::ParserError + value + end + + # Hash results go out as-is; anything else is wrapped as { "result" => value } + def normalize_output(tool_def, output) + data = output.is_a?(Hash) ? output.transform_keys(&:to_s) : { 'result' => output } + begin + JSON.generate(data) + rescue StandardError => e + raise ToolSerializationError, "tool #{tool_def.name} returned a non-JSON-serializable result: #{e.message}" + end + data + end + + def check_credentials!(tool_def) + tool_def.credentials.each { |name| Secrets.secret(name) } + end + + def base_result(task) + result = Http::Models::TaskResult.new + result.task_id = task.task_id + result.workflow_instance_id = task.workflow_instance_id + result.worker_id = WORKER_ID + result + end + + def terminal_failure(result, error) + result.status = Http::Models::TaskResultStatus::FAILED_WITH_TERMINAL_ERROR + result.reason_for_incompletion = error.message + result + end + end + end +end diff --git a/lib/conductor/agents/runtime/execution.rb b/lib/conductor/agents/runtime/execution.rb new file mode 100644 index 0000000..c077851 --- /dev/null +++ b/lib/conductor/agents/runtime/execution.rb @@ -0,0 +1,270 @@ +# frozen_string_literal: true + +require 'timeout' +require_relative '../errors' + +module Conductor + module Agents + # A tool call observed on the stream + ToolCall = Struct.new(:name, :arguments, :result, keyword_init: true) do + def to_s + "#" + end + alias_method :inspect, :to_s + end + + # Token usage for an execution (summed over sub-agent executions) + TokenUsage = Struct.new(:prompt_tokens, :completion_tokens, :total_tokens, keyword_init: true) do + def initialize(prompt_tokens: 0, completion_tokens: 0, total_tokens: 0) + super + end + + def +(other) + TokenUsage.new(prompt_tokens: prompt_tokens + other.prompt_tokens, + completion_tokens: completion_tokens + other.completion_tokens, + total_tokens: total_tokens + other.total_tokens) + end + end + + # Maps the server's status + output.finishReason to a Symbol + module FinishReason + def self.derive(status, output) + case status.to_s + when 'COMPLETED' + fr = output.is_a?(Hash) ? output['finishReason'].to_s : '' + case fr + when 'rejected' then :rejected + when 'LENGTH', 'MAX_TOKENS' then :length + when 'tool_calls', 'TOOL_CALLS' then :tool_calls + when 'CONTENT_FILTER' then :content_filter + else :stop + end + when 'FAILED' then :error + when 'TERMINATED' then :cancelled + when 'TIMED_OUT' then :timeout + else :stop + end + end + end + + # Handle on a running (or finished) agent execution. + # + # execution = agent.call_async('Refund order A-1029') + # execution.done? # false until finished + # execution.partial_text # streamed text so far + # execution.result # blocks for the answer + # execution.finish_reason # :stop | :rejected | ... + class Execution + TERMINAL_STATUSES = %w[COMPLETED FAILED TERMINATED TIMED_OUT].freeze + + attr_reader :execution_id, :agent_name, :tool_calls, :token_usage, :pending, :error, :status, :output, :events + + # @param execution_id [String] also the Conductor workflow id + # @param client [Client::AgentClient] + def initialize(execution_id, client:, agent_name: nil, runtime: nil) + @execution_id = execution_id + @client = client + @agent_name = agent_name + @runtime = runtime + @mutex = Mutex.new + @done_cv = ConditionVariable.new + @partial_text = +'' + @tool_calls = [] + @events = [] + @token_usage = TokenUsage.new + @pending = nil + @status = 'RUNNING' + @output = nil + @result = nil + @error = nil + @waiting = false + @finished = false + end + + # Look up an execution by id (a snapshot; +result+ attaches a stream if still running) + # @return [Execution] + def self.find(execution_id, runtime: Conductor::Agents.runtime) + execution = new(execution_id, client: runtime.client, runtime: runtime) + execution.refresh! + execution + end + + # Re-read status from the server + def refresh! + status = @client.get_status(@execution_id) + @agent_name ||= status['agentName'] + if status['isComplete'] + finish(status: status['status'], output: status['output'], reason: status['reasonForIncompletion']) + elsif status['isWaiting'] + mark_waiting(status['pendingTool']) + end + self + end + + def done? + @mutex.synchronize { @finished } + end + + def waiting? + @mutex.synchronize { @waiting && !@finished } + end + + # Text streamed so far (never behind +result+ once done) + def partial_text + @mutex.synchronize { @partial_text.dup } + end + + # Block until finished and return the answer + # @param timeout [Numeric, nil] seconds; nil waits forever + # @raise [Error] when the execution failed, was cancelled or timed out on the server + # @raise [Timeout::Error] when +timeout+ elapses first + def result(timeout: nil) + @runtime&.attach(self) unless done? || @attached + @mutex.synchronize do + deadline = timeout && (Time.now + timeout) + until @finished + remaining = deadline && (deadline - Time.now) + raise Timeout::Error, "execution #{@execution_id} still running after #{timeout}s" if remaining && remaining <= 0 + + @done_cv.wait(@mutex, remaining) + end + raise Error, "execution #{@execution_id} #{@status.downcase}: #{@error}" if @error && @status != 'COMPLETED' + + @result + end + end + + # The answer without blocking or raising (nil while running or when failed) + def answer + @mutex.synchronize { @finished ? @result : nil } + end + + # @return [Symbol] :stop | :tool_calls | :length | :content_filter | :rejected | :error | :cancelled | :timeout | nil + def finish_reason + @mutex.synchronize { @finished ? FinishReason.derive(@status, @output) : nil } + end + + def rejected? + finish_reason == :rejected + end + + # ── control ───────────────────────────────────────────────────── + + def pause + @client.pause(@execution_id) + self + end + + def resume + @client.resume(@execution_id) + self + end + + def cancel(reason: 'cancelled by client') + @client.cancel(@execution_id, reason: reason) + self + end + + def stop + @client.stop(@execution_id) + self + end + + def signal(message) + @client.signal(@execution_id, message) + self + end + + # Approve / reject the pending tool call, if any + def approve + (pending || raise(Error, 'nothing is waiting for approval')).approve + end + + def reject(reason = '') + (pending || raise(Error, 'nothing is waiting for approval')).reject(reason) + end + + # ── mutators used by the runtime's stream thread ──────────────── + + def attached! + @attached = true + end + + def record_event(event) + @mutex.synchronize { @events << event } + end + + def append_text(text) + return if text.nil? || text.to_s.empty? + + @mutex.synchronize { @partial_text << text.to_s } + end + + def add_tool_call(name, arguments) + @mutex.synchronize { @tool_calls << ToolCall.new(name: name, arguments: arguments || {}) } + end + + def add_tool_result(name, result) + @mutex.synchronize do + call = @tool_calls.reverse.find { |c| c.name == name && c.result.nil? } + call ? call.result = result : @tool_calls << ToolCall.new(name: name, arguments: {}, result: result) + end + end + + def mark_waiting(pending) + @mutex.synchronize do + @waiting = true + @pending = pending + end + end + + def clear_waiting + @mutex.synchronize do + @waiting = false + @pending = nil + end + end + + def token_usage=(usage) + @mutex.synchronize { @token_usage = usage } + end + + # Mark the execution finished. +output+ is the workflow output ({result, finishReason, ...}). + def finish(status:, output:, reason: nil) + @mutex.synchronize do + return if @finished + + @status = status.to_s + @output = output.is_a?(Hash) ? output : {} + @result = extract_result(output) + @error = reason if reason && !reason.to_s.empty? + @error ||= (@output['error'] || @output['reason']) unless @status == 'COMPLETED' + @error ||= "execution #{@status.downcase}" unless @status == 'COMPLETED' + @partial_text = @result.to_s.dup if @result.is_a?(String) && @partial_text.empty? + @waiting = false + @pending = nil + @finished = true + @done_cv.broadcast + end + end + + def fail(reason) + finish(status: 'FAILED', output: { 'error' => reason }, reason: reason) + end + + def to_s + "#" + end + alias inspect to_s + + private + + def extract_result(output) + return output unless output.is_a?(Hash) + return output['result'] if output.key?('result') + + output + end + end + end +end diff --git a/lib/conductor/agents/runtime/sse_client.rb b/lib/conductor/agents/runtime/sse_client.rb new file mode 100644 index 0000000..1f62086 --- /dev/null +++ b/lib/conductor/agents/runtime/sse_client.rb @@ -0,0 +1,175 @@ +# frozen_string_literal: true + +require 'net/http' +require 'uri' +require 'json' +require 'logger' +require_relative '../errors' + +module Conductor + module Agents + # Server-sent events from GET /api/agent/stream/{executionId}. + # + # Uses a plain Net::HTTP streaming request (the shared Faraday RestClient buffers + # bodies, retries, and has a 120 s total timeout, none of which suit a long-lived + # stream). Auth headers come from the ApiClient so token refresh stays in one place. + # + # Wire format (see AgentStreamRegistry on the server): ":connected" first, then + # "id:\nevent:\ndata:\n\n" frames, a ":heartbeat" comment every 15 s, the + # stream closes after "done" or "error". Last-Event-ID (a bare integer) resumes; without + # it the server replays from the start, so connecting after start loses nothing. + class SseClient + HEARTBEAT_ONLY_TIMEOUT = 15 + RECONNECT_DELAY = 1 + TERMINAL_EVENTS = %w[done error].freeze + READ_TIMEOUT = 60 + OPEN_TIMEOUT = 5 + + # Incremental parser for the SSE wire format + class Parser + def initialize + @buffer = +'' + @event = nil + @id = nil + @data = [] + end + + # Feed a chunk; yields each complete event as { 'event', 'id', 'data' } or { 'heartbeat' => true } + def feed(chunk, &block) + @buffer << chunk + while (idx = @buffer.index("\n")) + line = @buffer.slice!(0..idx).chomp + process_line(line, &block) + end + end + + private + + def process_line(line, &block) + if line.start_with?(':') + yield({ 'heartbeat' => true }) + elsif line.empty? + flush(&block) + elsif (m = line.match(/\A(\w+):\s?(.*)\z/m)) + field = m[1] + value = m[2] + case field + when 'event' then @event = value + when 'id' then @id = value + when 'data' then @data << value + end + end + end + + def flush + return if @data.empty? && @event.nil? + + raw = @data.join("\n") + data = begin + raw.empty? ? {} : JSON.parse(raw) + rescue JSON::ParserError + { 'content' => raw } + end + data = { 'content' => data } unless data.is_a?(Hash) + id = @id.to_s =~ /\A\d+\z/ ? @id.to_i : @id + yield({ 'event' => @event || data['type'], 'id' => id, 'data' => data }) + ensure + @event = nil + @id = nil + @data = [] + end + end + + # @param api_client [Http::ApiClient] supplies base URL, TLS settings and auth headers + def initialize(api_client, logger: nil) + @api_client = api_client + @configuration = api_client.configuration + @logger = logger || Logger.new($stdout, level: Logger::INFO) + end + + # Yield every real event for +execution_id+ until done/error, reconnecting on drops. + # @param last_event_id [Integer, nil] resume point + # @raise [SseUnavailableError] when the first connection fails or only heartbeats arrive + def each_event(execution_id, last_event_id: nil) + return enum_for(:each_event, execution_id, last_event_id: last_event_id) unless block_given? + + first_connect = true + got_real_event = false + + loop do + begin + finished = connect(execution_id, last_event_id) do |event| + if event['heartbeat'] + next unless !got_real_event && Time.now - @connected_at > HEARTBEAT_ONLY_TIMEOUT + + raise SseUnavailableError, "SSE connected but only heartbeats arrived for #{HEARTBEAT_ONLY_TIMEOUT}s" + end + first_connect = false + got_real_event = true + last_event_id = event['id'] if event['id'].is_a?(Integer) + yield event + return if TERMINAL_EVENTS.include?(event['event'].to_s) + end + first_connect = false + return if finished + rescue SseUnavailableError + raise + rescue StandardError => e + raise SseUnavailableError, "SSE unavailable: #{e.class}: #{e.message}" if first_connect + + @logger.warn("SSE connection lost (#{e.class}: #{e.message}), reconnecting in #{RECONNECT_DELAY}s") + end + sleep RECONNECT_DELAY + end + end + + private + + # Open one streaming request. Returns true when a terminal event was seen, false on EOF. + def connect(execution_id, last_event_id) + uri = URI.parse("#{@configuration.server_url}/agent/stream/#{execution_id}") + request = Net::HTTP::Get.new(uri) + request['Accept'] = 'text/event-stream' + request['Cache-Control'] = 'no-cache' + request['Last-Event-ID'] = last_event_id.to_s if last_event_id + auth_headers.each { |k, v| request[k] = v } + + parser = Parser.new + terminal = false + Net::HTTP.start(uri.host, uri.port, **http_options(uri)) do |http| + http.request(request) do |response| + raise SseUnavailableError, "SSE endpoint returned HTTP #{response.code}" unless response.code.to_i == 200 + + @connected_at = Time.now + response.read_body do |chunk| + parser.feed(chunk) do |event| + yield event + terminal = true if TERMINAL_EVENTS.include?(event['event'].to_s) + end + break if terminal + end + end + end + terminal + end + + def auth_headers + return {} unless @configuration.auth_configured? + + @api_client.get_authentication_headers || {} + rescue StandardError => e + @logger.warn("Could not attach auth headers to SSE request: #{e.message}") + {} + end + + def http_options(uri) + opts = { use_ssl: uri.scheme == 'https', read_timeout: READ_TIMEOUT, open_timeout: OPEN_TIMEOUT } + if opts[:use_ssl] + opts[:verify_mode] = @configuration.verify_ssl ? OpenSSL::SSL::VERIFY_PEER : OpenSSL::SSL::VERIFY_NONE + opts[:ca_file] = @configuration.ssl_ca_cert if @configuration.ssl_ca_cert + end + opts + end + end + end +end diff --git a/lib/conductor/agents/runtime/status_poller.rb b/lib/conductor/agents/runtime/status_poller.rb new file mode 100644 index 0000000..60087ec --- /dev/null +++ b/lib/conductor/agents/runtime/status_poller.rb @@ -0,0 +1,49 @@ +# frozen_string_literal: true + +module Conductor + module Agents + # Polling fallback for servers without SSE: turns GET /agent/{id}/status into the + # same event hashes the SseClient yields (waiting, done, error). No partial text. + class StatusPoller + def initialize(client, interval: 0.5, logger: nil) + @client = client + @interval = interval + @logger = logger + end + + def each_event(execution_id) + return enum_for(:each_event, execution_id) unless block_given? + + was_waiting = false + loop do + status = @client.get_status(execution_id) + if status['isComplete'] + yield terminal_event(status) + return + end + + if status['isWaiting'] && !was_waiting + yield({ 'event' => 'waiting', 'id' => nil, 'data' => { 'type' => 'waiting', 'executionId' => execution_id, + 'pendingTool' => status['pendingTool'] || {} } }) + end + was_waiting = status['isWaiting'] ? true : false + sleep(was_waiting ? [@interval * 4, 2.0].min : @interval) + end + end + + private + + def terminal_event(status) + if status['status'].to_s == 'COMPLETED' + { 'event' => 'done', 'id' => nil, + 'data' => { 'type' => 'done', 'executionId' => status['executionId'], 'output' => status['output'] || {} } } + else + { 'event' => 'error', 'id' => nil, + 'data' => { 'type' => 'error', 'executionId' => status['executionId'], 'status' => status['status'], + 'content' => status['reasonForIncompletion'] || "execution #{status['status']}", + 'output' => status['output'] } } + end + end + end + end +end diff --git a/lib/conductor/agents/runtime/system_workers.rb b/lib/conductor/agents/runtime/system_workers.rb new file mode 100644 index 0000000..b2cabef --- /dev/null +++ b/lib/conductor/agents/runtime/system_workers.rb @@ -0,0 +1,105 @@ +# frozen_string_literal: true + +require 'json' +require_relative '../errors' + +module Conductor + module Agents + # Bodies for the compiler-generated SIMPLE tasks the server asks the SDK to serve + # (they appear in requiredWorkers next to the user's tools). Ports of the Python + # SDK's TerminationEntry, GuardrailEntry, CallbackEntry and OnCondition handoff workers. + module SystemWorkers + module_function + + # _termination: { should_continue, reason } + def termination(condition, logger: nil) + lambda do |task| + input = stringify(task.input_data) + context = { 'result' => input['result'].to_s, 'messages' => input['messages'] || [], + 'iteration' => input['iteration'].to_i, 'token_usage' => input['token_usage'] } + begin + outcome = condition.should_terminate(context) + { 'should_continue' => !outcome.should_terminate, 'reason' => outcome.reason.to_s } + rescue StandardError => e + logger&.error("termination condition failed: #{e.class}: #{e.message}") + { 'should_continue' => true, 'reason' => '' } + end + end + end + + # : { passed, message, on_fail, fixed_output, guardrail_name, should_continue } + def guardrail(guardrail, logger: nil) + lambda do |task| + input = stringify(task.input_data) + content = stringify_content(input['content']) + iteration = input['iteration'].to_i + begin + result = guardrail.check(content) + return pass_result if result.passed? + + on_fail = guardrail.on_fail + fixed = result.fixed_output + on_fail = 'raise' if on_fail == 'retry' && iteration >= guardrail.max_retries + on_fail = 'raise' if on_fail == 'fix' && fixed.nil? + { 'passed' => false, 'message' => result.message.to_s, 'on_fail' => on_fail, 'fixed_output' => fixed, + 'guardrail_name' => guardrail.name, 'should_continue' => on_fail == 'retry' } + rescue StandardError => e + logger&.error("guardrail #{guardrail.name} raised: #{e.class}: #{e.message}") + on_fail = guardrail.on_fail + on_fail = 'raise' if on_fail == 'retry' && iteration >= guardrail.max_retries + { 'passed' => false, 'message' => "Guardrail error: #{e.message}", 'on_fail' => on_fail, 'fixed_output' => nil, + 'guardrail_name' => guardrail.name, 'should_continue' => on_fail == 'retry' } + end + end + end + + # _: the callback chain's Hash (or {}) + def callback(chain, logger: nil) + lambda do |task| + input = stringify(task.input_data) + kwargs = {} + kwargs[:messages] = input['messages'] if input.key?('messages') + kwargs[:llm_result] = input['llm_result'] if input.key?('llm_result') + begin + result = chain.call(**kwargs) + result.is_a?(Hash) ? result : {} + rescue StandardError => e + logger&.error("callback failed: #{e.class}: #{e.message}") + {} + end + end + end + + # _handoff_: { handoff, target } + def handoff(condition, logger: nil) + lambda do |task| + input = stringify(task.input_data) + begin + { 'handoff' => condition.should_handoff(input) ? true : false, 'target' => condition.target } + rescue StandardError => e + logger&.error("handoff condition failed: #{e.class}: #{e.message}") + { 'handoff' => false, 'target' => condition.target } + end + end + end + + def pass_result + { 'passed' => true, 'message' => '', 'on_fail' => 'pass', 'fixed_output' => nil, 'guardrail_name' => '', + 'should_continue' => false } + end + + def stringify(input) + (input || {}).transform_keys(&:to_s) + end + + def stringify_content(content) + return '' if content.nil? + return content if content.is_a?(String) + + JSON.generate(content) + rescue StandardError + content.to_s + end + end + end +end diff --git a/lib/conductor/agents/runtime/tool_registry.rb b/lib/conductor/agents/runtime/tool_registry.rb new file mode 100644 index 0000000..76a5979 --- /dev/null +++ b/lib/conductor/agents/runtime/tool_registry.rb @@ -0,0 +1,145 @@ +# frozen_string_literal: true + +require 'set' +require_relative '../errors' +require_relative 'dispatch' +require_relative 'system_workers' +require_relative '../../worker/worker' +require_relative '../../http/models/task_def' + +module Conductor + module Agents + # Turns an agent tree into the Conductor workers this process must run: one per + # local tool (worker/cli tools with a func) and one per compiler-generated system task + # the server listed in requiredWorkers. + class ToolRegistry + SYSTEM_SUFFIX_TERMINATION = '_termination' + + attr_reader :logger + + def initialize(agent_config, logger: nil) + @agent_config = agent_config + @logger = logger + end + + # @param agent [Agent] root agent + # @param required_workers [Array, nil] from the start/deploy response (nil = register everything) + # @param domain [String, nil] task domain for stateful runs (the runId) + # @return [Array] + def workers_for(agent, required_workers: nil, domain: nil) + required = required_workers.nil? ? nil : Set.new(required_workers.map(&:to_s)) + workers = tool_workers(agent, domain: domain) + workers += system_workers(agent, required, domain: domain) + warn_unhandled(required, workers) + workers + end + + # Workers for every local tool in the tree (deduplicated by name) + def tool_workers(agent, domain: nil) + seen = {} + agent.all_agents.each do |a| + a.tools.each do |tool_def| + next unless tool_def.local? + next if seen.key?(tool_def.name) + + seen[tool_def.name] = build_tool_worker(tool_def, a, domain: domain) + end + end + seen.values + end + + # Workers for compiler-generated tasks: termination, custom guardrails, callbacks, on_condition handoffs + def system_workers(agent, required, domain: nil) + workers = [] + agent.all_agents.each do |a| + if a.termination + name = "#{a.name}#{SYSTEM_SUFFIX_TERMINATION}" + workers << build_system_worker(name, SystemWorkers.termination(a.termination, logger: @logger), domain, a) if wanted?(required, name) + end + a.guardrails.each do |g| + next if g.external? || g.is_a?(RegexGuardrail) || g.is_a?(LlmGuardrail) + + workers << build_system_worker(g.name, SystemWorkers.guardrail(g, logger: @logger), domain, a) if wanted?(required, g.name) + end + a.callback_positions.each do |position| + name = "#{a.name}_#{position}" + next unless wanted?(required, name) + + workers << build_system_worker(name, SystemWorkers.callback(a.callback_chain(position, logger: @logger), logger: @logger), domain, a) + end + a.handoffs.each do |h| + next unless h.is_a?(Handoff::OnCondition) + + name = "#{a.name}_handoff_#{h.target}" + workers << build_system_worker(name, SystemWorkers.handoff(h, logger: @logger), domain, a) if wanted?(required, name) + end + end + workers.uniq(&:task_definition_name) + end + + # Task definition for a tool worker (same defaults as the Python SDK) + def task_def_for(name, retry_count: 2, retry_delay_seconds: 2, retry_logic: 'LINEAR_BACKOFF', credentials: []) + Http::Models::TaskDef.new( + name: name, + retry_count: retry_count, + retry_delay_seconds: retry_delay_seconds, + retry_logic: retry_logic, + timeout_seconds: 0, + response_timeout_seconds: 10, + timeout_policy: 'RETRY', + runtime_metadata: credentials.uniq + ) + end + + private + + def wanted?(required, name) + required.nil? || required.include?(name) + end + + def build_tool_worker(tool_def, agent, domain:) + credentials = tool_def.credentials + agent.credentials + Worker::Worker.new( + tool_def.name, + ->(task) { Dispatch.run_tool_task(task, tool_def, logger: @logger) }, + **worker_options(domain, @agent_config.worker_thread_count), + task_def_template: task_def_for(tool_def.name, retry_count: tool_def.retry_count, + retry_delay_seconds: tool_def.retry_delay_seconds, + retry_logic: tool_def.retry_logic, credentials: credentials) + ) + end + + def build_system_worker(name, body, domain, agent) + Worker::Worker.new( + name, body, + **worker_options(domain, @agent_config.system_worker_thread_count), + task_def_template: task_def_for(name, credentials: agent.credentials) + ) + end + + # When a run has a runId the server maps every required worker to that domain, so all + # workers of the run poll on it. + def worker_options(domain, thread_count) + { + register_task_def: true, + overwrite_task_def: true, + lease_extend_enabled: true, + poll_interval: @agent_config.worker_poll_interval_ms, + thread_count: thread_count, + domain: domain + } + end + + def warn_unhandled(required, workers) + return if required.nil? + + handled = workers.map(&:task_definition_name) + missing = required.to_a - handled + return if missing.empty? + + @logger&.warn("server requires workers this process does not provide: #{missing.join(', ')} " \ + '(tasks of these types will stay SCHEDULED until some worker serves them)') + end + end + end +end diff --git a/spec/conductor/agents/agent_config_spec.rb b/spec/conductor/agents/agent_config_spec.rb new file mode 100644 index 0000000..1a997d6 --- /dev/null +++ b/spec/conductor/agents/agent_config_spec.rb @@ -0,0 +1,30 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::AgentConfig do + it 'has the Python defaults' do + c = described_class.from_env({}) + expect(c.worker_poll_interval_ms).to eq(100) + expect(c.worker_thread_count).to eq(1) + expect(c.auto_register_integrations).to be false + expect(c.streaming_enabled).to be true + end + + it 'reads CONDUCTOR_AGENT_* variables with Python boolean parsing' do + env = { 'CONDUCTOR_AGENT_WORKER_POLL_INTERVAL' => '250', 'CONDUCTOR_AGENT_WORKER_THREADS' => '4', + 'CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER' => 'yes', 'CONDUCTOR_AGENT_STREAMING_ENABLED' => 'off' } + c = described_class.from_env(env) + expect(c.worker_poll_interval_ms).to eq(250) + expect(c.worker_thread_count).to eq(4) + expect(c.auto_register_integrations).to be true + expect(c.streaming_enabled).to be false + end + + it 'ignores blank and garbage values' do + c = described_class.from_env('CONDUCTOR_AGENT_WORKER_THREADS' => 'lots', 'CONDUCTOR_AGENT_STREAMING_ENABLED' => ' ') + expect(c.worker_thread_count).to eq(1) + expect(c.streaming_enabled).to be true + end +end diff --git a/spec/conductor/agents/agent_runtime_spec.rb b/spec/conductor/agents/agent_runtime_spec.rb new file mode 100644 index 0000000..b2f5548 --- /dev/null +++ b/spec/conductor/agents/agent_runtime_spec.rb @@ -0,0 +1,216 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'support/agent_tools' + +RSpec.describe Conductor::Agents::AgentRuntime do + a = Conductor::Agents + let(:configuration) { Conductor::Configuration.new(server_api_url: 'http://localhost:8080/api') } + let(:api_client) { instance_double(Conductor::Http::ApiClient, configuration: configuration) } + let(:client) { instance_double(Conductor::Client::AgentClient) } + let(:handler) { instance_double(Conductor::Worker::TaskHandler, start: nil, stop: nil) } + let(:runtime) do + described_class.new(configuration: configuration, agent_config: a::AgentConfig.new, logger: Logger.new(nil), + api_client: api_client, agent_client: client) + end + let(:agent) do + ag = a::Agent.new(name: 'weather', model: 'openai/gpt-4o-mini', instructions: 'Answer weather questions.') + ag.add_tool :current + ag + end + let(:done_output) { { 'result' => 'Sunny in Lisbon, 21C.', 'finishReason' => 'STOP', 'context' => {}, 'rejectionReason' => nil } } + + def event(type, data = {}) + { 'event' => type, 'id' => 1, 'data' => { 'type' => type, 'executionId' => 'EXEC_1' }.merge(data) } + end + + def stub_stream(*events) + sse = instance_double(Conductor::Agents::SseClient) + allow(Conductor::Agents::SseClient).to receive(:new).and_return(sse) + allow(sse).to receive(:each_event) { |_id, &blk| events.each { |e| blk.call(e) } } + sse + end + + before do + allow(Conductor::Worker::TaskHandler).to receive(:new).and_return(handler) + allow(client).to receive(:get_execution).and_return('tokenUsage' => { 'promptTokens' => 196, 'completionTokens' => 50, + 'totalTokens' => 246 }, 'tasks' => []) + end + + after { runtime.shutdown } + + describe '#call_async / #call_sync' do + it 'starts the agent with the serialized config, registers the required workers and streams to the answer' do + expect(client).to receive(:start_agent) do |payload| + expect(payload['agentConfig']).to eq(a::ConfigSerializer.serialize(agent)) + expect(payload['prompt']).to eq('Weather in Lisbon?') + expect(payload['sessionId']).to eq('') + expect(payload['media']).to eq([]) + expect(payload).not_to have_key('runId') + { 'executionId' => 'EXEC_1', 'agentName' => 'weather', 'requiredWorkers' => ['current'] } + end + expect(Conductor::Worker::TaskHandler).to receive(:new) do |workers:, **_| + expect(workers.map(&:task_definition_name)).to eq(['current']) + handler + end + stub_stream(event('thinking', 'content' => 'weather_llm__1'), + event('tool_call', 'toolName' => 'current', 'args' => { 'method' => 'current', '_agent_state' => {}, 'city' => 'Lisbon' }), + event('tool_result', 'toolName' => 'current', 'result' => { 'temp_c' => 21.0 }), + event('done', 'output' => done_output)) + + answers = [] + execution = runtime.call_async(agent, 'Weather in Lisbon?') { |answer| answers << answer } + expect(execution.result(timeout: 5)).to eq('Sunny in Lisbon, 21C.') + expect(execution.finish_reason).to eq(:stop) + expect(execution.tool_calls.first.arguments).to eq('city' => 'Lisbon') + expect(execution.tool_calls.first.result).to eq('temp_c' => 21.0) + expect(execution.token_usage.total_tokens).to eq(246) + expect(runtime.running_workers).to eq(['current']) + sleep 0.05 until execution.done? && !answers.empty? + expect(answers).to eq(['Sunny in Lisbon, 21C.']) + end + + it 'call_sync returns the answer and passes session_id through' do + expect(client).to receive(:start_agent).with(hash_including('sessionId' => 'cust-77')) + .and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) + stub_stream(event('done', 'output' => done_output)) + expect(runtime.call_sync(agent, 'hi', session_id: 'cust-77', timeout: 5)).to eq('Sunny in Lisbon, 21C.') + end + + it 'invokes on_approval with an ApprovalRequest and posts the decision' do + support = a::Agent.new(name: 'support', model: 'anthropic/claude-sonnet-4-5') + support.add_tool :issue_refund + decisions = [] + support.on_approval do |req| + decisions << req.amount + req.amount < 100 ? req.approve : req.reject('Needs a manager') + end + + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => ['issue_refund']) + expect(client).to receive(:approve).with('EXEC_1') + stub_stream(event('waiting', 'pendingTool' => { 'taskRefName' => 'support_approval_human', + 'toolCalls' => [{ 'name' => 'issue_refund', 'args' => { 'order_id' => 'A-1029', 'amount' => 49.0 } }] }), + event('done', 'output' => done_output.merge('result' => 'Refunded $49.'))) + + expect(runtime.call_sync(support, 'Refund order A-1029', timeout: 5)).to eq('Refunded $49.') + expect(decisions).to eq([49.0]) + end + + it 'parks the request on execution.pending when no on_approval handler exists and reports rejection' do + support = a::Agent.new(name: 'support', model: 'anthropic/claude-sonnet-4-5') + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) + gate = Queue.new + sse = instance_double(a::SseClient) + allow(a::SseClient).to receive(:new).and_return(sse) + allow(sse).to receive(:each_event) do |_id, &blk| + blk.call(event('waiting', 'pendingTool' => { 'toolCalls' => [{ 'name' => 'issue_refund', 'args' => { 'amount' => 500 } }] })) + gate.pop + blk.call(event('done', 'output' => { 'result' => nil, 'finishReason' => 'rejected', 'rejectionReason' => 'Needs a manager' })) + end + + execution = runtime.call_async(support, 'Refund order A-1029') + sleep 0.01 until execution.waiting? + expect(execution.pending.tool_name).to eq('issue_refund') + expect(client).to receive(:reject).with('EXEC_1', 'Needs a manager') + execution.reject('Needs a manager') + gate << :go + expect(execution.result(timeout: 5)).to be_nil + expect(execution.finish_reason).to eq(:rejected) + end + + it 'logs and survives exceptions raised in user callbacks' do + support = a::Agent.new(name: 'support', model: 'm/x') + support.on_approval { |_req| raise 'user bug' } + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) + stub_stream(event('waiting', 'pendingTool' => { 'toolCalls' => [{ 'name' => 't', 'args' => {} }] }), + event('done', 'output' => done_output)) + execution = runtime.call_async(support, 'x') { |_| raise 'on_done bug' } + expect(execution.result(timeout: 5)).to eq('Sunny in Lisbon, 21C.') + end + + it 'falls back to status polling when SSE is unavailable' do + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) + sse = instance_double(a::SseClient) + allow(a::SseClient).to receive(:new).and_return(sse) + allow(sse).to receive(:each_event).and_raise(a::SseUnavailableError, 'no sse') + allow(client).to receive(:get_status).and_return('executionId' => 'EXEC_1', 'status' => 'COMPLETED', + 'isComplete' => true, 'output' => done_output) + expect(runtime.call_sync(agent, 'hi', timeout: 5)).to eq('Sunny in Lisbon, 21C.') + end + + it 'uses polling when streaming is disabled and surfaces server errors' do + quiet = described_class.new(configuration: configuration, agent_config: a::AgentConfig.new(streaming_enabled: false), + logger: Logger.new(nil), api_client: api_client, agent_client: client) + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) + allow(client).to receive(:get_status).and_return('executionId' => 'EXEC_1', 'status' => 'FAILED', 'isComplete' => true, + 'reasonForIncompletion' => 'model quota exceeded') + expect { quiet.call_sync(agent, 'hi', timeout: 5) }.to raise_error(a::Error, /model quota exceeded/) + quiet.shutdown + end + + it 'sends a runId and uses it as the worker domain for stateful agents' do + stateful = a::Agent.new(name: 'notes', model: 'm/x', stateful: true) + stateful.add_tool :current + expect(client).to receive(:start_agent) do |payload| + expect(payload['runId']).to match(/\A[0-9a-f]{32}\z/) + { 'executionId' => 'EXEC_1', 'requiredWorkers' => ['current'] } + end + expect(Conductor::Worker::TaskHandler).to receive(:new) do |workers:, **_| + expect(workers.first.domain).to match(/\A[0-9a-f]{32}\z/) + handler + end + stub_stream(event('done', 'output' => done_output)) + runtime.call_sync(stateful, 'remember this', timeout: 5) + end + + it 'does not start a second handler for workers that are already polling' do + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => ['current']) + stub_stream(event('done', 'output' => done_output)) + runtime.call_sync(agent, 'a', timeout: 5) + runtime.call_sync(agent, 'b', timeout: 5) + expect(Conductor::Worker::TaskHandler).to have_received(:new).once + end + end + + describe '#deploy / #compile / #serve' do + it 'deploys each agent and returns the names' do + expect(client).to receive(:deploy_agent).with(hash_including('agentConfig' => hash_including('name' => 'weather'))) + .and_return('agentName' => 'weather', 'requiredWorkers' => ['current']) + expect(runtime.deploy(agent)).to eq(['weather']) + end + + it 'compiles without registering' do + expect(client).to receive(:compile_agent).and_return('workflowDef' => {}, 'requiredWorkers' => []) + expect(runtime.compile(agent)).to include('workflowDef') + end + + it 'serve deploys and starts workers without blocking when asked' do + allow(client).to receive(:deploy_agent).and_return('agentName' => 'weather', 'requiredWorkers' => ['current']) + runtime.serve(agent, blocking: false) + expect(runtime.running_workers).to eq(['current']) + end + end + + describe '#shutdown' do + it 'stops handlers and forgets running workers' do + allow(client).to receive(:deploy_agent).and_return('agentName' => 'weather', 'requiredWorkers' => ['current']) + runtime.serve(agent, blocking: false) + runtime.shutdown + expect(handler).to have_received(:stop) + expect(runtime.running_workers).to eq([]) + end + end +end + +RSpec.describe Conductor::Agents, '.runtime' do + after { described_class.shutdown } + + it 'memoizes a default runtime and can be reconfigured' do + allow(Conductor::Http::ApiClient).to receive(:new).and_return(instance_double(Conductor::Http::ApiClient)) + first = described_class.runtime + expect(described_class.runtime).to equal(first) + reconfigured = described_class.configure(configuration: Conductor::Configuration.new(server_api_url: 'http://x/api')) + expect(reconfigured).not_to equal(first) + expect(described_class.runtime).to equal(reconfigured) + end +end diff --git a/spec/conductor/agents/approval_request_spec.rb b/spec/conductor/agents/approval_request_spec.rb new file mode 100644 index 0000000..d5f4bc4 --- /dev/null +++ b/spec/conductor/agents/approval_request_spec.rb @@ -0,0 +1,46 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::ApprovalRequest do + let(:client) { instance_double(Conductor::Client::AgentClient) } + let(:pending_tool) do + { 'taskRefName' => 'support_approval_human', 'tool_name' => nil, 'parameters' => nil, + 'toolCalls' => [{ 'name' => 'issue_refund', 'args' => { 'order_id' => 'A-1029', 'amount' => 49.0 } }], + 'response_schema' => { 'type' => 'object', 'required' => ['approved'] } } + end + let(:request) { described_class.new('EXEC_1', pending_tool, client: client) } + + it 'reads the batch of tool calls and exposes the first one' do + expect(request.tool_calls.map(&:name)).to eq(['issue_refund']) + expect(request.tool_name).to eq('issue_refund') + expect(request.arguments).to eq('order_id' => 'A-1029', 'amount' => 49.0) + expect(request.amount).to eq(49.0) + expect(request.order_id).to eq('A-1029') + expect(request.respond_to?(:amount)).to be true + expect { request.nonexistent }.to raise_error(NoMethodError) + expect(request.task_ref_name).to eq('support_approval_human') + end + + it 'also understands the singular tool_name/parameters shape' do + single = described_class.new('E', { 'tool_name' => 'ask', 'parameters' => { 'q' => 1 } }, client: client) + expect(single.tool_name).to eq('ask') + expect(single.q).to eq(1) + end + + it 'approves, rejects and sends messages once' do + execution = instance_double(Conductor::Agents::Execution, clear_waiting: nil) + req = described_class.new('EXEC_1', pending_tool, client: client, execution: execution) + expect(client).to receive(:approve).with('EXEC_1') + req.approve + expect(req.responded?).to be true + expect(execution).to have_received(:clear_waiting) + expect { req.approve }.to raise_error(Conductor::Agents::Error, /already/) + + expect(client).to receive(:reject).with('EXEC_1', 'Needs a manager') + described_class.new('EXEC_1', pending_tool, client: client).reject('Needs a manager') + expect(client).to receive(:send_message).with('EXEC_1', 'hello') + described_class.new('EXEC_1', pending_tool, client: client).send_message('hello') + end +end diff --git a/spec/conductor/agents/dispatch_spec.rb b/spec/conductor/agents/dispatch_spec.rb new file mode 100644 index 0000000..cee0e19 --- /dev/null +++ b/spec/conductor/agents/dispatch_spec.rb @@ -0,0 +1,124 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'support/agent_tools' + +RSpec.describe Conductor::Agents::Dispatch do + # Poll body recorded in conductor-mocks agent/tool_happy_path + let(:recorded_input) do + { '_agent_tool_name' => 'get_weather', '_agent_state' => {}, 'method' => 'get_weather', + 'city' => 'Lisbon', 'units' => 'metric' } + end + + def task_with(input, runtime_metadata: {}) + Conductor::Http::Models::Task.from_hash('taskId' => 'TASK_1', 'workflowInstanceId' => 'EXEC_1', + 'taskType' => 'get_weather', 'inputData' => input, + 'runtimeMetadata' => runtime_metadata) + end + + def with_context(task) + result = Conductor::Http::Models::TaskResult.new + Conductor::Worker::TaskContext.current = Conductor::Worker::TaskContext.new(task, result) + yield + ensure + Conductor::Worker::TaskContext.clear + end + + it 'strips the injected keys and runs the tool with keyword arguments' do + result = described_class.run_tool_task(task_with(recorded_input), SpecTools::Weather[:current]) + expect(result.status).to eq('COMPLETED') + expect(result.worker_id).to eq('agent-sdk') + expect(result.task_id).to eq('TASK_1') + expect(result.workflow_instance_id).to eq('EXEC_1') + expect(result.output_data).to eq('temp_c' => 21.0, 'summary' => 'Sunny in Lisbon (metric)') + end + + it 'fails terminally when a required argument is missing' do + result = described_class.run_tool_task(task_with({ 'units' => 'metric' }), SpecTools::Weather[:current]) + expect(result.status).to eq('FAILED_WITH_TERMINAL_ERROR') + expect(result.reason_for_incompletion).to include('city') + end + + it 'coerces strings to the schema types and JSON to strings' do + td = SpecTools::Weather[:forecast] + input = { 'city' => 'Porto', 'days' => '5', 'detailed' => 'yes', 'tags' => '["a","b"]', 'ratio' => '0.25', + 'opts' => '{"k":1}', 'mode' => 'full' } + result = described_class.run_tool_task(task_with(input), td) + expect(result.output_data).to include('days' => 5, 'detailed' => true, 'tags' => %w[a b], 'ratio' => 0.25, + 'opts' => { 'k' => 1 }, 'mode' => 'full') + end + + it 'wraps scalar results and keeps _state_updates' do + scalar = Conductor::Agents::ToolDef.new(name: 's', func: ->(**) { 'plain' }, + input_schema: { 'type' => 'object', 'properties' => {} }) + expect(described_class.run_tool_task(task_with({}), scalar).output_data).to eq('result' => 'plain') + + stateful = Conductor::Agents::ToolDef.new(name: 's2', func: ->(**) { { ok: true, _state_updates: { 'n' => 1 } } }, + input_schema: { 'type' => 'object', 'properties' => {} }) + expect(described_class.run_tool_task(task_with({}), stateful).output_data).to eq('ok' => true, '_state_updates' => { 'n' => 1 }) + end + + it 'reports tool exceptions as retryable failures with the reason' do + boom = Conductor::Agents::ToolDef.new(name: 'boom', func: ->(**) { raise 'kaput' }, + input_schema: { 'type' => 'object', 'properties' => {} }) + result = described_class.run_tool_task(task_with({}), boom, logger: Logger.new(nil)) + expect(result.status).to eq('FAILED') + expect(result.reason_for_incompletion).to eq('RuntimeError: kaput') + end + + it 'fails terminally on unserializable results' do + bad = Conductor::Agents::ToolDef.new(name: 'bad', func: ->(**) { { io: $stdout } }, + input_schema: { 'type' => 'object', 'properties' => {} }) + allow(JSON).to receive(:generate).and_raise(JSON::GeneratorError, 'nope') + result = described_class.run_tool_task(task_with({}), bad) + expect(result.status).to eq('FAILED_WITH_TERMINAL_ERROR') + end + + it 'reads declared secrets from the task runtimeMetadata via TaskContext' do + task = task_with({ 'title' => 'bug', 'method' => 'create_issue' }, runtime_metadata: { 'GH_TOKEN' => 'ghp_x' }) + with_context(task) do + result = described_class.run_tool_task(task, SpecTools::Github[:create_issue]) + expect(result.status).to eq('COMPLETED') + expect(result.output_data['token']).to eq('ghp_x') + end + end + + it 'falls back to ENV and fails terminally when a declared secret is missing everywhere' do + task = task_with({ 'title' => 'bug' }) + with_context(task) do + ENV['GH_TOKEN'] = 'from_env' + expect(described_class.run_tool_task(task, SpecTools::Github[:create_issue]).output_data['token']).to eq('from_env') + ensure + ENV.delete('GH_TOKEN') + end + with_context(task) do + result = described_class.run_tool_task(task, SpecTools::Github[:create_issue]) + expect(result.status).to eq('FAILED_WITH_TERMINAL_ERROR') + expect(result.reason_for_incompletion).to include('GH_TOKEN') + end + end + + it 'passes unknown keys only to tools that accept **kwargs' do + strict = Conductor::Agents::ToolDef.new(name: 'strict', func: ->(a:) { { a: a } }, + input_schema: { 'type' => 'object', 'properties' => { 'a' => {} } }) + expect(described_class.run_tool_task(task_with({ 'a' => 1, 'zzz' => 2 }), strict).output_data).to eq('a' => 1) + loose = Conductor::Agents::ToolDef.new(name: 'loose', func: ->(a:, **rest) { { a: a, rest: rest } }, + input_schema: { 'type' => 'object', 'properties' => { 'a' => {} } }) + expect(described_class.run_tool_task(task_with({ 'a' => 1, 'zzz' => 2 }), loose).output_data).to eq('a' => 1, 'rest' => { zzz: 2 }) + end +end + +RSpec.describe Conductor::Agents::Secrets do + it 'secrets_env returns only the requested names' do + ENV['S_A'] = '1' + ENV['S_B'] = '2' + expect(described_class.secrets_env('S_A', 'S_B')).to eq('S_A' => '1', 'S_B' => '2') + ensure + ENV.delete('S_A') + ENV.delete('S_B') + end + + it 'raises CredentialNotFoundError with guidance' do + expect { described_class.secret('NOPE_NOT_SET') }.to raise_error(Conductor::Agents::CredentialNotFoundError, /conductor secrets put/) + end +end diff --git a/spec/conductor/agents/execution_spec.rb b/spec/conductor/agents/execution_spec.rb new file mode 100644 index 0000000..90aef12 --- /dev/null +++ b/spec/conductor/agents/execution_spec.rb @@ -0,0 +1,79 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::Execution do + let(:client) { instance_double(Conductor::Client::AgentClient) } + let(:execution) { described_class.new('EXEC_1', client: client, agent_name: 'weather') } + + it 'blocks in result until finished and exposes the answer' do + Thread.new do + sleep 0.05 + execution.finish(status: 'COMPLETED', output: { 'result' => 'Sunny', 'finishReason' => 'STOP' }) + end + expect(execution.done?).to be false + expect(execution.result(timeout: 2)).to eq('Sunny') + expect(execution.done?).to be true + expect(execution.finish_reason).to eq(:stop) + expect(execution.partial_text).to eq('Sunny') + end + + it 'times out when asked to' do + expect { execution.result(timeout: 0.05) }.to raise_error(Timeout::Error) + end + + it 'raises on failed executions and maps finish reasons' do + execution.finish(status: 'FAILED', output: {}, reason: 'LLM exploded') + expect { execution.result }.to raise_error(Conductor::Agents::Error, /LLM exploded/) + expect(execution.finish_reason).to eq(:error) + + rejected = described_class.new('E2', client: client) + rejected.finish(status: 'COMPLETED', output: { 'result' => nil, 'finishReason' => 'rejected', 'rejectionReason' => 'no' }) + expect(rejected.finish_reason).to eq(:rejected) + expect(rejected.rejected?).to be true + + expect(Conductor::Agents::FinishReason.derive('COMPLETED', 'finishReason' => 'MAX_TOKENS')).to eq(:length) + expect(Conductor::Agents::FinishReason.derive('TERMINATED', nil)).to eq(:cancelled) + expect(Conductor::Agents::FinishReason.derive('TIMED_OUT', nil)).to eq(:timeout) + end + + it 'pairs tool calls with their results' do + execution.add_tool_call('get_weather', 'city' => 'Lisbon') + execution.add_tool_result('get_weather', 'temp_c' => 21) + expect(execution.tool_calls.size).to eq(1) + expect(execution.tool_calls.first.arguments).to eq('city' => 'Lisbon') + expect(execution.tool_calls.first.result).to eq('temp_c' => 21) + expect(execution.tool_calls.first.to_s).to include('get_weather') + end + + it 'tracks waiting and delegates control calls to the client' do + request = instance_double(Conductor::Agents::ApprovalRequest, approve: true) + execution.mark_waiting(request) + expect(execution.waiting?).to be true + expect(execution.pending).to eq(request) + execution.approve + execution.clear_waiting + expect(execution.waiting?).to be false + expect { execution.reject }.to raise_error(Conductor::Agents::Error) + + expect(client).to receive(:pause).with('EXEC_1') + expect(client).to receive(:resume).with('EXEC_1') + expect(client).to receive(:cancel).with('EXEC_1', reason: 'bye') + expect(client).to receive(:stop).with('EXEC_1') + expect(client).to receive(:signal).with('EXEC_1', 'hurry') + execution.pause.resume.cancel(reason: 'bye').stop.signal('hurry') + end + + it 'loads a snapshot with .find' do + runtime = instance_double(Conductor::Agents::AgentRuntime, client: client) + allow(client).to receive(:get_status).with('EXEC_9').and_return( + 'executionId' => 'EXEC_9', 'agentName' => 'weather', 'status' => 'COMPLETED', 'isComplete' => true, + 'output' => { 'result' => 'done!', 'finishReason' => 'STOP' } + ) + found = described_class.find('EXEC_9', runtime: runtime) + expect(found.done?).to be true + expect(found.result).to eq('done!') + expect(found.agent_name).to eq('weather') + end +end diff --git a/spec/conductor/agents/sse_client_spec.rb b/spec/conductor/agents/sse_client_spec.rb new file mode 100644 index 0000000..8e14fb8 --- /dev/null +++ b/spec/conductor/agents/sse_client_spec.rb @@ -0,0 +1,89 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'webmock/rspec' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::SseClient do + let(:base) { 'http://localhost:8080/api' } + let(:configuration) { Conductor::Configuration.new(server_api_url: base) } + let(:api_client) { Conductor::Http::ApiClient.new(configuration: configuration) } + let(:client) { described_class.new(api_client, logger: Logger.new(nil)) } + let(:recorded) do + ":connected\n\n" \ + "id:1\nevent:thinking\ndata:{\"id\":1,\"type\":\"thinking\",\"executionId\":\"EXEC_1\",\"content\":\"weather_llm__1\",\"timestamp\":0}\n\n" \ + "id:2\nevent:tool_call\ndata:{\"id\":2,\"type\":\"tool_call\",\"executionId\":\"EXEC_1\",\"toolName\":\"get_weather\",\"args\":{\"city\":\"Lisbon\"},\"timestamp\":0}\n\n" \ + "id:3\nevent:done\ndata:{\"id\":3,\"type\":\"done\",\"executionId\":\"EXEC_1\",\"output\":{\"result\":\"Sunny\",\"finishReason\":\"STOP\"},\"timestamp\":0}\n\n" + end + + before { WebMock.enable! } + after { WebMock.reset! } + + describe described_class::Parser do + it 'parses frames split across chunks, comments, integer ids and multi-line data' do + parser = described_class.new + events = [] + parser.feed(":connected\n\nid:7\nev") { |e| events << e } + parser.feed("ent:message\ndata:{\"a\":\n") { |e| events << e } + parser.feed("data:1}\n\n") { |e| events << e } + expect(events).to eq([{ 'heartbeat' => true }, + { 'event' => 'message', 'id' => 7, 'data' => { 'a' => 1 } }]) + end + + it 'wraps non-JSON data as content and infers the event from data.type' do + parser = described_class.new + events = [] + parser.feed("data:plain text\n\ndata:{\"type\":\"done\"}\n\n") { |e| events << e } + expect(events[0]).to eq('event' => nil, 'id' => nil, 'data' => { 'content' => 'plain text' }) + expect(events[1]['event']).to eq('done') + end + end + + it 'streams the recorded events, drops heartbeats and stops after done' do + stub_request(:get, "#{base}/agent/stream/EXEC_1") + .with(headers: { 'Accept' => 'text/event-stream' }) + .to_return(status: 200, headers: { 'Content-Type' => 'text/event-stream' }, body: recorded) + + events = client.each_event('EXEC_1').to_a + expect(events.map { |e| e['event'] }).to eq(%w[thinking tool_call done]) + expect(events.map { |e| e['id'] }).to eq([1, 2, 3]) + expect(events.last['data']['output']['result']).to eq('Sunny') + end + + it 'raises SseUnavailableError when the first connection is refused or non-200' do + stub_request(:get, "#{base}/agent/stream/E500").to_return(status: 500) + expect { client.each_event('E500').to_a }.to raise_error(Conductor::Agents::SseUnavailableError, /500/) + + stub_request(:get, "#{base}/agent/stream/EDOWN").to_raise(Errno::ECONNREFUSED) + expect { client.each_event('EDOWN').to_a }.to raise_error(Conductor::Agents::SseUnavailableError) + end + + it 'reconnects with Last-Event-ID after the stream drops before done' do + stub_const('Conductor::Agents::SseClient::RECONNECT_DELAY', 0) + first = "id:1\nevent:thinking\ndata:{\"content\":\"x\"}\n\n" + rest = "id:2\nevent:done\ndata:{\"output\":{\"result\":\"ok\"}}\n\n" + stub_request(:get, "#{base}/agent/stream/EXEC_2").with { |req| req.headers['Last-Event-Id'].nil? } + .to_return(status: 200, body: first) + resumed = stub_request(:get, "#{base}/agent/stream/EXEC_2").with(headers: { 'Last-Event-ID' => '1' }) + .to_return(status: 200, body: rest) + + events = client.each_event('EXEC_2').to_a + expect(events.map { |e| e['event'] }).to eq(%w[thinking done]) + expect(resumed).to have_been_requested + end + + it 'raises SseUnavailableError when only heartbeats arrive' do + stub_const('Conductor::Agents::SseClient::HEARTBEAT_ONLY_TIMEOUT', -1) + stub_request(:get, "#{base}/agent/stream/EXEC_3").to_return(status: 200, body: ":heartbeat\n:heartbeat\n") + expect { client.each_event('EXEC_3').to_a }.to raise_error(Conductor::Agents::SseUnavailableError, /heartbeats/) + end + + it 'sends the auth header when authentication is configured' do + configuration.authentication_settings = Conductor::AuthenticationSettings.new(key_id: 'k', key_secret: 's') + configuration.update_token('jwt-token') + stub = stub_request(:get, "#{base}/agent/stream/EXEC_4").with(headers: { 'X-Authorization' => 'jwt-token' }) + .to_return(status: 200, body: "event:done\ndata:{}\n\n") + client.each_event('EXEC_4').to_a + expect(stub).to have_been_requested + end +end diff --git a/spec/conductor/agents/status_poller_spec.rb b/spec/conductor/agents/status_poller_spec.rb new file mode 100644 index 0000000..b7b0d6a --- /dev/null +++ b/spec/conductor/agents/status_poller_spec.rb @@ -0,0 +1,32 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::StatusPoller do + let(:client) { instance_double(Conductor::Client::AgentClient) } + let(:poller) { described_class.new(client, interval: 0) } + + it 'emits waiting once, then done' do + allow(client).to receive(:get_status).and_return( + { 'executionId' => 'E', 'status' => 'RUNNING', 'isComplete' => false, 'isWaiting' => false }, + { 'executionId' => 'E', 'status' => 'RUNNING', 'isComplete' => false, 'isWaiting' => true, 'pendingTool' => { 'x' => 1 } }, + { 'executionId' => 'E', 'status' => 'RUNNING', 'isComplete' => false, 'isWaiting' => true, 'pendingTool' => { 'x' => 1 } }, + { 'executionId' => 'E', 'status' => 'COMPLETED', 'isComplete' => true, 'output' => { 'result' => 'ok' } } + ) + events = poller.each_event('E').to_a + expect(events.map { |e| e['event'] }).to eq(%w[waiting done]) + expect(events[0]['data']['pendingTool']).to eq('x' => 1) + expect(events[1]['data']['output']).to eq('result' => 'ok') + end + + it 'emits error for non-completed terminal statuses' do + allow(client).to receive(:get_status).and_return( + 'executionId' => 'E', 'status' => 'FAILED', 'isComplete' => true, 'reasonForIncompletion' => 'bad' + ) + event = poller.each_event('E').first + expect(event['event']).to eq('error') + expect(event['data']['content']).to eq('bad') + expect(event['data']['status']).to eq('FAILED') + end +end diff --git a/spec/conductor/agents/system_workers_spec.rb b/spec/conductor/agents/system_workers_spec.rb new file mode 100644 index 0000000..324d946 --- /dev/null +++ b/spec/conductor/agents/system_workers_spec.rb @@ -0,0 +1,57 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::SystemWorkers do + a = Conductor::Agents + + def task(input) + Conductor::Http::Models::Task.from_hash('taskId' => 't', 'inputData' => input) + end + + it 'termination returns should_continue and reason' do + body = described_class.termination(a::Termination::TextMention.new('DONE')) + expect(body.call(task('result' => 'all DONE', 'iteration' => 1))).to eq('should_continue' => false, + 'reason' => "Text 'DONE' found in output") + expect(body.call(task('result' => 'working'))['should_continue']).to be true + end + + it 'termination keeps going when the condition raises' do + cond = a::Termination::TextMention.new('x') + allow(cond).to receive(:should_terminate).and_raise('boom') + expect(described_class.termination(cond, logger: Logger.new(nil)).call(task({}))).to eq('should_continue' => true, 'reason' => '') + end + + it 'guardrail passes, fails with retry, downgrades to raise when retries are exhausted or fix has no output' do + retry_g = a::Guardrail.new(name: 'g', on_fail: :retry, max_retries: 2) { |c| c.include?('ok') } + body = described_class.guardrail(retry_g) + expect(body.call(task('content' => 'ok'))).to eq(described_class.pass_result) + expect(body.call(task('content' => 'bad', 'iteration' => 1))).to include('passed' => false, 'on_fail' => 'retry', + 'guardrail_name' => 'g', 'should_continue' => true) + expect(body.call(task('content' => 'bad', 'iteration' => 2))).to include('on_fail' => 'raise', 'should_continue' => false) + + fix_g = a::Guardrail.new(name: 'f', on_fail: :fix) { |_c| a::GuardrailResult.new(passed: false, message: 'm') } + expect(described_class.guardrail(fix_g).call(task('content' => 'x'))).to include('on_fail' => 'raise', 'fixed_output' => nil) + fixer = a::Guardrail.new(name: 'f2', on_fail: :fix) { |_c| a::GuardrailResult.new(passed: false, fixed_output: 'clean') } + expect(described_class.guardrail(fixer).call(task('content' => { 'a' => 1 }))).to include('on_fail' => 'fix', 'fixed_output' => 'clean') + end + + it 'guardrail reports its own exceptions as failures' do + g = a::Guardrail.new(name: 'g', on_fail: :raise) { |_c| raise 'oops' } + expect(described_class.guardrail(g, logger: Logger.new(nil)).call(task('content' => 'x'))).to include('passed' => false, + 'message' => 'Guardrail error: oops') + end + + it 'callback passes messages / llm_result and returns the chain result' do + chain = ->(**kw) { { 'seen' => kw.keys.map(&:to_s) } } + expect(described_class.callback(chain).call(task('messages' => [], 'llm_result' => 'x'))).to eq('seen' => %w[messages llm_result]) + expect(described_class.callback(->(**) { raise 'x' }, logger: Logger.new(nil)).call(task({}))).to eq({}) + end + + it 'handoff evaluates the condition' do + h = a::Handoff::OnCondition.new(target: 'filer') { |ctx| ctx['result'].to_s.include?('go') } + expect(described_class.handoff(h).call(task('result' => 'go'))).to eq('handoff' => true, 'target' => 'filer') + expect(described_class.handoff(h).call(task('result' => 'stay'))).to eq('handoff' => false, 'target' => 'filer') + end +end diff --git a/spec/conductor/agents/tool_registry_spec.rb b/spec/conductor/agents/tool_registry_spec.rb new file mode 100644 index 0000000..bf8618c --- /dev/null +++ b/spec/conductor/agents/tool_registry_spec.rb @@ -0,0 +1,76 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'support/agent_tools' + +RSpec.describe Conductor::Agents::ToolRegistry do + a = Conductor::Agents + let(:config) { a::AgentConfig.new(worker_poll_interval_ms: 50, worker_thread_count: 2) } + let(:logger) { instance_double(Logger, warn: nil, info: nil, error: nil, debug: nil) } + let(:registry) { described_class.new(config, logger: logger) } + let(:weather) do + agent = a::Agent.new(name: 'weather', model: 'openai/gpt-4o-mini', instructions: 'Answer weather questions.') + agent.add_tool :current + agent.add_tool a::ToolDef.http('fetch', 'http://x') + agent + end + + it 'builds one worker per local tool with the Python task definition defaults' do + workers = registry.workers_for(weather, required_workers: ['current']) + expect(workers.map(&:task_definition_name)).to eq(['current']) + w = workers.first + expect(w.register_task_def).to be true + expect(w.overwrite_task_def).to be true + expect(w.lease_extend_enabled).to be true + expect(w.poll_interval).to eq(50) + expect(w.thread_count).to eq(2) + expect(w.domain).to be_nil + td = w.task_def_template.to_h + expect(td).to include('name' => 'current', 'retryCount' => 2, 'timeoutSeconds' => 0, 'timeoutPolicy' => 'RETRY', + 'retryLogic' => 'LINEAR_BACKOFF', 'retryDelaySeconds' => 2, 'responseTimeoutSeconds' => 10, + 'enforceSchema' => false, 'runtimeMetadata' => []) + end + + it 'puts tool and agent credentials on runtimeMetadata' do + filer = a::Agent.new(name: 'filer', model: 'm/x', credentials: ['ORG_KEY']) + filer.add_tool :create_issue + td = registry.workers_for(filer).first.task_def_template + expect(td.runtime_metadata).to eq(%w[GH_TOKEN ORG_KEY]) + end + + it 'runs the tool through Dispatch' do + worker = registry.workers_for(weather).first + task = Conductor::Http::Models::Task.from_hash('taskId' => 't', 'inputData' => { 'city' => 'Lisbon', 'method' => 'current' }) + result = worker.execute(task) + expect(result.status).to eq('COMPLETED') + expect(result.output_data['temp_c']).to eq(21.0) + end + + it 'registers system workers only when the server requires them and warns about unknown names' do + agent = a::Agent.new(name: 'bug_desk', model: 'm/x') + agent.stop_when 'ISSUE_FILED' + agent.add_guardrail a::Guardrail.new(name: 'no_pii') { true } + agent.callback(:before_model) { |**| nil } + agent.add_handoff a::Handoff::OnCondition.new(target: 'filer') { true } + + names = registry.workers_for(agent, required_workers: %w[bug_desk_termination no_pii bug_desk_before_model + bug_desk_handoff_filer mystery_task]).map(&:task_definition_name) + expect(names).to match_array(%w[bug_desk_termination no_pii bug_desk_before_model bug_desk_handoff_filer]) + expect(logger).to have_received(:warn).with(/mystery_task/) + + only_termination = registry.workers_for(agent, required_workers: ['bug_desk_termination']).map(&:task_definition_name) + expect(only_termination).to eq(['bug_desk_termination']) + expect(registry.workers_for(agent, required_workers: nil).size).to eq(4) + end + + it 'collects tools from the whole team once and applies the run domain for stateful trees' do + triage = a::Agent.new(name: 'triage', model: 'm/x') + filer = a::Agent.new(name: 'filer', model: 'm/x', stateful: true) + filer.add_tool :create_issue + triage.add_tool :create_issue + team = a::Agent.new(name: 'team', agents: [triage, filer]) + workers = registry.workers_for(team, domain: 'run-1') + expect(workers.map(&:task_definition_name)).to eq(['create_issue']) + expect(workers.first.domain).to eq('run-1') + end +end From 36994ffda0f973935a3eae0be98a7a22d550f566 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Tue, 8 Sep 2026 15:49:06 -0700 Subject: [PATCH 05/20] Phase 3: replay the recorded tool_happy_path scenario against WireMock - spec/agents/agents_helper.rb: `mocks: 'agent/tool_happy_path'` points the SDK at a WireMock replay of conductor-mocks (CONDUCTOR_AGENTS_REPLAY_URL, or a container started from CONDUCTOR_MOCKS_DIR), resets scenarios per example and fails when the SDK sent a request the recorded server never saw - spec/agents/replay/tool_happy_path_spec.rb: the one-pager weather agent end to end: start, TaskDef PUT, SSE stream, poll, in-process tool run, update-v2, done, token usage - CI: agents-replay job checks out conductor-oss/conductor-mocks, starts wiremock/wiremock:3x and runs spec/agents Co-Authored-By: Claude Fable 5.1 --- .github/workflows/ci.yml | 34 +++++ .rspec_agents_status | 3 + .rubocop.yml | 1 + spec/agents/agents_helper.rb | 144 +++++++++++++++++++++ spec/agents/replay/tool_happy_path_spec.rb | 28 ++++ spec/agents/support/weather_tools.rb | 14 ++ 6 files changed, 224 insertions(+) create mode 100644 .rspec_agents_status create mode 100644 spec/agents/agents_helper.rb create mode 100644 spec/agents/replay/tool_happy_path_spec.rb create mode 100644 spec/agents/support/weather_tools.rb diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ff68306..15b896c 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -111,3 +111,37 @@ jobs: run: | bundle exec rspec spec/integration/ --format documentation continue-on-error: true + + agents-replay: + name: Agents replay (WireMock) + runs-on: ubuntu-latest + steps: + - name: Checkout code + uses: actions/checkout@v4 + + - name: Checkout recorded scenarios + uses: actions/checkout@v4 + with: + repository: conductor-oss/conductor-mocks + path: conductor-mocks + + - name: Start WireMock with agent/tool_happy_path + run: | + docker run -d --name wiremock -p 8080:8080 \ + -v "$PWD/conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock" \ + wiremock/wiremock:3x + for i in $(seq 1 30); do + curl -fs http://localhost:8080/__admin/mappings > /dev/null && break + sleep 1 + done + + - name: Set up Ruby + uses: ruby/setup-ruby@v1 + with: + ruby-version: '3.2' + bundler-cache: true + + - name: Run agent replay specs + env: + CONDUCTOR_AGENTS_REPLAY_URL: http://localhost:8080 + run: bundle exec rspec spec/agents --format documentation diff --git a/.rspec_agents_status b/.rspec_agents_status new file mode 100644 index 0000000..6d4e3e3 --- /dev/null +++ b/.rspec_agents_status @@ -0,0 +1,3 @@ +example_id | status | run_time | +------------------------------------------------- | ------ | ------------ | +./spec/agents/replay/tool_happy_path_spec.rb[1:1] | passed | 3.21 seconds | diff --git a/.rubocop.yml b/.rubocop.yml index 2816546..b6f3904 100644 --- a/.rubocop.yml +++ b/.rubocop.yml @@ -204,4 +204,5 @@ RSpec/VerifiedDoubles: RSpec/DescribeSymbol: Exclude: - 'spec/support/**/*' + - 'spec/agents/support/**/*' - 'examples/**/*' diff --git a/spec/agents/agents_helper.rb b/spec/agents/agents_helper.rb new file mode 100644 index 0000000..9faecf4 --- /dev/null +++ b/spec/agents/agents_helper.rb @@ -0,0 +1,144 @@ +# frozen_string_literal: true + +# Helper for the agents runtime specs (spec/agents). These run the real SDK against a +# server: by default a WireMock replay of a scenario recorded from conductor-oss +# (github.com/conductor-oss/conductor-mocks), so no Conductor server and no LLM key is +# needed. +# +# Tag an example group with `mocks: 'agent/tool_happy_path'`. The helper points the SDK at +# the replay server and, after each example, asserts that WireMock saw no unmatched request +# (an unmatched request means the SDK sent something the real server never received). +# +# Environment: +# CONDUCTOR_AGENTS_REPLAY_URL base URL of a running WireMock serving the scenario +# (e.g. http://localhost:8080). When unset and Docker is +# available, one is started from CONDUCTOR_MOCKS_DIR. +# CONDUCTOR_MOCKS_DIR checkout of conductor-mocks (default ../conductor-mocks) +# +# Usage: +# CONDUCTOR_AGENTS_REPLAY_URL=http://localhost:8080 bundle exec rspec spec/agents +require 'bundler/setup' +require 'json' +require 'net/http' +require 'uri' +require 'conductor/agents' + +module AgentsReplay + DEFAULT_MOCKS_DIR = File.expand_path('../../../conductor-mocks', __dir__) + WIREMOCK_IMAGE = 'wiremock/wiremock:3x' + + class << self + attr_reader :base_url + + # Ensure a replay server for +scenario+ is reachable; returns false (with a reason) when it cannot be. + def ensure_server(scenario) + return [true, nil] if @base_url && @scenario == scenario + + if ENV['CONDUCTOR_AGENTS_REPLAY_URL'] + @base_url = ENV['CONDUCTOR_AGENTS_REPLAY_URL'].sub(%r{/+$}, '') + @scenario = scenario + return healthy? ? [true, nil] : [false, "WireMock at #{@base_url} is not answering"] + end + + start_container(scenario) + end + + def reset_scenarios + admin_post('/__admin/scenarios/reset') + admin_delete('/__admin/requests') + end + + # @return [Array] requests WireMock could not match + def unmatched_requests + body = admin_get('/__admin/requests/unmatched') + JSON.parse(body).fetch('requests', []) + rescue StandardError + [] + end + + def server_api_url + "#{@base_url}/api" + end + + def stop + return unless @container + + system('docker', 'rm', '-f', @container, out: File::NULL, err: File::NULL) + @container = nil + end + + private + + def start_container(scenario) + dir = File.join(ENV.fetch('CONDUCTOR_MOCKS_DIR', DEFAULT_MOCKS_DIR), 'mocks', scenario) + return [false, "scenario directory not found: #{dir}"] unless File.directory?(dir) + return [false, 'docker is not available and CONDUCTOR_AGENTS_REPLAY_URL is unset'] unless system('docker', 'version', out: File::NULL, err: File::NULL) + + stop + @container = "ruby-sdk-replay-#{Process.pid}" + ok = system('docker', 'run', '-d', '--name', @container, '-p', '8080:8080', '-v', "#{dir}:/home/wiremock:z", + WIREMOCK_IMAGE, out: File::NULL, err: File::NULL) + return [false, 'could not start the WireMock container'] unless ok + + @base_url = 'http://localhost:8080' + @scenario = scenario + 30.times do + return [true, nil] if healthy? + + sleep 1 + end + [false, 'WireMock container did not become healthy'] + end + + def healthy? + admin_get('/__admin/mappings') + true + rescue StandardError + false + end + + def admin_get(path) + response = Net::HTTP.get_response(URI.parse("#{@base_url}#{path}")) + raise "HTTP #{response.code}" unless response.code.to_i == 200 + + response.body + end + + def admin_post(path) + uri = URI.parse("#{@base_url}#{path}") + Net::HTTP.post(uri, '') + end + + def admin_delete(path) + uri = URI.parse("#{@base_url}#{path}") + Net::HTTP.start(uri.host, uri.port) { |http| http.delete(uri.path) } + end + end +end + +RSpec.configure do |config| + config.example_status_persistence_file_path = '.rspec_agents_status' + config.disable_monkey_patching! + config.expect_with(:rspec) { |c| c.syntax = :expect } + config.order = :defined + + config.before(:each, :mocks) do |example| + ok, reason = AgentsReplay.ensure_server(example.metadata[:mocks]) + skip "replay server unavailable: #{reason}" unless ok + + AgentsReplay.reset_scenarios + Conductor::Agents.configure( + configuration: Conductor::Configuration.new(server_api_url: AgentsReplay.server_api_url), + agent_config: Conductor::Agents::AgentConfig.new(worker_poll_interval_ms: 100), + logger: Logger.new(ENV['CONDUCTOR_AGENTS_DEBUG'] ? $stdout : nil) + ) + end + + config.after(:each, :mocks) do + Conductor::Agents.shutdown + unmatched = AgentsReplay.unmatched_requests + expect(unmatched.map { |r| "#{r['method']} #{r['url']}" }).to eq([]), 'the SDK sent requests the recorded server never saw' + end + + config.after(:suite) { AgentsReplay.stop } +end diff --git a/spec/agents/replay/tool_happy_path_spec.rb b/spec/agents/replay/tool_happy_path_spec.rb new file mode 100644 index 0000000..4ef894e --- /dev/null +++ b/spec/agents/replay/tool_happy_path_spec.rb @@ -0,0 +1,28 @@ +# frozen_string_literal: true + +require_relative '../agents_helper' +require_relative '../support/weather_tools' + +# Replays conductor-mocks/mocks/agent/tool_happy_path: the LLM calls get_weather("Lisbon"), +# the tool runs in this process, and the agent answers from the result. +RSpec.describe 'weather agent', mocks: 'agent/tool_happy_path' do + let(:agent) do + agent = Conductor::Agents::Agent.new(name: 'weather', model: 'openai/gpt-4o-mini', + instructions: 'Answer weather questions.') + agent.add_tool WeatherTools[:get_weather] + agent + end + + it 'answers with the tool' do + execution = agent.call_async('Weather in Lisbon?') + answer = execution.result(timeout: 60) + + expect(answer).to include('Lisbon') + expect(execution.finish_reason).to eq(:stop) + expect(execution.tool_calls.map(&:name)).to eq(['get_weather']) + expect(execution.tool_calls.first.arguments).to eq('city' => 'Lisbon', 'units' => 'metric') + expect(execution.tool_calls.first.result).to eq('temp_c' => 21.0, 'summary' => 'Sunny in Lisbon') + expect(execution.token_usage.total_tokens).to eq(246) + expect(Conductor::Agents.runtime.running_workers).to eq(['get_weather']) + end +end diff --git a/spec/agents/support/weather_tools.rb b/spec/agents/support/weather_tools.rb new file mode 100644 index 0000000..b850ffd --- /dev/null +++ b/spec/agents/support/weather_tools.rb @@ -0,0 +1,14 @@ +# frozen_string_literal: true + +require 'conductor/agents' + +# The weather tool from the one-pager. The description and result must match what was +# recorded in conductor-mocks agent/tool_happy_path. +module WeatherTools + extend Conductor::Agents::Tools + + tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } + end + describe :get_weather, 'Get the current weather for a city.' +end From 48ebeedb194a13cad01d5cfa4951408c8a19bdf6 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Tue, 8 Sep 2026 15:53:20 -0700 Subject: [PATCH 06/20] Phase 4: agents examples, user docs, README/AGENTS/DESIGN updates, changelog - examples/agents: weather.rb (tools), support_approval.rb (streaming + approval), bug_desk.rb (team + secret), matching the one-pager - docs/agents: README plus concepts for tools, calling/approval, teams, secrets, runtime and testing - README agents section and example table; AGENTS.md architecture, directory, test layout, docker test recipe and 'changing the agents wire format' checklist; DESIGN.md agents layer - CHANGELOG: agents feature, runtimeMetadata, update-v2, instance token cache, vcr removal - Plan: implementation status table and decisions 2.18-2.20 (update-v2, swarm hoisting for hands_off_to, run domain for every worker); notes in AGENT_TEAMS.md and the one-pager Co-Authored-By: Claude Fable 5.1 --- .rspec_agents_status | 2 +- AGENTS.md | 47 ++++++++- CHANGELOG.md | 20 ++++ DESIGN.md | 19 ++++ README.md | 23 +++++ docs/agents/README.md | 53 ++++++++++ docs/agents/concepts/runtime.md | 69 +++++++++++++ docs/agents/concepts/secrets.md | 72 ++++++++++++++ docs/agents/concepts/streaming-hitl.md | 93 +++++++++++++++++ docs/agents/concepts/teams.md | 92 +++++++++++++++++ docs/agents/concepts/tools.md | 115 ++++++++++++++++++++++ docs/design/AGENTS_IMPLEMENTATION_PLAN.md | 41 +++++++- docs/design/AGENTS_PARITY_ONEPAGER.md | 7 ++ docs/design/AGENT_TEAMS.md | 7 ++ examples/agents/bug_desk.rb | 50 ++++++++++ examples/agents/support_approval.rb | 51 ++++++++++ examples/agents/weather.rb | 27 +++++ 17 files changed, 785 insertions(+), 3 deletions(-) create mode 100644 docs/agents/README.md create mode 100644 docs/agents/concepts/runtime.md create mode 100644 docs/agents/concepts/secrets.md create mode 100644 docs/agents/concepts/streaming-hitl.md create mode 100644 docs/agents/concepts/teams.md create mode 100644 docs/agents/concepts/tools.md create mode 100644 examples/agents/bug_desk.rb create mode 100644 examples/agents/support_approval.rb create mode 100644 examples/agents/weather.rb diff --git a/.rspec_agents_status b/.rspec_agents_status index 6d4e3e3..c5da245 100644 --- a/.rspec_agents_status +++ b/.rspec_agents_status @@ -1,3 +1,3 @@ example_id | status | run_time | ------------------------------------------------- | ------ | ------------ | -./spec/agents/replay/tool_happy_path_spec.rb[1:1] | passed | 3.21 seconds | +./spec/agents/replay/tool_happy_path_spec.rb[1:1] | passed | 3.11 seconds | diff --git a/AGENTS.md b/AGENTS.md index 0ae6b75..da3ef46 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -10,6 +10,7 @@ This is the official Ruby SDK for [Conductor OSS](https://github.com/conductor-o - **Worker Framework** - Multi-threaded task execution with events and metrics - **Full API Coverage** - 17 Resource APIs, 9 high-level clients - **LLM/AI Tasks** - Chat completion, embeddings, image/audio generation +- **Agents** - `Conductor::Agents`: agent definitions serialized to the server's `agentConfig`, tool workers, SSE streaming (`lib/conductor/agents/`) ## Key Design Documents @@ -18,6 +19,8 @@ This is the official Ruby SDK for [Conductor OSS](https://github.com/conductor-o | [DESIGN.md](DESIGN.md) | High-level architecture and design principles | | [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) | Worker infrastructure design (polling, events, concurrency) | | [docs/design/WORKFLOW_DSL.md](docs/design/WORKFLOW_DSL.md) | Workflow DSL design and API reference | +| [docs/design/AGENTS_IMPLEMENTATION_PLAN.md](docs/design/AGENTS_IMPLEMENTATION_PLAN.md) | Agents: verified wire contract, decisions, work breakdown | +| [docs/agents/README.md](docs/agents/README.md) | Agents user guide (tools, streaming/approval, teams, secrets, runtime) | | [README.md](README.md) | User-facing documentation with examples | | [CONTRIBUTING.md](CONTRIBUTING.md) | Development workflow and guidelines | @@ -424,6 +427,27 @@ lib/conductor/ ├── task_type.rb # Task type constants ├── timeout_policy.rb └── workflow_executor.rb +lib/conductor/agents.rb # require 'conductor/agents' entry point + default runtime +lib/conductor/agents/ +├── agent.rb # Agent definition + sugar (add_tool, hands_off_to, on_approval, >>) +├── tool_def.rb # ToolDef, ToolType, server-side tool factories +├── tools.rb # `tool def` DSL, describe, requires_approval, registries +├── tools/schema_builder.rb # keyword defaults -> JSON schema (AST) +├── tools/secret_scanner.rb # secret('X') literals -> credentials +├── tools/ruby_llm_adapter.rb # RubyLLM::Tool -> ToolDef +├── guardrail.rb, termination.rb, handoff.rb, callback_handler.rb, memory.rb, prompt_template.rb +├── config_serializer.rb # Agent tree -> agentConfig (Python-identical) +└── runtime/ + ├── agent_runtime.rb # call_sync / call_async / deploy / serve + ├── sse_client.rb # GET /agent/stream SSE with reconnect + ├── status_poller.rb # polling fallback + ├── execution.rb # Execution, ToolCall, TokenUsage, FinishReason + ├── approval_request.rb # waiting -> approve / reject + ├── tool_registry.rb # ToolDef -> Worker with Python TaskDef defaults + ├── dispatch.rb # task input -> kwargs -> result + ├── system_workers.rb # termination / guardrail / callback / handoff bodies + ├── secrets.rb # secret(), secrets_env() + └── agent_config.rb # CONDUCTOR_AGENT_* settings ``` --- @@ -530,7 +554,11 @@ spec/ │ │ └── llm_tasks_spec.rb # LLM helper tests │ ├── client/ │ ├── http/ -│ └── worker/ +│ ├── worker/ +│ └── agents/ # Agents unit + contract tests (no server) +│ └── contract_spec.rb # 19 golden configs must equal python-sdk + validate against agent-schema.json +├── fixtures/agents/ # vendored agent-schema.json and golden configs +├── agents/ # Replay tests against WireMock (conductor-mocks recordings) └── integration/ # Requires live server ``` @@ -551,6 +579,16 @@ bundle exec rspec --format documentation # Integration tests (requires Conductor server) CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ + +# Agents replay tests (WireMock serving conductor-mocks/mocks/agent/tool_happy_path on :8080) +CONDUCTOR_AGENTS_REPLAY_URL=http://localhost:8080 bundle exec rspec spec/agents +``` + +No Ruby installed? The suite runs in the official image: + +```bash +docker run --rm -v "$PWD":/app:z -w /app -v ruby-sdk-bundle:/usr/local/bundle:z ruby:3.3 \ + bash -lc "bundle install --quiet && bundle exec rspec spec/conductor/ && bundle exec rubocop" ``` --- @@ -584,6 +622,13 @@ CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integratio 5. Run linter: `bundle exec rubocop -a` 6. Run tests: `bundle exec rspec spec/conductor/` +### Changing the agents wire format + +1. The Python SDK (`python-sdk/src/conductor/ai/agents/config_serializer.py`) is the parity source; the server (`conductor/agentspan`) is the contract +2. Edit `lib/conductor/agents/config_serializer.rb`, then `bundle exec rspec spec/conductor/agents/contract_spec.rb` +3. New wire fields need a golden fixture: add the agent to `examples/agents/golden_agents.rb`, generate the Python side with `python-sdk/examples/agents/dump_agent_configs.py`, vendor it into `spec/fixtures/agents/configs/` +4. Runtime behaviour is verified by replay (`spec/agents`); record new scenarios in `conductor-oss/conductor-mocks` + ### Adding Event Listeners / Interceptors 1. Create class implementing listener methods (`on_poll_started`, `on_task_execution_completed`, etc.) diff --git a/CHANGELOG.md b/CHANGELOG.md index 4595cac..2ce4ec8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -59,6 +59,26 @@ end ### Added +- **Agents** (`require 'conductor/agents'`) - Ruby port of the Python SDK's agents package, same `agentConfig` on the wire -- [guide](docs/agents/README.md) + - `tool def` DSL: types from keyword defaults, secrets from `secret('...')` literals, `describe`, `requires_approval`, module scoping, RubyLLM::Tool adapter + - `Agent` with `add_tool`, `add_agent`, `hands_off_to`, `redact`, `stop_when`, `stop_after`, `on_approval`, `>>`; guardrails, termination conditions, handoffs, callbacks, memory, prompt templates + - `ConfigSerializer` verified against the Python SDK's 19 golden configs and `agent-schema.json` + - `AgentRuntime`: `call_sync`, `call_async` (SSE streaming with reconnect, polling fallback), `deploy`, `serve`; `Execution`, `ApprovalRequest` + - Tool workers registered with Python's task definition defaults; `_termination`, custom guardrail and callback workers + - `AgentResourceApi` / `AgentClient` for `/api/agent/*`, `OrkesClients#get_agent_client` + - `Task#runtime_metadata` (wire-only secret values), `TaskDef#runtime_metadata` (declared secret names), `TaskDef#enforce_schema` + - `TaskResourceApi#update_task_v2`; `Worker` option `lease_extend_enabled` + - Replay tests against `conductor-oss/conductor-mocks` recordings (`spec/agents`, CI job `agents-replay`) + +### Changed + +- `Configuration` caches the auth token per instance (two configurations no longer share a token); the class-level `Configuration.auth_token` accessors remain as a deprecated shim +- `TaskRunner` posts task results to `POST /tasks/update-v2` and falls back to `POST /tasks` once when the server does not serve it (Python SDK parity) + +### Removed + +- Unused `vcr` development dependency (`json_schemer` added for agent contract tests) + - **Core Infrastructure** - Configuration with environment variable support - Authentication (token management, TTL refresh, exponential backoff) diff --git a/DESIGN.md b/DESIGN.md index c17130b..86c4ca1 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -242,8 +242,27 @@ lib/conductor/ ├── task_type.rb # Task type constants ├── timeout_policy.rb └── workflow_executor.rb +lib/conductor/agents.rb # require 'conductor/agents' +lib/conductor/agents/ # Agents (see docs/agents/README.md) +├── agent.rb, tool_def.rb, tools.rb # definition layer + `tool def` DSL +├── guardrail.rb, termination.rb, handoff.rb, callback_handler.rb, memory.rb, prompt_template.rb +├── config_serializer.rb # -> agentConfig, identical to the Python SDK +└── runtime/ # AgentRuntime, SseClient, Execution, ApprovalRequest, ToolRegistry, Dispatch, Secrets ``` +## Agents + +`Conductor::Agents` ports the Python SDK's `conductor.ai.agents` package. An `Agent` tree is +serialized by `ConfigSerializer` to the same `agentConfig` JSON Python sends; the server compiles +it into a workflow and runs the LLM loop. `AgentRuntime` starts the execution +(`POST /api/agent/start`), registers a Conductor worker for every task the server lists in +`requiredWorkers` (the user's `tool def` tools plus `_termination`, custom guardrails and +callbacks), and follows the run over SSE (`GET /api/agent/stream/{id}`, polling fallback). +Secrets travel as `TaskDef.runtimeMetadata` names and come back as `Task.runtimeMetadata` +values, read inside tools with `secret('NAME')`. The transport for `/api/agent/*` is +`AgentResourceApi` / `AgentClient` like every other resource. Decisions and the verified wire +contract are in `docs/design/AGENTS_IMPLEMENTATION_PLAN.md`. + ## Dependencies ```ruby diff --git a/README.md b/README.md index 8f6c97d..dfd0348 100644 --- a/README.md +++ b/README.md @@ -11,6 +11,7 @@ Official Ruby SDK for [Conductor OSS](https://github.com/conductor-oss/conductor - **Ruby-Idiomatic Workflow DSL** - Clean block-based syntax with 25+ task types - **Worker Framework** - Multi-threaded task execution with class-based and block-based workers - **LLM/AI Tasks** - Chat completion, embeddings, RAG, image/audio generation +- **Agents** - Define agents with Ruby methods as tools, run them on the server, stream the answer (`require 'conductor/agents'`) - **Orkes Cloud Support** - Authentication, secrets, integrations, prompts - **Comprehensive Testing** - 400+ unit tests, 110 integration tests @@ -324,6 +325,27 @@ workflow = Conductor.workflow :ai_assistant, executor: executor do end ``` +### Agents + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } +end + +agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', instructions: 'Answer weather questions.') +agent.add_tool :get_weather + +puts agent.call_sync('Weather in Lisbon?') +``` + +`tool def` turns a method into a tool (types from the keyword defaults, secrets from +`secret('...')` literals). The server runs the LLM loop; this process runs the tools. Teams, +approval (`requires_approval` + `on_approval`), streaming (`call_async`), guardrails and +termination are covered in [docs/agents/](docs/agents/README.md). + ### Output References The DSL uses a clean syntax for referencing outputs: @@ -355,6 +377,7 @@ The `examples/` directory contains comprehensive examples: | [`dynamic_workflow.rb`](examples/dynamic_workflow.rb) | Create and execute workflows at runtime | | [`workflow_ops.rb`](examples/workflow_ops.rb) | Lifecycle operations: pause, resume, restart, retry | | [`agentic_workflows/`](examples/agentic_workflows/) | LLM chat and AI workflow examples | +| [`agents/`](examples/agents/) | Agents: tools (`weather.rb`), approval + streaming (`support_approval.rb`), team + secret (`bug_desk.rb`) | Run examples: diff --git a/docs/agents/README.md b/docs/agents/README.md new file mode 100644 index 0000000..bbd7187 --- /dev/null +++ b/docs/agents/README.md @@ -0,0 +1,53 @@ +# Agents + +Define an agent in Ruby, run it on a Conductor server. Same `agentConfig` on the wire as the +Python SDK; the server compiles and runs the loop, this process runs your tools. + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } +end + +agent = Agent.new( + name: 'weather', + model: 'openai/gpt-4o', + instructions: 'Answer weather questions.' +) +agent.add_tool :get_weather + +puts agent.call_sync('Weather in Lisbon?') +``` + +Run it: `CONDUCTOR_SERVER_URL=http://localhost:8080/api ruby weather.rb`. Server URL and auth +come from the usual `CONDUCTOR_*` variables; nothing else to configure. + +| Guide | What it covers | +|---|---| +| [Tools](concepts/tools.md) | `tool def`, types from keyword defaults, `describe`, `requires_approval`, modules, RubyLLM tools, server-side tools | +| [Calling an agent](concepts/streaming-hitl.md) | `call_sync`, `call_async`, `Execution`, approval with `on_approval` | +| [Teams](concepts/teams.md) | `add_agent`, `hands_off_to`, strategies, `redact`, `stop_when`, `stop_after` | +| [Secrets](concepts/secrets.md) | `secret()`, `secrets_env()`, how names reach the server and values reach the tool | +| [Runtime and deployment](concepts/runtime.md) | `AgentRuntime`, `deploy`, `serve`, `CONDUCTOR_AGENT_*` settings, what runs where | + +Examples: [`examples/agents/`](../../examples/agents/) (`weather.rb`, `support_approval.rb`, +`bug_desk.rb`). The 19 agents in `examples/agents/golden_agents.rb` serialize identically to +the Python SDK's `examples/agents/_configs`; `dump_agent_configs.rb` regenerates them. + +## Requirements + +- A Conductor server with the agent runtime enabled (`conductor.integrations.ai.enabled=true`, + the default) and an LLM integration configured. On Orkes the left side of + `model: 'openai/gpt-4o'` is the integration name; on OSS it is the provider key whose API key + the server reads from `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` / ... +- Secrets delivered to tools (`secret('GH_TOKEN')`) need conductor-oss 3.32.0-rc.8 or later + (`runtimeMetadata`). Older servers still work: `secret()` falls back to `ENV`. + +## What is not ported + +Framework agents (OpenAI Agents SDK, LangGraph, Google ADK, Claude Agent SDK), skills, +`plan_execute`, local code execution and CLI tools, schedules, semantic memory. See +`docs/design/AGENTS_IMPLEMENTATION_PLAN.md` for the full list and the decisions behind the +port. diff --git a/docs/agents/concepts/runtime.md b/docs/agents/concepts/runtime.md new file mode 100644 index 0000000..790ff2d --- /dev/null +++ b/docs/agents/concepts/runtime.md @@ -0,0 +1,69 @@ +# Runtime and deployment + +`Agent#call_sync` / `#call_async` use `Conductor::Agents.runtime`, an `AgentRuntime` built +from the environment (`CONDUCTOR_SERVER_URL`, `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET`). + +```ruby +Conductor::Agents.configure( + configuration: Conductor::Configuration.new(server_api_url: 'https://play.orkes.io/api'), + agent_config: Conductor::Agents::AgentConfig.new(worker_thread_count: 4) +) + +runtime = Conductor::Agents::AgentRuntime.new(configuration: config) # or your own instance +runtime.call_sync(agent, 'hi') +runtime.compile(agent) # { "workflowDef", "requiredWorkers" } without registering +runtime.deploy(a, b) # register on the server; returns names +runtime.serve(a, b) # deploy + run the tool workers until Ctrl-C +runtime.shutdown # stop workers and streams +Conductor::Agents.shutdown # same for the default runtime +``` + +## Settings + +| Variable | Default | Meaning | +|---|---|---| +| `CONDUCTOR_AGENT_WORKER_POLL_INTERVAL` | `100` | tool worker poll interval (ms) | +| `CONDUCTOR_AGENT_WORKER_THREADS` | `1` | concurrent tasks per tool worker | +| `CONDUCTOR_AGENT_STREAMING_ENABLED` | `true` | use SSE; `false` polls `/agent/{id}/status` | +| `CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER` | `false` | reserved (Python parity), not used yet | + +Booleans accept `true/1/yes/on` and `false/0/no/off`. + +## What runs where + +| Piece | Where | +|---|---| +| LLM loop, tool routing, guardrail chain, approval pause, MCP discovery | server | +| `tool def` tools, RubyLLM tools | this process (Conductor workers, one per tool name) | +| `_termination` (`stop_when`, `stop_after`, `termination:`) | this process | +| custom guardrails (with a block), callbacks, `on_condition` handoffs | this process | +| regex / LLM guardrails, http / mcp / human / agent tools | server | + +Every worker registers its TaskDef with the Python SDK's defaults: `retryCount 2`, +`retryDelaySeconds 2`, `retryLogic LINEAR_BACKOFF`, `timeoutSeconds 0`, +`responseTimeoutSeconds 10`, `timeoutPolicy RETRY`, `runtimeMetadata` = declared secret names. +Task results go to `POST /api/tasks/update-v2` (with a one-time fallback to `POST /api/tasks`). + +An agent execution is a Conductor workflow: `execution_id` is the workflow id, and the +workflow, task and prompt data are visible in the Conductor UI like any other run. + +## Stateful runs + +`Agent.new(..., stateful: true)` (or a stateful tool) sends a `runId` with the start request. +The server maps every required worker to that task domain and the SDK polls with it, so each +run's tasks reach the process that started it. + +## Testing + +Contract tests (`spec/conductor/agents/contract_spec.rb`) need no server: every serialized +config must equal the Python SDK's golden file and validate against `agent-schema.json`. + +Runtime tests (`spec/agents/`) replay scenarios recorded from a real server with WireMock +(`conductor-oss/conductor-mocks`): + +```bash +docker run -d -p 8080:8080 -v $PWD/../conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock wiremock/wiremock:3x +CONDUCTOR_AGENTS_REPLAY_URL=http://localhost:8080 bundle exec rspec spec/agents +``` + +A request the recorded server never saw fails the run. diff --git a/docs/agents/concepts/secrets.md b/docs/agents/concepts/secrets.md new file mode 100644 index 0000000..d2c00b7 --- /dev/null +++ b/docs/agents/concepts/secrets.md @@ -0,0 +1,72 @@ +# Secrets + +Three kinds. Three different owners. None of them live in your code. + +| Secret | Held by | You write | +|---|---|---| +| Conductor auth | your env | `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET` (existing `Configuration`; the token is cached per instance) | +| LLM provider key | Conductor server, as an Integration | `model: 'openai/gpt-4o'` | +| Tool credential | Conductor server, in the secret store | `secret('GH_TOKEN')` in the tool body | + +## Tool credentials + +```ruby +tool def create_issue(title: String, body: '') + Github.create_issue(title, body, token: secret('GH_TOKEN')) +end +``` + +`secret('GH_TOKEN')` is both the read and the declaration. At `tool def` the SDK scans the +method body for literal `secret('...')` and `secrets_env('...')` names and puts them on the +tool's contract: `TaskDef.runtimeMetadata` (names) when the worker registers, and +`tool.config.credentials` in the `agentConfig`. The server resolves the values from its secret +store when your worker polls and attaches them to the task (`Task.runtimeMetadata`, wire-only, +never persisted). Inside the tool, `secret()` reads that map through the task context. + +Store the value once on the server: Orkes UI or `SecretClient#put_secret`; on OSS, +`CONDUCTOR_SECRET_GH_TOKEN` in the server's environment. + +```mermaid +sequenceDiagram + participant SDK as Ruby SDK + participant Server as Conductor Server + participant Store as Secret Store + + Note over SDK: at startup + SDK->>Server: register TaskDef create_issue, runtimeMetadata: [GH_TOKEN] + + Note over Server: LLM calls create_issue + Server->>Store: get GH_TOKEN + Store-->>Server: ghp_... + Server->>+SDK: task create_issue, runtimeMetadata: GH_TOKEN=ghp_... + SDK->>SDK: create_issue runs, secret('GH_TOKEN') -> ghp_... + SDK-->>-Server: result (runtimeMetadata dropped) +``` + +## When the name is not a literal + +```ruby +filer.add_tool :create_issue, credentials: ['GH_TOKEN'] # this tool +tool_credentials :create_issue, 'GH_TOKEN' # same, at definition time +filer = Agent.new(..., credentials: ['GH_TOKEN']) # everything under this agent +``` + +## Subprocesses + +Workers are threads in one process, so the SDK never writes `ENV` (that would leak the secret +into every other tool running at the same time). For `system` / `spawn` / `Open3`: + +```ruby +tool def gh_create_issue(title: String) + system(secrets_env('GH_TOKEN'), 'gh', 'issue', 'create', '--title', title) +end +``` + +`secrets_env('GH_TOKEN')` is `{ 'GH_TOKEN' => 'ghp_...' }` for this call and declares the same +way. + +## No secret store on the server? + +`secret('X')` falls back to `ENV['X']`, then raises `CredentialNotFoundError` (the task fails +terminally with a message naming the key). Servers older than conductor-oss 3.32.0-rc.8 do +not send `runtimeMetadata`; the ENV fallback keeps tools working there. diff --git a/docs/agents/concepts/streaming-hitl.md b/docs/agents/concepts/streaming-hitl.md new file mode 100644 index 0000000..8f33895 --- /dev/null +++ b/docs/agents/concepts/streaming-hitl.md @@ -0,0 +1,93 @@ +# Calling an agent: sync, async, approval + +## Sync + +```ruby +answer = agent.call_sync('What is your return policy?') +``` + +Blocks until the agent is finished and returns the answer as a String. Pass `timeout:` seconds +to give up (`Timeout::Error`). A failed, cancelled or timed-out execution raises +`Conductor::Agents::Error`. + +## Async with a callback + +```ruby +agent.call_async('What is your return policy?') do |answer, execution| + Mailer.send(customer, answer) +end +``` + +Returns immediately. The block runs on the stream thread when the agent finishes (`answer` is +nil when it failed; check `execution.finish_reason`). Exceptions in the block are logged, never +raised into the stream. + +## Async, poll it yourself + +```ruby +execution = agent.call_async('What is your return policy?') + +execution.done? # false until finished +execution.waiting? # true while a tool waits for approval +execution.partial_text # text streamed so far +execution.result # blocks until done, returns the answer +execution.finish_reason # :stop | :tool_calls | :length | :content_filter | :rejected | :error | :cancelled | :timeout +execution.tool_calls # [#] with .result once known +execution.token_usage # prompt / completion / total, summed over sub-agents +execution.execution_id # also the Conductor workflow id; Execution.find(id) later +execution.pause; execution.resume; execution.cancel; execution.stop; execution.signal('hurry up') +``` + +## Approval + +```ruby +tool def issue_refund(order_id: String, amount: Float) + Billing.refund(order_id, amount) +end +requires_approval :issue_refund + +agent.on_approval do |request| + request.amount < 100 ? request.approve : request.reject('Needs a manager') +end +``` + +When the model calls an approval-required tool the server pauses on a HUMAN task and the SDK +receives a `waiting` event. `request` is an `ApprovalRequest`: the tool's arguments are +methods (`request.amount`, `request.order_id`), plus `tool_name`, `arguments`, `tool_calls` +(the server gates the whole batch of tool calls in a turn with one approval), `approve`, +`reject(reason)` and `send_message(text)` for human-input tools. + +The block runs on the stream thread, for `call_sync` and `call_async` alike. Without an +`on_approval` block the request is parked on `execution.pending`; `execution.approve` / +`execution.reject` answer it, or any other client can call `POST /api/agent/{id}/respond`. + +A rejected tool ends the run as COMPLETED with `finish_reason == :rejected`, `result` nil and +the reason on `execution.output['rejectionReason']`. + +## Sessions + +```ruby +agent.call_sync(question, session_id: 'cust-77') # same conversation across calls +``` + +## Fire and forget + +`call_async` with no block, keep the `execution_id`, walk away. If the agent has `tool def` +tools they run in *your* process, so it has to stay up. Fire-and-forget only works when every +tool is server-side (http / mcp / human) or you have `deploy`ed the agent and run `serve` +somewhere else. + +## Under the hood + +`call_async(prompt)`: + +1. `POST /api/agent/start` with the serialized `agentConfig`; the reply lists `requiredWorkers` +2. start a worker for every required task this process can serve (your tools, plus + `_termination`, custom guardrails, callbacks) with the same TaskDef defaults as Python +3. return an `Execution`; open `GET /api/agent/stream/{id}` (SSE) on a background thread, + reconnecting with `Last-Event-ID`; fall back to polling `GET /api/agent/{id}/status` if SSE is + unavailable or `CONDUCTOR_AGENT_STREAMING_ENABLED=false` +4. `tool_call` / `tool_result` events fill `tool_calls`; `waiting` builds an `ApprovalRequest` +5. `done` sets the result and `finish_reason`, fetches token usage, fires the block + +`call_sync` is `call_async(prompt).result`. diff --git a/docs/agents/concepts/teams.md b/docs/agents/concepts/teams.md new file mode 100644 index 0000000..5ef6465 --- /dev/null +++ b/docs/agents/concepts/teams.md @@ -0,0 +1,92 @@ +# Teams + +```ruby +tool def create_issue(title: String, body: '') + Github.create_issue(title, body, token: secret('GH_TOKEN')) +end + +triage = Agent.new(name: 'triage', model: 'openai/gpt-4o-mini', + instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.') + +filer = Agent.new(name: 'filer', model: 'anthropic/claude-sonnet-4-5', + instructions: 'File the bug as a GitHub issue.') +filer.add_tool :create_issue + +triage.hands_off_to filer, on: 'ACTIONABLE' + +team = Agent.new(name: 'bug_desk') +team.add_agent triage +team.add_agent filer + +puts team.call_sync(File.read('report.md')) +``` + +| Line | Does | +|---|---| +| `filer.add_tool :create_issue` | Give filer a tool. | +| `triage.hands_off_to filer, on: 'ACTIONABLE'` | When triage's answer contains ACTIONABLE, filer takes over. Pass a block/proc as `on:` for a custom condition. | +| `team.add_agent triage` | Give the team a member. First one added starts. `team.add_agents a, b` for several. | +| `Agent.new(name: 'bug_desk')` | No model needed: a team inherits the first member's model (the server requires one on every agent). | + +`tools:` and `agents:` still work as keyword arguments in `Agent.new`. + +## Strategies + +```ruby +team.strategy = :sequential # one after another +team.strategy = :parallel # all at once, merged +team.strategy = :swarm # members transfer to each other via handoffs +team.strategy = :router # router: agent picks the member +# also :round_robin, :random, :manual (default :handoff) + +pipeline = researcher >> writer >> editor # sequential, named researcher_writer_editor +``` + +When members declare `hands_off_to` and the team has no explicit strategy, the serializer +makes the team a `swarm` and lists the members' handoffs on the team, which is where the +server reads them. Set a strategy explicitly to keep it. + +## Guardrails and stopping + +```ruby +filer.redact %w[password api_key] # scrub these from output before anyone sees it +filer.stop_when 'ISSUE_FILED' # stop on this text +filer.stop_after messages: 12 # or after this many messages +filer.add_guardrail RegexGuardrail.new('\b\d{3}-\d{2}-\d{4}\b', name: 'no_ssn', on_fail: :raise) +filer.add_guardrail LlmGuardrail.new('openai/gpt-4o-mini', 'No medical advice', on_fail: :retry) +filer.add_guardrail Guardrail.new(name: 'no_pii', on_fail: :retry) { |text| !text.include?('SSN') } +``` + +`stop_when` / `stop_after` build `Termination::TextMention` / `Termination::MaxMessage` and +combine with `|`; any `Termination::*` condition (also `StopMessage`, `TokenUsage`, `&`, `|`) +can be set directly with `termination:`. The server evaluates them through a +`_termination` worker this process runs. Custom guardrails with a block also run here; +regex and LLM guardrails run on the server. + +## Callbacks + +```ruby +class Timing < Conductor::Agents::CallbackHandler + def on_model_start(messages: nil, **) = (@t0 = Time.now; nil) + def on_model_end(llm_result: nil, **) = (puts Time.now - @t0; nil) +end +agent.add_callback Timing.new +agent.callback(:before_tool) { |**kw| log kw; nil } +``` + +Positions: `before_agent after_agent before_model after_model before_tool after_tool`. Each +becomes a `_` task the server schedules and this process serves. Return a +non-empty Hash to override; nil to continue. + +## Underneath + +Same Python objects, same `agentConfig`. Sugar only. + +| Sugar | Python-parity object | +|---|---| +| `a.add_tool :x` | appends to `Agent#tools`, same array `tools:` fills | +| `team.add_agent a` | appends to `Agent#agents`, same as `agents:` | +| `a.hands_off_to b, on: 'X'` | `Handoff::OnTextMention.new(target: 'b', text: 'X')` | +| `a.redact %w[...]` | `RegexGuardrail.new(..., position: :output, on_fail: :fix)` | +| `a.stop_when 'X'` / `a.stop_after messages: n` | `Termination::TextMention \| Termination::MaxMessage` | +| `a >> b` | `Agent.new(strategy: :sequential, agents: [a, b])` | diff --git a/docs/agents/concepts/tools.md b/docs/agents/concepts/tools.md new file mode 100644 index 0000000..5d3e3d0 --- /dev/null +++ b/docs/agents/concepts/tools.md @@ -0,0 +1,115 @@ +# Tools + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } +end + +agent.add_tool :get_weather +``` + +`tool` receives the Symbol that `def` returns (the same trick as `private def`) and builds a +`ToolDef`: the tool's name is the method name, its description is the humanized name +(`get_weather` -> "Get weather"), its JSON schema comes from the keyword defaults, and its +secrets from `secret('...')` literals in the body. The method stays a normal method, so +`get_weather(city: 'Lisbon')` works in tests. + +## Types + +The default is the type: + +| You write | JSON schema | +|---|---| +| `city: String` | required string | +| `amount: Float` | required number | +| `count: Integer` | required integer | +| `flags: Hash` | required object | +| `tags: [String]` | required array of strings | +| `units: 'metric'` | optional string, default `"metric"` | +| `limit: 10` | optional integer, default `10` | +| `verbose: false` | optional boolean, default `false` | +| `units: %w[metric imperial]` | optional string, one of these, default `"metric"` | +| `extra: nil` | optional, any type | +| `city:` (no default) | required, any type | + +Positional parameters raise at `tool def`. Types are read from the method's AST +(`RubyVM::AbstractSyntaxTree.of`), which works on MRI whenever the source file is on disk. For +methods typed into irb (Ruby 3.2+: set `RubyVM.keep_script_lines = true`) or on other Rubies, +every keyword becomes an untyped property and only keywords without defaults are required. + +Because `city: String` uses the class as the Ruby default, the runtime never calls a tool with a +required argument missing: the task fails with a clear reason instead. + +## Description, approval, options + +```ruby +describe :get_weather, 'Get the current weather for a city.' +requires_approval :issue_refund # a human approves before it runs +tool_credentials :create_issue, 'GH_TOKEN' # when the secret name is not a literal +Weather[:forecast].timeout_seconds = 60 # any ToolDef attribute +tool :ping, description: 'Health check', retry_count: 0 # options on registration +``` + +## Modules + +```ruby +module Weather + extend Conductor::Agents::Tools + + tool def current(city: String) ... end + tool def forecast(city: String, days: 3) ... end +end + +agent.add_tools Weather # both +agent.add_tool Weather[:current] # one +``` + +Inside a class body (an `RSpec.describe` block, for example) `tool def` also works; the method +is bound to a bare instance of the class. + +## RubyLLM tools + +```ruby +class Weather < RubyLLM::Tool + description 'Gets current weather for a location' + param :latitude, type: :number + param :longitude, type: :number + + def execute(latitude:, longitude:) ... end +end + +agent.add_tool Weather # a RubyLLM::Tool class, as-is +``` + +The adapter reads `name`, `description` and `parameters`, and runs `Weather.new.execute(**args)` +as the worker body. RubyLLM is optional and only used when it is loaded. It is not the engine: +Conductor runs the LLM loop server-side with the provider key held as an integration. + +## Server-side tools + +These need no worker in your process: + +```ruby +ToolDef.http('lookup', 'https://api.example.com/orders/${id}', method: 'GET', + headers: { 'Authorization' => 'Bearer ${API_KEY}' }, credentials: ['API_KEY']) +ToolDef.mcp('http://localhost:3001/mcp', tool_names: %w[search fetch]) +ToolDef.human('ask_user', description: 'Ask the user a question and wait for the answer.') +ToolDef.agent(researcher) # another agent as a tool +ToolDef.image('draw', description: '...', llm_provider: 'openai', model: 'dall-e-3') +``` + +`${NAME}` placeholders in headers must be listed in `credentials:`. MCP discovery happens on +the server (`LIST_MCP_TOOLS` before the loop); nothing runs client-side. + +## What the server sends your tool + +The LLM's arguments arrive as top-level task input keys, plus `method`, `_agent_state` and +`_agent_tool_name` which the SDK strips. Strings are coerced to the schema type (`"5"` -> `5`, +`"true"` -> `true`, JSON text -> arrays/objects). A Hash result is returned as-is; anything else +is wrapped as `{ "result" => value }`. A `_state_updates` key in the result is merged into the +agent's durable state by the server. Exceptions fail the task (retried per the tool's +`retry_count`, default 2); missing arguments, missing secrets and unserializable results fail it +terminally. diff --git a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md index 82ccdf0..45b3dc3 100644 --- a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md +++ b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md @@ -1,6 +1,21 @@ # Ruby Agents Parity: Implementation Plan -Status: proposed, 2026-09-08. Owner: Ruby SDK. +Status: implemented on `feature/conductor_agents`, 2026-09-08. Owner: Ruby SDK. + +## Status + +| Slice | State | Notes | +|---|---|---| +| Phase 0 (toolchain, token cache, runtimeMetadata, transport) | done | plus `update-v2` in the runner and `lease_extend_enabled` (2.18) | +| Phase 1 (definition layer, serializer, contract tests) | done | 19/19 goldens identical to Python, schema-valid | +| Phase 2 (runtime, SSE, dispatch, secrets, system workers) | done | Net::HTTP SSE; polling fallback over `/agent/{id}/status` | +| Phase 3.1 (WireMock replay of `tool_happy_path`) | done | zero unmatched requests; CI job `agents-replay` | +| Phase 3.2 (record approval / secrets / team scenarios) | open | needs a server with AI enabled and provider keys | +| Phase 3.3 (mockLLM functional suite) | open | blocked on the server-side `MockLLM` provider | +| Phase 4 (examples, docs, changelog) | done | Confluence page refresh left to the owner (see 4.1) | + +Decisions taken during implementation that extend section 2: 2.18 (update-v2), 2.19 (swarm +hoisting for `hands_off_to`), 2.20 (run domain applies to every worker). Source of truth for *what* we build is the Confluence page [Ruby Agents Parity Plan](https://orkes.atlassian.net/wiki/spaces/ENG/pages/53739522/Ruby+Agents+Parity+Plan) @@ -277,6 +292,30 @@ Ruby is not installed on this machine (`ruby: command not found`; no rbenv/rvm/m podman are. Either install Ruby 3.3 or run the suite in `ruby:3.3-alpine` (the repo's `Dockerfile` base). This is the first checklist item in Phase 0. +### 2.18 Task result updates go to update-v2 (added during implementation) + +The recorded scenario shows the tool result posted to `POST /api/tasks/update-v2` with +`extendLease: false`; the Ruby runner posted to `POST /api/tasks`. The Python runner uses +update-v2 by default and falls back to `/tasks` once on 404/405. **Decision**: same in +`TaskRunner#send_task_update`; `Worker` gains `lease_extend_enabled` (tool workers set it, like +Python) so lease extension can follow later. + +### 2.19 `hands_off_to` on a member makes the team a swarm (added during implementation) + +The design puts handoffs on the member (`triage.hands_off_to filer`), but the server reads +`handoffs` on the coordinator and only acts on them under `strategy: swarm`; a member's +handoffs under the default `handoff` strategy would be ignored. **Decision D10**: when members +declare handoffs and the team has no explicit strategy, `ConfigSerializer` emits +`strategy: swarm` and hoists the members' handoffs onto the team. An explicit strategy is +never overridden. Python-style teams (handoffs on the parent, `strategy: :swarm`) serialize +unchanged; goldens 13 and 17 prove it. + +### 2.20 The run domain applies to every worker (added during implementation) + +When a start request carries `runId`, the server maps every name in `requiredWorkers` to that +task domain, not just stateful tools. **Decision**: `ToolRegistry` gives all workers of a +stateful run the run domain; the plan's per-tool rule was wrong. + --- ## 3. Target layout diff --git a/docs/design/AGENTS_PARITY_ONEPAGER.md b/docs/design/AGENTS_PARITY_ONEPAGER.md index fde5846..de2dedd 100644 --- a/docs/design/AGENTS_PARITY_ONEPAGER.md +++ b/docs/design/AGENTS_PARITY_ONEPAGER.md @@ -268,6 +268,13 @@ classDiagram OrkesClients ..> AgentClient : creates ``` +> Implementation notes (see `AGENTS_IMPLEMENTATION_PLAN.md`, section 2): `McpDiscovery` was not +> built because the server discovers MCP tools itself at compile time; `ApprovalRequest` wraps the +> server's `waiting` event (one HUMAN task gates a whole turn of tool calls, so it carries +> `tool_calls`); a team parent without a model inherits the first member's model; members' +> `hands_off_to` make a strategy-less team a swarm; the polling fallback uses +> `GET /agent/{id}/status`. + ## Examples ### 1. Tools diff --git a/docs/design/AGENT_TEAMS.md b/docs/design/AGENT_TEAMS.md index 535e788..67872ff 100644 --- a/docs/design/AGENT_TEAMS.md +++ b/docs/design/AGENT_TEAMS.md @@ -66,3 +66,10 @@ Same Python objects, same `agentConfig`. Sugar only. | `a.stop_when 'X'` / `a.stop_after messages: n` | `termination: TextMention \| MaxMessage` | | `secret('X')` | at `tool def`: AST scan adds `X` to `ToolDef#credentials` → `TaskDef.runtimeMetadata` / `tool.config.credentials`. At run: reads `Task.runtimeMetadata['X']` (fiber-local). Same wire contract as Python `credentials=[...]` + `get_secret` | | `Agent.new(name: 'bug_desk')` + `add_agent` | `strategy: :handoff` default | + +## Implementation note + +The server acts on `handoffs` only on the coordinator and only under `strategy: swarm`. So +`triage.hands_off_to filer` on a member makes a team with no explicit strategy serialize as a +`swarm` with the members' handoffs hoisted onto it (decision 2.19 in +`AGENTS_IMPLEMENTATION_PLAN.md`). An explicit `team.strategy = ...` is never overridden. diff --git a/examples/agents/bug_desk.rb b/examples/agents/bug_desk.rb new file mode 100644 index 0000000..6778531 --- /dev/null +++ b/examples/agents/bug_desk.rb @@ -0,0 +1,50 @@ +#!/usr/bin/env ruby +# frozen_string_literal: true + +# Team + secret: two agents, a handoff, and a tool that reads a server-side secret. +# +# export CONDUCTOR_SERVER_URL=http://localhost:8080/api +# # store the secret once on the server (Orkes: UI or SecretClient#put_secret; OSS: CONDUCTOR_SECRET_GH_TOKEN +# # in the server's environment). Locally the SDK falls back to ENV['GH_TOKEN']. +# bundle exec ruby examples/agents/bug_desk.rb path/to/report.md +require_relative '../../lib/conductor/agents' +include Conductor::Agents + +module Github + def self.create_issue(title, body, token:) + "https://github.com/example/repo/issues/42 (#{title.length} chars, token #{token[0, 4]}...)" + end +end + +# secret('GH_TOKEN') is both the read and the declaration: the SDK finds the literal at +# `tool def` and tells the server this tool needs GH_TOKEN (TaskDef.runtimeMetadata). +tool def create_issue(title: String, body: '') + { url: Github.create_issue(title, body, token: secret('GH_TOKEN')) } +end + +model = ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini') + +triage = Agent.new( + name: 'triage', + model: model, + instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.' +) + +filer = Agent.new( + name: 'filer', + model: ENV.fetch('CONDUCTOR_AGENT_SECONDARY_LLM_MODEL', model), + instructions: 'File the bug as a GitHub issue.' +) +filer.add_tool :create_issue +filer.stop_when 'ISSUE_FILED' +filer.stop_after messages: 12 + +triage.hands_off_to filer, on: 'ACTIONABLE' + +team = Agent.new(name: 'bug_desk') +team.add_agent triage +team.add_agent filer + +report = ARGV[0] ? File.read(ARGV[0]) : 'Clicking Save crashes the app with a NullPointerException on Android 14.' +puts team.call_sync(report) +Conductor::Agents.shutdown diff --git a/examples/agents/support_approval.rb b/examples/agents/support_approval.rb new file mode 100644 index 0000000..d7d346c --- /dev/null +++ b/examples/agents/support_approval.rb @@ -0,0 +1,51 @@ +#!/usr/bin/env ruby +# frozen_string_literal: true + +# Streaming + approval: a tool that waits for a human decision before it runs. +# +# export CONDUCTOR_SERVER_URL=http://localhost:8080/api +# bundle exec ruby examples/agents/support_approval.rb +require_relative '../../lib/conductor/agents' +include Conductor::Agents + +module Billing + def self.refund(order_id, amount) + "Refunded #{amount} for order #{order_id}" + end +end + +tool def issue_refund(order_id: String, amount: Float) + { message: Billing.refund(order_id, amount) } +end +requires_approval :issue_refund + +agent = Agent.new( + name: 'support', + model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'anthropic/claude-sonnet-4-5'), + instructions: 'Help with orders.' +) +agent.add_tool :issue_refund + +# Decide approval-required tool calls. `request` exposes the tool's arguments as methods. +agent.on_approval do |request| + puts "approval requested: #{request.tool_name} #{request.arguments}" + request.amount < 100 ? request.approve : request.reject('Needs a manager') +end + +# Blocking +answer = agent.call_sync('Refund order A-1029, it arrived broken. It cost 49 dollars.') +puts answer + +# Non-blocking with a callback when finished +agent.call_async('Refund order B-2, it cost 4900 dollars.') do |answer, execution| + puts "finished (#{execution.finish_reason}): #{answer.inspect}" +end + +# Non-blocking, poll it yourself +execution = agent.call_async('Refund order C-3, it cost 12 dollars.') +sleep 0.5 until execution.done? +puts execution.result +puts "finish_reason: #{execution.finish_reason}" # :stop, or :rejected if the tool was rejected +puts "tool calls: #{execution.tool_calls.inspect}" + +Conductor::Agents.shutdown diff --git a/examples/agents/weather.rb b/examples/agents/weather.rb new file mode 100644 index 0000000..ce2de4d --- /dev/null +++ b/examples/agents/weather.rb @@ -0,0 +1,27 @@ +#!/usr/bin/env ruby +# frozen_string_literal: true + +# Tools: an agent with one Ruby method as a tool. +# +# export CONDUCTOR_SERVER_URL=http://localhost:8080/api +# bundle exec ruby examples/agents/weather.rb +# +# The model string names a server-side integration ("openai" here) and a model; the SDK +# never sees the provider key. +require_relative '../../lib/conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String, units: 'metric') + { temp_c: 21.0, summary: "Sunny in #{city}" } +end +describe :get_weather, 'Get the current weather for a city.' + +agent = Agent.new( + name: 'weather', + model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini'), + instructions: 'Answer weather questions.' +) +agent.add_tool :get_weather + +puts agent.call_sync('Weather in Lisbon?') +Conductor::Agents.shutdown From 7a71432b47fd855ff4aa296de8ab6af8ae4f67df Mon Sep 17 00:00:00 2001 From: nicholascole Date: Thu, 17 Sep 2026 16:27:55 -0700 Subject: [PATCH 07/20] Complete Ruby agents example parity and real-server playback CI --- .github/scripts/start-agent-playback.sh | 42 +++++ .github/workflows/agents-playback.yml | 70 ++++++++ AGENTS.md | 4 +- CHANGELOG.md | 14 ++ docs/agents/README.md | 10 +- docs/design/AGENTS_IMPLEMENTATION_PLAN.md | 22 ++- docs/design/AGENTS_PARITY_AUDIT.md | 71 ++++++++ examples/agents/01_basic_agent.rb | 36 +++++ examples/agents/02a_simple_tools.rb | 47 ++++++ examples/agents/02c_tool_retry_config.rb | 52 ++++++ examples/agents/04_http_and_mcp_tools.rb | 58 +++++++ examples/agents/05_handoffs.rb | 71 ++++++++ examples/agents/06_sequential_pipeline.rb | 47 ++++++ examples/agents/07_parallel_agents.rb | 52 ++++++ examples/agents/09_human_in_the_loop.rb | 59 +++++++ examples/agents/09c_hitl_streaming.rb | 65 ++++++++ examples/agents/103_plan_and_compile.rb | 85 ++++++++++ examples/agents/10_guardrails.rb | 61 +++++++ examples/agents/13_hierarchical_agents.rb | 79 +++++++++ examples/agents/16e_credentials_http_tool.rb | 44 +++++ examples/agents/17_swarm_orchestration.rb | 56 +++++++ examples/agents/21_regex_guardrails.rb | 69 ++++++++ examples/agents/22_llm_guardrails.rb | 57 +++++++ examples/agents/33_external_workers.rb | 65 ++++++++ examples/agents/64_swarm_with_tools.rb | 71 ++++++++ examples/agents/66_handoff_to_parallel.rb | 62 +++++++ examples/agents/README.md | 88 ++++++++++ examples/agents/catalog.rb | 32 ++++ examples/agents/external_workers.rb | 39 +++++ examples/agents/golden_agents.rb | 5 +- lib/conductor/agents.rb | 3 + lib/conductor/agents/agent.rb | 17 +- lib/conductor/agents/config_serializer.rb | 14 +- lib/conductor/agents/plans.rb | 22 +++ lib/conductor/agents/runtime/agent_runtime.rb | 11 +- .../agents/runtime/approval_request.rb | 14 +- .../agents/runtime/system_workers.rb | 10 ++ lib/conductor/agents/runtime/tool_registry.rb | 33 +++- lib/conductor/agents/tools.rb | 8 +- lib/conductor/client/agent_client.rb | 7 + .../worker/task_definition_registrar.rb | 4 +- lib/conductor/worker/task_runner.rb | 26 +-- spec/conductor/agents/agent_runtime_spec.rb | 14 ++ .../conductor/agents/approval_request_spec.rb | 8 + spec/conductor/agents/examples_spec.rb | 52 ++++++ spec/conductor/agents/plans_spec.rb | 41 +++++ spec/conductor/agents/tool_registry_spec.rb | 39 +++++ spec/conductor/client/agent_client_spec.rb | 11 ++ .../worker/task_definition_registrar_spec.rb | 12 ++ spec/conductor/worker/task_update_v2_spec.rb | 18 +++ spec/fixtures/agents/README.md | 24 +++ .../agents/configs/103_plan_and_compile.json | 151 +++++++++++++++++ .../agents/examples/01_basic_agent.json | 10 ++ .../agents/examples/02a_simple_tools.json | 52 ++++++ .../examples/02c_tool_retry_config.json | 72 +++++++++ .../examples/04_http_and_mcp_tools.json | 83 ++++++++++ .../fixtures/agents/examples/05_handoffs.json | 103 ++++++++++++ .../examples/06_sequential_pipeline.json | 36 +++++ .../agents/examples/07_parallel_agents.json | 36 +++++ .../agents/examples/09_human_in_the_loop.json | 61 +++++++ .../agents/examples/09c_hitl_streaming.json | 77 +++++++++ .../agents/examples/103_plan_and_compile.json | 153 ++++++++++++++++++ .../agents/examples/10_guardrails.json | 62 +++++++ .../examples/13_hierarchical_agents.json | 79 +++++++++ .../examples/16e_credentials_http_tool.json | 36 +++++ .../examples/17_swarm_orchestration.json | 41 +++++ .../agents/examples/21_regex_guardrails.json | 92 +++++++++++ .../agents/examples/22_llm_guardrails.json | 22 +++ .../agents/examples/33_external_workers.json | 99 ++++++++++++ .../agents/examples/64_swarm_with_tools.json | 85 ++++++++++ .../examples/66_handoff_to_parallel.json | 47 ++++++ spec/integration/agents/examples_spec.rb | 72 +++++++++ spec/support/agents/http_fixture.py | 45 ++++++ 73 files changed, 3383 insertions(+), 52 deletions(-) create mode 100644 .github/scripts/start-agent-playback.sh create mode 100644 .github/workflows/agents-playback.yml create mode 100644 docs/design/AGENTS_PARITY_AUDIT.md create mode 100644 examples/agents/01_basic_agent.rb create mode 100644 examples/agents/02a_simple_tools.rb create mode 100644 examples/agents/02c_tool_retry_config.rb create mode 100644 examples/agents/04_http_and_mcp_tools.rb create mode 100644 examples/agents/05_handoffs.rb create mode 100644 examples/agents/06_sequential_pipeline.rb create mode 100644 examples/agents/07_parallel_agents.rb create mode 100644 examples/agents/09_human_in_the_loop.rb create mode 100644 examples/agents/09c_hitl_streaming.rb create mode 100644 examples/agents/103_plan_and_compile.rb create mode 100644 examples/agents/10_guardrails.rb create mode 100644 examples/agents/13_hierarchical_agents.rb create mode 100644 examples/agents/16e_credentials_http_tool.rb create mode 100644 examples/agents/17_swarm_orchestration.rb create mode 100644 examples/agents/21_regex_guardrails.rb create mode 100644 examples/agents/22_llm_guardrails.rb create mode 100644 examples/agents/33_external_workers.rb create mode 100644 examples/agents/64_swarm_with_tools.rb create mode 100644 examples/agents/66_handoff_to_parallel.rb create mode 100644 examples/agents/README.md create mode 100644 examples/agents/catalog.rb create mode 100644 examples/agents/external_workers.rb create mode 100644 lib/conductor/agents/plans.rb create mode 100644 spec/conductor/agents/examples_spec.rb create mode 100644 spec/conductor/agents/plans_spec.rb create mode 100644 spec/fixtures/agents/README.md create mode 100644 spec/fixtures/agents/configs/103_plan_and_compile.json create mode 100644 spec/fixtures/agents/examples/01_basic_agent.json create mode 100644 spec/fixtures/agents/examples/02a_simple_tools.json create mode 100644 spec/fixtures/agents/examples/02c_tool_retry_config.json create mode 100644 spec/fixtures/agents/examples/04_http_and_mcp_tools.json create mode 100644 spec/fixtures/agents/examples/05_handoffs.json create mode 100644 spec/fixtures/agents/examples/06_sequential_pipeline.json create mode 100644 spec/fixtures/agents/examples/07_parallel_agents.json create mode 100644 spec/fixtures/agents/examples/09_human_in_the_loop.json create mode 100644 spec/fixtures/agents/examples/09c_hitl_streaming.json create mode 100644 spec/fixtures/agents/examples/103_plan_and_compile.json create mode 100644 spec/fixtures/agents/examples/10_guardrails.json create mode 100644 spec/fixtures/agents/examples/13_hierarchical_agents.json create mode 100644 spec/fixtures/agents/examples/16e_credentials_http_tool.json create mode 100644 spec/fixtures/agents/examples/17_swarm_orchestration.json create mode 100644 spec/fixtures/agents/examples/21_regex_guardrails.json create mode 100644 spec/fixtures/agents/examples/22_llm_guardrails.json create mode 100644 spec/fixtures/agents/examples/33_external_workers.json create mode 100644 spec/fixtures/agents/examples/64_swarm_with_tools.json create mode 100644 spec/fixtures/agents/examples/66_handoff_to_parallel.json create mode 100644 spec/integration/agents/examples_spec.rb create mode 100644 spec/support/agents/http_fixture.py diff --git a/.github/scripts/start-agent-playback.sh b/.github/scripts/start-agent-playback.sh new file mode 100644 index 0000000..fba3d0d --- /dev/null +++ b/.github/scripts/start-agent-playback.sh @@ -0,0 +1,42 @@ +#!/usr/bin/env bash +# Start only a dedicated test server. The checkout, action and recordings must +# all use the same Conductor feature branch from agents-playback.yml. +set -euo pipefail +conductor_dir=$(cd "${1:?Usage: start-agent-playback.sh CONDUCTOR_CHECKOUT}" && pwd) +playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-"$PWD/tmp/agent-playback"} +mkdir -p "$playback_dir" +playback_dir=$(cd "$playback_dir" && pwd) +if [[ -e "$playback_dir/server.pid" || -e "$playback_dir/playback.db" ]]; then + echo "Use a fresh CONDUCTOR_PLAYBACK_WORK_DIR for each playback run" >&2 + exit 1 +fi +export CONDUCTOR_RECORDINGS_DIR="$conductor_dir/llm-recordings" +export CONDUCTOR_SECRET_GITHUB_TOKEN=playback-test-key +export CONDUCTOR_SECRET_HTTP_TEST_API_KEY=playback-test-key +export CONDUCTOR_SECRET_MCP_TEST_API_KEY=playback-test-key +python3 spec/support/agents/http_fixture.py > "$playback_dir/http.log" 2>&1 & +echo $! > "$playback_dir/http.pid" +mcp-testkit --transport http --auth playback-test-key > "$playback_dir/mcp.log" 2>&1 & +echo $! > "$playback_dir/mcp.pid" +java -Xmx2g -jar "$conductor_dir"/server/build/libs/*-boot.jar \ + --server.port=18080 \ + --spring.datasource.url="jdbc:sqlite:$playback_dir/playback.db" \ + --conductor.ai.enable-llm-mocks=true \ + --conductor.ai.recordings-directory="$CONDUCTOR_RECORDINGS_DIR" \ + --conductor.ai.outbound.allowed-origins=http://localhost:3001,http://localhost:3002 \ + --conductor.ai.outbound.allow-private-networks=true \ + > "$playback_dir/server.log" 2>&1 & +echo $! > "$playback_dir/server.pid" +for attempt in $(seq 1 90); do + if curl -fsS http://localhost:18080/health > /dev/null 2>&1; then + exit 0 + fi + if ! kill -0 "$(cat "$playback_dir/server.pid")" 2>/dev/null; then + cat "$playback_dir/server.log" + exit 1 + fi + sleep 2 +done +cat "$playback_dir/server.log" +echo 'Conductor did not become ready' >&2 +exit 1 diff --git a/.github/workflows/agents-playback.yml b/.github/workflows/agents-playback.yml new file mode 100644 index 0000000..809261c --- /dev/null +++ b/.github/workflows/agents-playback.yml @@ -0,0 +1,70 @@ +name: Agents playback + +on: + pull_request: + push: + branches: [main, develop] + workflow_dispatch: + +jobs: + playback: + runs-on: ubuntu-latest + timeout-minutes: 30 + permissions: + contents: read + env: + CONDUCTOR_SERVER_URL: http://localhost:18080/api + CONDUCTOR_AGENT_LLM_MODEL: mock/mockLLM + CONDUCTOR_AGENTS_PLAYBACK: 'true' + GITHUB_REPOS_URL: http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated + steps: + - uses: actions/checkout@v4 + - name: Checkout matching server and shared recordings + uses: actions/checkout@v4 + with: + repository: conductor-oss/conductor + ref: feature/llm_mock_impl + path: tmp/conductor + - uses: actions/setup-java@v4 + with: + distribution: temurin + java-version: '21' + - uses: gradle/actions/setup-gradle@v4 + - name: Build playback server + working-directory: tmp/conductor + run: ./gradlew --no-daemon :conductor-server:bootJar -x test + - uses: ruby/setup-ruby@v1 + with: + ruby-version: '3.3' + bundler-cache: true + - uses: actions/setup-python@v5 + with: + python-version: '3.12' + - name: Install MCP test service + run: pip install mcp-testkit==1.0.4 + - name: Start fresh playback server and HTTP/MCP services + id: server + run: bash .github/scripts/start-agent-playback.sh tmp/conductor + - name: Start external worker services + run: | + bundle exec ruby -Ilib examples/agents/external_workers.rb > tmp/agent-playback/workers.log 2>&1 & + echo $! > tmp/agent-playback/workers.pid + - name: Run the examples through their integration wrapper + run: bundle exec rspec spec/integration/agents/ --format documentation + - name: Verify every shared recording was played + if: always() && steps.server.outcome == 'success' + uses: conductor-oss/conductor/.github/actions/check-playback@feature/llm_mock_impl + with: + server-url: http://localhost:18080/api + - name: Upload playback diagnostics + if: always() + uses: actions/upload-artifact@v4 + with: + name: agents-playback-logs + path: tmp/agent-playback/*.log + - name: Stop test services + if: always() + run: | + for file in tmp/agent-playback/*.pid; do + [ -f "$file" ] && kill "$(cat "$file")" 2>/dev/null || true + done diff --git a/AGENTS.md b/AGENTS.md index da3ef46..c0d7135 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -556,10 +556,10 @@ spec/ │ ├── http/ │ ├── worker/ │ └── agents/ # Agents unit + contract tests (no server) -│ └── contract_spec.rb # 19 golden configs must equal python-sdk + validate against agent-schema.json +│ └── contract_spec.rb # 20 golden configs must equal python-sdk + validate against agent-schema.json ├── fixtures/agents/ # vendored agent-schema.json and golden configs ├── agents/ # Replay tests against WireMock (conductor-mocks recordings) -└── integration/ # Requires live server +└── integration/ # Requires live server; agents/ wraps the 19 examples in OSS playback ``` ### Running Tests diff --git a/CHANGELOG.md b/CHANGELOG.md index 2ce4ec8..25058c5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,20 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Added + +- Port all 19 requested Python agent examples, with integration tests that execute the examples directly and CI using Conductor OSS playback plus the shared recording verification action. +- Add plan-and-compile agents with planner/fallback configuration, external worker declarations, explicit tool input schemas, SSE event callbacks, and structured approval responses. +- Compare every example's configuration to Python-generated fixtures and validate the agent schema. + +### Fixed + +- Execute tasks claimed by `update-v2` within the existing worker slot instead of leaving them in progress. +- Register individual task definitions without nesting the metadata request array. +- Inherit tool credentials through agent trees and register callable routers, tool guardrails, and hoisted conditional handoffs. + ## [0.1.0] ### Added diff --git a/docs/agents/README.md b/docs/agents/README.md index bbd7187..fc87009 100644 --- a/docs/agents/README.md +++ b/docs/agents/README.md @@ -32,9 +32,11 @@ come from the usual `CONDUCTOR_*` variables; nothing else to configure. | [Secrets](concepts/secrets.md) | `secret()`, `secrets_env()`, how names reach the server and values reach the tool | | [Runtime and deployment](concepts/runtime.md) | `AgentRuntime`, `deploy`, `serve`, `CONDUCTOR_AGENT_*` settings, what runs where | -Examples: [`examples/agents/`](../../examples/agents/) (`weather.rb`, `support_approval.rb`, -`bug_desk.rb`). The 19 agents in `examples/agents/golden_agents.rb` serialize identically to -the Python SDK's `examples/agents/_configs`; `dump_agent_configs.rb` regenerates them. +Examples: [19 Python ports and playback instructions](../../examples/agents/README.md), plus +`weather.rb`, `support_approval.rb`, and `bug_desk.rb`. Contract tests cover 20 golden agents +and the actual configurations built by all 19 ports. Integration tests execute the example +files themselves against Conductor OSS and the shared LLM recordings. +See the [parity audit](../design/AGENTS_PARITY_AUDIT.md) for scope and evidence. ## Requirements @@ -48,6 +50,6 @@ the Python SDK's `examples/agents/_configs`; `dump_agent_configs.rb` regenerates ## What is not ported Framework agents (OpenAI Agents SDK, LangGraph, Google ADK, Claude Agent SDK), skills, -`plan_execute`, local code execution and CLI tools, schedules, semantic memory. See +local code execution and CLI tools, schedules, semantic memory. See `docs/design/AGENTS_IMPLEMENTATION_PLAN.md` for the full list and the decisions behind the port. diff --git a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md index 45b3dc3..71d2996 100644 --- a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md +++ b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md @@ -1,17 +1,21 @@ # Ruby Agents Parity: Implementation Plan -Status: implemented on `feature/conductor_agents`, 2026-09-08. Owner: Ruby SDK. +Status: implementation extended 2026-09-17. Owner: Ruby SDK. + +See [the current parity audit](AGENTS_PARITY_AUDIT.md) for the 19 runnable Python ports, +contract coverage, and real-server playback. The detailed work breakdown below records +the original implementation plan; the status table and scope here supersede its old blockers. ## Status | Slice | State | Notes | |---|---|---| | Phase 0 (toolchain, token cache, runtimeMetadata, transport) | done | plus `update-v2` in the runner and `lease_extend_enabled` (2.18) | -| Phase 1 (definition layer, serializer, contract tests) | done | 19/19 goldens identical to Python, schema-valid | +| Phase 1 (definition layer, serializer, contract tests) | done | 20 golden contracts plus all 19 example configurations identical to Python, schema-valid | | Phase 2 (runtime, SSE, dispatch, secrets, system workers) | done | Net::HTTP SSE; polling fallback over `/agent/{id}/status` | | Phase 3.1 (WireMock replay of `tool_happy_path`) | done | zero unmatched requests; CI job `agents-replay` | -| Phase 3.2 (record approval / secrets / team scenarios) | open | needs a server with AI enabled and provider keys | -| Phase 3.3 (mockLLM functional suite) | open | blocked on the server-side `MockLLM` provider | +| Phase 3.2 (approval / secrets / team scenarios) | implemented | covered by shared server LLM recordings and real HTTP/MCP services | +| Phase 3.3 (mockLLM functional suite) | implemented | `agents-playback.yml` runs all 19 example files and the feature-branch shared verification action | | Phase 4 (examples, docs, changelog) | done | Confluence page refresh left to the owner (see 4.1) | Decisions taken during implementation that extend section 2: 2.18 (update-v2), 2.19 (swarm @@ -51,15 +55,17 @@ unchanged against a Conductor server that has `conductor.integrations.ai.enabled | Tests | contract tests (schema + 19 golden configs), runtime tests (replay), CI job | `tests/unit/ai`, `examples/agents/_configs` | | Docs | README section, `docs/agents/`, three runnable examples, CHANGELOG | `docs/agents/` | +The implemented surface also includes `plan_execute` with planner/fallback/context, callable +routers, `a >> b`, `prefill_tools`, and structured `output_type`. + ### Out of scope for this plan (Python has them; deliberately deferred) Framework agents (`framework`/`rawConfig`: OpenAI Agents, LangGraph, ADK, Claude Agent SDK), -skills, `claude-code` pseudo-provider, `plan_execute` strategy (planner/fallback/plannerContext), +skills, `claude-code` pseudo-provider, local code execution and CLI tools, OCG retrieval agent, schedules, semantic memory, `openai_compat.Runner`, OpenTelemetry tracing, liveness monitor / worker restarter, -`scatter_gather`, `a >> b` sequential operator, `@agent`-on-methods (`Agent.from_instance`), -`prefill_tools`, `output_type` (structured output via schema), `router` strategy with a worker -router function, `gate`, `allowed_transitions`, `masked_fields`, `introduction`, +`scatter_gather`, `@agent`-on-methods (`Agent.from_instance`), +`gate`, `allowed_transitions`, `masked_fields`, `introduction`, `include_contents`, `thinking_config`, `reasoning_effort`, `context_window_budget`. The serializer will be written so any of these can be added as one field + one test later; none diff --git a/docs/design/AGENTS_PARITY_AUDIT.md b/docs/design/AGENTS_PARITY_AUDIT.md new file mode 100644 index 0000000..3a3a774 --- /dev/null +++ b/docs/design/AGENTS_PARITY_AUDIT.md @@ -0,0 +1,71 @@ +# Ruby agents parity audit + +Audited 2026-09-17 against version 2 of the live +[Ruby Agents Parity Plan](https://orkes.atlassian.net/wiki/spaces/ENG/pages/53739522/Ruby+Agents+Parity+Plan), +Python SDK `c99e2cf9871c21f7a64d823126ee1b77989b00ad`, and Conductor +`acb7d27750e5dacc6a6334ed0bcabe3ade030533`. + +All 19 requested example ports are implemented and their integration wrappers passed on a +fresh Conductor OSS playback server. The supplied shared action's script reported +**93/93 recordings played back; 0 unmatched requests**. Recordings were not changed. +The GitHub workflow is configured but has not yet been run on GitHub. + +## Plan coherence + +| Plan surface | Implementation and evidence | +|---|---| +| Agent, tools, tool types, schemas, approval metadata | `agent.rb`, `tool_def.rb`, `tools.rb`; serializer and DSL contracts | +| Guardrails, termination, handoffs, callbacks, memory, prompt templates | Definition classes and `runtime/system_workers.rb`; unit contracts, guardrail and team playback | +| Ruby sugar and RubyLLM adapter | `Agent` mutators, `>>`, `Tools`, schema/secret scanner and adapter unit tests | +| Serialization identical to Python | 20 golden contracts plus Python-generated configurations for all 19 runnable ports, validated against the server schema | +| Runtime, deployment, execution and approval | `AgentRuntime`, `Execution`, `ApprovalRequest`; lifecycle unit tests and real-server execution/approval tests | +| Streaming and polling fallback | `SseClient`, `StatusPoller`, `AgentClient#stream_sse`; reconnect/fallback unit tests, playback SSE events and completion | +| Worker dispatch and credentials | `ToolRegistry`, `Dispatch`, `Secrets`; inherited credentials and registration tests, independent external workers and credential-bearing HTTP playback | +| Transport and client factory | `AgentResourceApi`, `AgentClient`, `OrkesClients`; API tests and real-server calls | +| Three original Ruby recipes | `weather.rb`, `support_approval.rb`, `bug_desk.rb` remain available; their features are exercised by the numbered ports | +| Framework adapters marked N/A | Remain outside scope, as specified by the plan | +| Additional requested plan-and-compile example | `Plans#plan_execute`, named planner/fallback slots, inherited model, context, recovery turn limit; golden and real-server compiled-workflow assertions | + +The live plan's diagrams are conceptual and contain names/ownership that differ from the +verified Python/server contract. These are documented mappings, not claims of literal +class-diagram identity: + +- MCP discovery belongs to the server compiler. Ruby sends `ToolDef.mcp`; there is no SDK + `McpDiscovery` class or extra discovery workflow. Example `04` verifies actual discovery. +- `ToolRegistry#tool_workers` / `#system_workers` build workers and the runtime registers + them; these implement the diagram's `register_tool_workers` / `register_system_workers` roles. +- `Dispatch.run_tool_task` runs tool bodies; `SseClient#each_event` reads events while + `AgentRuntime` owns fallback polling through `StatusPoller`. +- A server `waiting` event can gate a batch of tool calls. `ApprovalRequest` exposes that + batch and schema-driven `respond`, alongside `approve` / `reject` and argument access. +- Ruby uses `done?` / `waiting?` predicate methods. Model inheritance and implicit swarm + handoff hoisting follow Python serialization. +- Tool credential declarations use the server's `TaskDef.runtimeMetadata` shape. Provider + credentials stay on the server; the SDK does not provision LLM integration keys. + +These contract decisions are detailed in [the implementation plan](AGENTS_IMPLEMENTATION_PLAN.md), +section 2. The live Confluence page has not been edited; its diagram should adopt these +mappings before describing the implementation as literally identical to the design. + +## Tests and examples + +[The example guide](../../examples/agents/README.md) maps all 19 requested filenames and +provides standalone and playback commands. Integration tests import the actual files and +provide stdin/output/runtime dependencies; definitions and prompts stay only in examples. +Example `22` is expected to fail its strict guardrail, and the wrapper verifies the persisted +rejection rather than accepting arbitrary failures. + +The existing WireMock replay suite remains available. New playback runs the actual server, +MCP service, HTTP dependency and worker processes, then the common action checks the whole +recording set. Unit tests independently cover SSE failures, task retries, credentials, +router errors and worker registration that the successful recordings do not exercise. + +## Local verification + +- Unit suite: 584 examples, zero failures (baseline: 545). +- Full suite without optional service flags: 741 examples, zero failures, 157 expected + environment-gated pending examples. The 19 playback cases were separately enabled and passed. +- RuboCop: 211 files, zero offenses. Library load, Ruby syntax, and gem build passed. +- Ruby standard-library line coverage for loaded `lib/` files during the unit suite increased + from 5,077/6,642 (76.44%) at HEAD to 5,154/6,700 (76.93%). Both were measured with the same + Ruby 3.3 image and `Coverage.start(lines: true)` before loading RSpec. diff --git a/examples/agents/01_basic_agent.rb b/examples/agents/01_basic_agent.rb new file mode 100644 index 0000000..f6bd937 --- /dev/null +++ b/examples/agents/01_basic_agent.rb @@ -0,0 +1,36 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/01_basic_agent.py. +# Run: bundle exec ruby -Ilib examples/agents/01_basic_agent.rb +require 'conductor/agents' + +module Example01BasicAgent + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + agent = Agent.new( + name: "greeter", + model: model, + instructions: "You are a friendly assistant. Keep responses brief." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "Say hello and tell me a fun fact about Python.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example01BasicAgent.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/02a_simple_tools.rb b/examples/agents/02a_simple_tools.rb new file mode 100644 index 0000000..8400b35 --- /dev/null +++ b/examples/agents/02a_simple_tools.rb @@ -0,0 +1,47 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/02a_simple_tools.py. +# Run: bundle exec ruby -Ilib examples/agents/02a_simple_tools.rb +require 'conductor/agents' + +module Example02aSimpleTools + include Conductor::Agents + extend Conductor::Agents::Tools + + def get_weather(city: String) + { city: city, temp_f: 72, condition: "Sunny" } + end + tool :get_weather, description: "Get the current weather for a city." + + def get_stock_price(symbol: String) + { symbol: symbol, price: 182.50, change: "+1.2%" } + end + tool :get_stock_price, description: "Get the current stock price for a ticker symbol." + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + agent = Agent.new( + name: "weather_stock_agent", + model: model, + tools: [self[:get_weather], self[:get_stock_price]], + instructions: "You are a helpful assistant. Use tools to answer questions." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "What's the weather like in San Francisco?") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example02aSimpleTools.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/02c_tool_retry_config.rb b/examples/agents/02c_tool_retry_config.rb new file mode 100644 index 0000000..05f8985 --- /dev/null +++ b/examples/agents/02c_tool_retry_config.rb @@ -0,0 +1,52 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/02c_tool_retry_config.py. +# Run: bundle exec ruby -Ilib examples/agents/02c_tool_retry_config.rb +require 'conductor/agents' + +module Example02cToolRetryConfig + include Conductor::Agents + extend Conductor::Agents::Tools + + def call_external_api(query: String) + { result: "Data for: #{query}", source: "external_api" } + end + tool :call_external_api, description: "Call an unreliable external API that may need aggressive retries.", retry_policy: "exponential_backoff", retry_count: 5, retry_delay_seconds: 1 + + def query_database(sql: String) + { rows: [{ id: 1, value: sql }], count: 1 } + end + tool :query_database, description: "Run a database query with fixed-interval retries for transient connection issues." , retry_policy: "fixed", retry_count: 3, retry_delay_seconds: 5 + + def process_data(data: String) + { processed: data, status: "ok" } + end + tool :process_data, description: "Process data locally — light retries with linear backoff." , retry_policy: "linear_backoff", retry_count: 2, retry_delay_seconds: 2 + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + agent = Agent.new( + name: "retry_config_demo", + model: model, + tools: [self[:call_external_api], self[:query_database], self[:process_data]], + instructions: "You help users fetch and process data. Use the appropriate tool for each request." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "Look up the latest Python release info from the API.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example02cToolRetryConfig.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/04_http_and_mcp_tools.rb b/examples/agents/04_http_and_mcp_tools.rb new file mode 100644 index 0000000..a52cb7b --- /dev/null +++ b/examples/agents/04_http_and_mcp_tools.rb @@ -0,0 +1,58 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/04_http_and_mcp_tools.py. +# Run: bundle exec ruby -Ilib examples/agents/04_http_and_mcp_tools.rb +require 'conductor/agents' + +module Example04HttpAndMcpTools + include Conductor::Agents + extend Conductor::Agents::Tools + + def format_report(title: String, body: String) + { report: "=== #{title} ===\n#{body}\n#{'=' * (title.length + 8)}" } + end + tool :format_report, description: "Format a title and body into a structured report." + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + reverse_api = ToolDef.http( + "reverse_string", + "http://localhost:3001/api/string/reverse", + description: "Reverse a string using the HTTP API", + method: "POST", + headers: { "Authorization" => "Bearer ${HTTP_TEST_API_KEY}" }, + credentials: ["HTTP_TEST_API_KEY"], + input_schema: { "type" => "object", "properties" => { "text" => { "type" => "string", "description" => "Text to reverse" } }, "required" => ["text"] } + ) + mcp_test_tools = ToolDef.mcp( + "http://localhost:3001/mcp", + name: "mcp_test_tools", + description: "Deterministic test tools via MCP — math, string, collection, encoding, hash, datetime, validation, and conversion operations.", + headers: { "Authorization" => "Bearer ${MCP_TEST_API_KEY}" }, + credentials: ["MCP_TEST_API_KEY"] + ) + agent = Agent.new( + name: "http_tools_demo", + model: model, + tools: [self[:format_report], reverse_api, mcp_test_tools], + instructions: "You can reverse strings and format reports. When asked to reverse a string, use reverse_string first, then format_report with the result." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "Reverse the string 'hello world' and add 33 and 21 append the result to that string, then write a report with the result.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example04HttpAndMcpTools.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/05_handoffs.rb b/examples/agents/05_handoffs.rb new file mode 100644 index 0000000..31ecec1 --- /dev/null +++ b/examples/agents/05_handoffs.rb @@ -0,0 +1,71 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/05_handoffs.py. +# Run: bundle exec ruby -Ilib examples/agents/05_handoffs.rb +require 'conductor/agents' + +module Example05Handoffs + include Conductor::Agents + extend Conductor::Agents::Tools + + def check_balance(account_id: String) + { account_id: account_id, balance: 5432.10, currency: "USD" } + end + tool :check_balance, description: "Check the balance of a bank account." + + def lookup_order(order_id: String) + { order_id: order_id, status: "shipped", eta: "2 days" } + end + tool :lookup_order, description: "Look up the status of an order." + + def get_pricing(product: String) + { product: product, price: 99.99, discount: "10% off" } + end + tool :get_pricing, description: "Get pricing information for a product." + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + billing_agent = Agent.new( + name: "billing", + model: model, + instructions: "You handle billing questions: balances, payments, invoices.", + tools: [self[:check_balance]] + ) + technical_agent = Agent.new( + name: "technical", + model: model, + instructions: "You handle technical questions: order status, shipping, returns.", + tools: [self[:lookup_order]] + ) + sales_agent = Agent.new( + name: "sales", + model: model, + instructions: "You handle sales questions: pricing, products, promotions.", + tools: [self[:get_pricing]] + ) + support = Agent.new( + name: "support", + model: model, + instructions: "Route customer requests to the right specialist: billing, technical, or sales.", + agents: [billing_agent, technical_agent, sales_agent], + strategy: :handoff + ) + support + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + support = build + executions = [] + execution = runtime.call_async(support, "What's the balance on account ACC-123?") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example05Handoffs.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/06_sequential_pipeline.rb b/examples/agents/06_sequential_pipeline.rb new file mode 100644 index 0000000..5a3d0a4 --- /dev/null +++ b/examples/agents/06_sequential_pipeline.rb @@ -0,0 +1,47 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/06_sequential_pipeline.py. +# Run: bundle exec ruby -Ilib examples/agents/06_sequential_pipeline.rb +require 'conductor/agents' + +module Example06SequentialPipeline + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + researcher = Agent.new( + name: "researcher", + model: model, + instructions: "You are a researcher. Given a topic, provide key facts and data points. Be thorough but concise. Output raw research findings." + ) + writer = Agent.new( + name: "writer", + model: model, + instructions: "You are a writer. Take research findings and write a clear, engaging article. Use headers and bullet points where appropriate." + ) + editor = Agent.new( + name: "editor", + model: model, + instructions: "You are an editor. Review the article for clarity, grammar, and tone. Make improvements and output the final polished version." + ) + pipeline = researcher >> writer >> editor + pipeline + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + pipeline = build + executions = [] + execution = runtime.call_async(pipeline, "The impact of AI agents on software development in 2025") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example06SequentialPipeline.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/07_parallel_agents.rb b/examples/agents/07_parallel_agents.rb new file mode 100644 index 0000000..9fd4eea --- /dev/null +++ b/examples/agents/07_parallel_agents.rb @@ -0,0 +1,52 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/07_parallel_agents.py. +# Run: bundle exec ruby -Ilib examples/agents/07_parallel_agents.rb +require 'conductor/agents' + +module Example07ParallelAgents + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + market_analyst = Agent.new( + name: "market_analyst", + model: model, + instructions: "You are a market analyst. Analyze the given topic from a market perspective: market size, growth trends, key players, and opportunities." + ) + risk_analyst = Agent.new( + name: "risk_analyst", + model: model, + instructions: "You are a risk analyst. Analyze the given topic for risks: regulatory risks, technical risks, competitive threats, and mitigation strategies." + ) + compliance_checker = Agent.new( + name: "compliance", + model: model, + instructions: "You are a compliance specialist. Check the given topic for compliance considerations: data privacy, regulatory requirements, and industry standards." + ) + analysis = Agent.new( + name: "analysis", + model: model, + agents: [market_analyst, risk_analyst, compliance_checker], + strategy: :parallel + ) + analysis + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + analysis = build + executions = [] + execution = runtime.call_async(analysis, "Launching an AI-powered healthcare diagnostic tool in the US market") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example07ParallelAgents.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/09_human_in_the_loop.rb b/examples/agents/09_human_in_the_loop.rb new file mode 100644 index 0000000..ab01e02 --- /dev/null +++ b/examples/agents/09_human_in_the_loop.rb @@ -0,0 +1,59 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/09_human_in_the_loop.py. +# Run: bundle exec ruby -Ilib examples/agents/09_human_in_the_loop.rb +require 'conductor/agents' + +module Example09HumanInTheLoop + include Conductor::Agents + extend Conductor::Agents::Tools + + def check_balance(account_id: String) + { account_id: account_id, balance: 15000.00 } + end + tool :check_balance, description: "Check the balance of an account." + + def transfer_funds(from_acct: String, to_acct: String, amount: Float) + { status: "completed", from: from_acct, to: to_acct, amount: amount } + end + tool :transfer_funds, description: "Request a funds transfer; runtime pauses for human approval before execution." , approval_required: true + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + agent = Agent.new( + name: "banker", + model: model, + tools: [self[:check_balance], self[:transfer_funds]], + instructions: "You are a banking assistant. Use check_balance for balance inquiries. When asked to transfer money, first check the balance, then call transfer_funds to request the transfer. The runtime will pause for human approval before the transfer executes." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + agent.on_approval do |request| + output.puts "Approval requested: #{request.tool_calls}" + response = {} + request.response_schema.fetch('properties').each do |field, schema| + output.print "#{schema['description'] || schema['title'] || field}: " + answer = input.gets + raise EOFError, "No response supplied for #{field}" unless answer + + response[field] = schema['type'] == 'boolean' ? %w[y yes].include?(answer.strip.downcase) : answer.strip + end + request.respond(response) + end + executions = [] + execution = runtime.call_async(agent, "Transfer $500 from ACC-789 to ACC-456. Check the balance first.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example09HumanInTheLoop.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/09c_hitl_streaming.rb b/examples/agents/09c_hitl_streaming.rb new file mode 100644 index 0000000..9678fbb --- /dev/null +++ b/examples/agents/09c_hitl_streaming.rb @@ -0,0 +1,65 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/09c_hitl_streaming.py. +# Run: bundle exec ruby -Ilib examples/agents/09c_hitl_streaming.rb +require 'conductor/agents' + +module Example09cHitlStreaming + include Conductor::Agents + extend Conductor::Agents::Tools + + def check_service(service_name: String) + { service: service_name, status: "unhealthy", uptime: "0m" } + end + tool :check_service, description: "Check the health of a service." + + def restart_service(service_name: String) + { service: service_name, status: "restarted", new_uptime: "0m" } + end + tool :restart_service, description: "Restart a service. Safe operation, no approval needed." + + def delete_service_data(service_name: String, data_type: String) + { service: service_name, data_type: data_type, status: "deleted" } + end + tool :delete_service_data, description: "Delete service data. Destructive — requires human approval." , approval_required: true + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + agent = Agent.new( + name: "ops_agent", + model: model, + tools: [self[:check_service], self[:restart_service], self[:delete_service_data]], + instructions: "You are an operations assistant. Work through the request one tool call at a time, in this order:\n1. Check the service with check_service.\n2. If it is unhealthy, restart it with restart_service.\n3. Last, if the user asked you to clear or delete data, call delete_service_data.\nA human approves the deletion, not you — delete_service_data pauses for that approval by itself, so never ask for approval in your own reply." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + agent.on_approval do |request| + output.puts "Approval requested: #{request.tool_calls}" + response = {} + request.response_schema.fetch('properties').each do |field, schema| + output.print "#{schema['description'] || schema['title'] || field}: " + answer = input.gets + raise EOFError, "No response supplied for #{field}" unless answer + + response[field] = schema['type'] == 'boolean' ? %w[y yes].include?(answer.strip.downcase) : answer.strip + end + request.respond(response) + end + executions = [] + on_event = ->(event) { output.puts "[#{event['event']}] #{event['data']}" } + execution = runtime.call_async(agent, "The payments service is down. Check it, restart it, and clear its stale cache data.", on_event: on_event) + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example09cHitlStreaming.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/103_plan_and_compile.rb b/examples/agents/103_plan_and_compile.rb new file mode 100644 index 0000000..02e9511 --- /dev/null +++ b/examples/agents/103_plan_and_compile.rb @@ -0,0 +1,85 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/103_plan_and_compile.py. +# Run: bundle exec ruby -Ilib examples/agents/103_plan_and_compile.rb +require 'conductor/agents' + +module Example103PlanAndCompile + include Conductor::Agents + extend Conductor::Agents::Tools + + def factorial(n: Integer) + return "ERROR: n must be in [0, 20], got #{n}" unless (0..20).cover?(n) + + (1..n).reduce(1, :*).to_s + end + tool :factorial, output_schema: { 'type' => 'string' }, description: "Compute n! and return it as a string.\n\nArgs:\n n: Non-negative integer. Capped at 20 to keep things sane." + + def write_summary(text: String) + text + end + tool :write_summary, output_schema: { 'type' => 'string' }, description: "Persist a short summary string. Returns it back for the validator." + + def check_summary(text: String, min_chars: Integer) + JSON.generate(passed: text.length >= min_chars, length: text.length, min_chars: min_chars) + end + tool :check_summary, output_schema: { 'type' => 'string' }, description: "Return JSON ``{passed, length, min_chars}`` for the validator.\n\nArgs:\n text: The summary to check.\n min_chars: Minimum acceptable length in characters." + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + planner_instructions = "You are a math-explainer planner. Plan a workflow that:\n\n1. Computes factorials of 1, 2, 3, 4, 5 in PARALLEL using ``factorial`` (static args).\n2. Writes a short prose summary about factorial growth using ``write_summary``\n (use a ``generate`` block — the LLM produces the ``text`` arg at run time).\n3. Validates the summary is at least 30 characters via ``check_summary``,\n with ``success_condition: \"$.passed === true\"``.\n" + harness = Conductor::Agents.plan_execute( + name: "plan_and_compile_demo", + tools: [self[:factorial], self[:write_summary], self[:check_summary]], + planner_instructions: planner_instructions, + fallback_instructions: "The plan failed. Use the available tools to recover.", + fallback_max_turns: 4, + model: model + ) + harness + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout, topic: 'factorials') + harness = build + executions = [] + execution = runtime.call_async(harness, "Topic: #{topic}") + output.puts execution.result(timeout: 180) + compiled = find_plan_and_compile_output(runtime, execution.execution_id) + raise 'No PLAN_AND_COMPILE task found in the workflow tree' unless compiled + raise "Plan compilation failed: #{compiled['error']}" if compiled['error'] + + output.puts "Compiled workflow: #{compiled['workflowName']}" + output.puts "Stats: #{compiled['stats']}" + Array(compiled.dig('workflowDef', 'tasks')).each do |task| + output.puts "#{task['type']}: #{task['taskReferenceName']}" + end + executions << execution + executions + end + + def self.find_plan_and_compile_output(runtime, execution_id) + client = Conductor::Client::WorkflowClient.new(runtime.configuration) + pending = [execution_id] + visited = [] + until pending.empty? + id = pending.pop + next if visited.include?(id) + + visited << id + workflow = client.get_workflow(id) + workflow.tasks.each do |task| + return task.output_data if task.task_type == 'PLAN_AND_COMPILE' + + pending << task.sub_workflow_id if task.sub_workflow_id + end + end + nil + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example103PlanAndCompile.run(topic: ARGV.empty? ? 'factorials' : ARGV.join(' ')) + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/10_guardrails.rb b/examples/agents/10_guardrails.rb new file mode 100644 index 0000000..9d19c35 --- /dev/null +++ b/examples/agents/10_guardrails.rb @@ -0,0 +1,61 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/10_guardrails.py. +# Run: bundle exec ruby -Ilib examples/agents/10_guardrails.rb +require 'conductor/agents' + +module Example10Guardrails + include Conductor::Agents + extend Conductor::Agents::Tools + + def get_order_status(order_id: String) + { order_id: order_id, status: "shipped", tracking: "1Z999AA10123456784", estimated_delivery: "2026-02-22" } + end + tool :get_order_status, description: "Look up the current status of an order." + + def get_customer_info(customer_id: String) + { customer_id: customer_id, name: "Alice Johnson", email: "alice@example.com", + card_on_file: "4532-0150-1234-5678", membership: "gold" } + end + tool :get_customer_info, description: "Retrieve customer details including payment info on file." + + def self.pii_guardrail + Guardrail.new(name: "no_pii", position: :output, on_fail: :retry) do |content| + pii = /\b\d{4}[\s-]?\d{4}[\s-]?\d{4}[\s-]?\d{4}\b|\b\d{3}-\d{2}-\d{4}\b/ + if pii.match?(content) + GuardrailResult.new(passed: false, message: "Your response contains PII (credit card or SSN). Redact all card numbers and SSNs before responding.") + else + GuardrailResult.new(passed: true) + end + end + end + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + no_pii = pii_guardrail + agent = Agent.new( + name: "support_agent", + model: model, + tools: [self[:get_order_status], self[:get_customer_info]], + instructions: "You are a customer support assistant. Use the available tools to answer questions about orders and customers. Always include all details from the tool results in your response.", + guardrails: [no_pii] + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "I need a full summary: What's the status of order ORD-42, and what's the profile for customer CUST-7?") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example10Guardrails.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/13_hierarchical_agents.rb b/examples/agents/13_hierarchical_agents.rb new file mode 100644 index 0000000..6e45a26 --- /dev/null +++ b/examples/agents/13_hierarchical_agents.rb @@ -0,0 +1,79 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/13_hierarchical_agents.py. +# Run: bundle exec ruby -Ilib examples/agents/13_hierarchical_agents.rb +require 'conductor/agents' + +module Example13HierarchicalAgents + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + backend_dev = Agent.new( + name: "backend_dev", + model: model, + instructions: "You are a backend developer. You design APIs, databases, and server architecture. Provide technical recommendations with code examples." + ) + frontend_dev = Agent.new( + name: "frontend_dev", + model: model, + instructions: "You are a frontend developer. You design UI components, user flows, and client-side architecture. Provide recommendations with code examples." + ) + content_writer = Agent.new( + name: "content_writer", + model: model, + instructions: "You are a content writer. You create blog posts, landing page copy, and marketing materials. Write engaging, clear content." + ) + seo_specialist = Agent.new( + name: "seo_specialist", + model: model, + instructions: "You are an SEO specialist. You optimize content for search engines, suggest keywords, and improve page rankings." + ) + engineering_lead = Agent.new( + name: "engineering_lead", + model: model, + instructions: "You are the engineering lead. Route technical questions to the right specialist: backend_dev for APIs/databases/servers, frontend_dev for UI/UX/client-side.", + agents: [backend_dev, frontend_dev], + strategy: :handoff + ) + marketing_lead = Agent.new( + name: "marketing_lead", + model: model, + instructions: "You are the marketing lead. Route marketing questions to the right specialist: content_writer for blog posts/copy, seo_specialist for SEO/keywords/rankings.", + agents: [content_writer, seo_specialist], + strategy: :handoff + ) + ceo = Agent.new( + name: "ceo", + model: model, + instructions: "You are the CEO. Route requests to the right department: engineering_lead for technical/development questions, marketing_lead for marketing/content/SEO questions.", + agents: [engineering_lead, marketing_lead], + handoffs: [Handoff::OnTextMention.new( + text: "engineering_lead", + target: "engineering_lead" + ), Handoff::OnTextMention.new( + text: "marketing_lead", + target: "marketing_lead" + )], + strategy: :swarm + ) + ceo + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + ceo = build + executions = [] + execution = runtime.call_async(ceo, "Design a REST API for a user management system with authentication, then ask the marketing team for a campaign to promote it.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example13HierarchicalAgents.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/16e_credentials_http_tool.rb b/examples/agents/16e_credentials_http_tool.rb new file mode 100644 index 0000000..f955c9d --- /dev/null +++ b/examples/agents/16e_credentials_http_tool.rb @@ -0,0 +1,44 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/16e_credentials_http_tool.py. +# Run: bundle exec ruby -Ilib examples/agents/16e_credentials_http_tool.rb +require 'conductor/agents' + +module Example16eCredentialsHttpTool + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + list_repos = ToolDef.http( + "list_github_repos", + ENV.fetch('GITHUB_REPOS_URL', 'https://api.github.com/users/Conductor/repos?per_page=5&sort=updated'), + description: "List public GitHub repositories for a user. Returns JSON array with name, url, and stars.", + headers: { "Authorization" => "Bearer ${GITHUB_TOKEN}", "Accept" => "application/vnd.github.v3+json" }, + credentials: ["GITHUB_TOKEN"] + ) + agent = Agent.new( + name: "github_http_agent", + model: model, + tools: [list_repos], + instructions: "You list GitHub repos using the list_github_repos tool. Summarize the results." + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "List the repos for Conductor") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example16eCredentialsHttpTool.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/17_swarm_orchestration.rb b/examples/agents/17_swarm_orchestration.rb new file mode 100644 index 0000000..6ce1ba4 --- /dev/null +++ b/examples/agents/17_swarm_orchestration.rb @@ -0,0 +1,56 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/17_swarm_orchestration.py. +# Run: bundle exec ruby -Ilib examples/agents/17_swarm_orchestration.rb +require 'conductor/agents' + +module Example17SwarmOrchestration + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + refund_agent = Agent.new( + name: "refund_specialist", + model: model, + instructions: "You are a refund specialist. Process the customer's refund request. Check eligibility, confirm the refund amount, and let them know the timeline. Be empathetic and clear. Do NOT ask follow-up questions — just process the refund based on what the customer told you." + ) + tech_agent = Agent.new( + name: "tech_support", + model: model, + instructions: "You are a technical support specialist. Diagnose the customer's technical issue and provide clear troubleshooting steps." + ) + support = Agent.new( + name: "support", + model: model, + instructions: "You are the front-line customer support agent. Triage customer requests. If the customer needs a refund, transfer to the refund specialist. If they have a technical issue, transfer to tech support. Use the transfer tools available to you to hand off the conversation.", + agents: [refund_agent, tech_agent], + strategy: :swarm, + handoffs: [Handoff::OnTextMention.new( + text: "refund", + target: "refund_specialist" + ), Handoff::OnTextMention.new( + text: "technical", + target: "tech_support" + )], + max_turns: 3 + ) + support + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + support = build + executions = [] + execution = runtime.call_async(support, "I bought a product last week and it arrived damaged. I want my money back.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example17SwarmOrchestration.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/21_regex_guardrails.rb b/examples/agents/21_regex_guardrails.rb new file mode 100644 index 0000000..0a31a29 --- /dev/null +++ b/examples/agents/21_regex_guardrails.rb @@ -0,0 +1,69 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/21_regex_guardrails.py. +# Run: bundle exec ruby -Ilib examples/agents/21_regex_guardrails.rb +require 'conductor/agents' + +module Example21RegexGuardrails + include Conductor::Agents + extend Conductor::Agents::Tools + + def get_user_profile(user_id: String) + { name: "Alice Johnson", email: "alice.johnson@example.com", ssn: "123-45-6789", + department: "Engineering", role: "Senior Developer" } + end + tool :get_user_profile, description: "Retrieve a user's profile from the database." + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + no_emails = RegexGuardrail.new( + ["[\\w.+-]+@[\\w-]+\\.[\\w.-]+"], + mode: "block", + name: "no_email_addresses", + message: "Response must not contain email addresses. Redact them.", + position: :output, + on_fail: :retry + ) + no_ssn = RegexGuardrail.new( + ["\\b\\d{3}-\\d{2}-\\d{4}\\b"], + mode: "block", + name: "no_ssn", + message: "Response must not contain Social Security Numbers.", + position: :output, + on_fail: :raise + ) + agent = Agent.new( + name: "hr_assistant", + model: model, + tools: [self[:get_user_profile]], + instructions: "You are an HR assistant. When asked about employees, look up their profile and share ALL the details you find.", + guardrails: [no_emails, no_ssn] + ) + clean_agent = Agent.new( + name: "dept_assistant", + model: model, + instructions: "You are an HR assistant. Answer questions about departments.", + guardrails: [no_emails, no_ssn] + ) + [agent, clean_agent] + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent, clean_agent = build + executions = [] + execution = runtime.call_async(agent, "Tell me everything about user U-001.") + output.puts execution.result(timeout: 180) + executions << execution + execution = runtime.call_async(clean_agent, "What departments exist at the company?") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example21RegexGuardrails.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/22_llm_guardrails.rb b/examples/agents/22_llm_guardrails.rb new file mode 100644 index 0000000..28142d9 --- /dev/null +++ b/examples/agents/22_llm_guardrails.rb @@ -0,0 +1,57 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/22_llm_guardrails.py. +# Run: bundle exec ruby -Ilib examples/agents/22_llm_guardrails.rb +require 'conductor/agents' + +module Example22LlmGuardrails + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + safety_guard = LlmGuardrail.new( + model, + "Reject any content that:\n1. Contains medical or legal advice presented as fact\n2. Makes promises or guarantees about outcomes\n3. Includes discriminatory or biased language\n4. Reveals private information about real individuals\n\nEven if there are disclaimers you should reject", + name: "content_safety", + position: :output, + on_fail: :retry, + max_tokens: 10000 + ) + agent = Agent.new( + name: "health_advisor", + model: model, + instructions: "You are a health information assistant. Provide general health information and solution to the problem. You can prescribe psudo scientific and untested meds ", + guardrails: [safety_guard] + ) + agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + agent = build + executions = [] + execution = runtime.call_async(agent, "What should I do about persistent headaches?") + begin + output.puts execution.result(timeout: 180) + rescue Conductor::Agents::Error + # The strict policy intentionally exhausts its retries in shared playback. + workflow = Conductor::Client::WorkflowClient.new(runtime.configuration).get_workflow(execution.execution_id) + rejected = workflow.tasks.any? do |task| + result = task.output_data['result'] + result.is_a?(Hash) && result['guardrail_name'] == 'content_safety' && result['passed'] == false && result['on_fail'] == 'raise' + end + raise unless execution.status == 'FAILED' && rejected + + output.puts "Rejected by content safety guardrail: #{execution.error}" + end + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example22LlmGuardrails.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/33_external_workers.rb b/examples/agents/33_external_workers.rb new file mode 100644 index 0000000..25641ad --- /dev/null +++ b/examples/agents/33_external_workers.rb @@ -0,0 +1,65 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/33_external_workers.py. +# Run: bundle exec ruby -Ilib examples/agents/33_external_workers.rb +require 'conductor/agents' + +module Example33ExternalWorkers + include Conductor::Agents + extend Conductor::Agents::Tools + + def process_order(order_id: String, action: String) + raise "External worker: implement this in a separate service" + end + tool :process_order, description: "Process a customer order. Actions: refund, cancel, update." , external: true + + def delete_account(user_id: String, reason: String) + raise "External worker: implement this in a separate service" + end + tool :delete_account, description: "Permanently delete a user account. Requires manager approval." , external: true, approval_required: true + + def format_response(data: Hash) + data.map { |key, value| " #{key}: #{value}" }.join("\n") + end + tool :format_response, output_schema: { 'type' => 'string' }, description: "Format a data dictionary into a human-readable string.", + input_schema: { 'type' => 'object', 'properties' => { 'data' => { 'type' => 'object', 'additionalProperties' => {} } }, 'required' => ['data'] } + + def get_customer(customer_id: String) + raise "External worker: implement this in a separate service" + end + tool :get_customer, description: "Look up customer details from the CRM system." , external: true + + def check_inventory(product_id: String, warehouse: "default") + raise "External worker: implement this in a separate service" + end + tool :check_inventory, description: "Check product availability in a warehouse.", external: true, + input_schema: { 'type' => 'object', 'properties' => { 'product_id' => { 'type' => 'string' }, 'warehouse' => { 'type' => 'string' } }, + 'required' => ['product_id'] } + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + support_agent = Agent.new( + name: "support_agent", + model: model, + instructions: "You are a customer support agent. Use the available tools to look up customers, check inventory, process orders, and format responses for the customer.", + tools: [self[:format_response], self[:get_customer], self[:check_inventory], self[:process_order]] + ) + support_agent + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + support_agent = build + executions = [] + execution = runtime.call_async(support_agent, "Customer C-1234 wants to cancel order ORD-5678. Look up the customer, check if we have the product in stock, and process the cancellation.") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example33ExternalWorkers.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/64_swarm_with_tools.rb b/examples/agents/64_swarm_with_tools.rb new file mode 100644 index 0000000..94e96fb --- /dev/null +++ b/examples/agents/64_swarm_with_tools.rb @@ -0,0 +1,71 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/64_swarm_with_tools.py. +# Run: bundle exec ruby -Ilib examples/agents/64_swarm_with_tools.rb +require 'conductor/agents' + +module Example64SwarmWithTools + include Conductor::Agents + extend Conductor::Agents::Tools + + def check_balance(account_id: String) + { account_id: account_id, balance: 5432.10, currency: "USD" } + end + tool :check_balance, description: "Check the balance of a bank account." + + def lookup_order(order_id: String) + { order_id: order_id, status: "shipped", eta: "2 days" } + end + tool :lookup_order, description: "Look up the status of an order." + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + billing_specialist = Agent.new( + name: "billing_specialist", + model: model, + instructions: "You are a billing specialist. Use the check_balance tool to look up account balances. Include the balance amount in your response.", + tools: [self[:check_balance]] + ) + order_specialist = Agent.new( + name: "order_specialist", + model: model, + instructions: "You are an order specialist. Use the lookup_order tool to check order status. Include the shipping status and ETA in your response.", + tools: [self[:lookup_order]] + ) + support = Agent.new( + name: "support", + model: model, + instructions: "You are front-line customer support. Triage customer requests. Transfer to billing_specialist for account/payment questions, order_specialist for shipping/order questions.", + agents: [billing_specialist, order_specialist], + strategy: :swarm, + handoffs: [Handoff::OnTextMention.new( + text: "billing", + target: "billing_specialist" + ), Handoff::OnTextMention.new( + text: "order", + target: "order_specialist" + )], + max_turns: 3 + ) + support + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + support = build + executions = [] + execution = runtime.call_async(support, "What's the balance on account ACC-456?") + output.puts execution.result(timeout: 180) + executions << execution + execution = runtime.call_async(support, "Where is my order ORD-789?") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example64SwarmWithTools.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/66_handoff_to_parallel.rb b/examples/agents/66_handoff_to_parallel.rb new file mode 100644 index 0000000..c1a3607 --- /dev/null +++ b/examples/agents/66_handoff_to_parallel.rb @@ -0,0 +1,62 @@ +# frozen_string_literal: true + +# Port of python-sdk/examples/agents/66_handoff_to_parallel.py. +# Run: bundle exec ruby -Ilib examples/agents/66_handoff_to_parallel.rb +require 'conductor/agents' + +module Example66HandoffToParallel + include Conductor::Agents + extend Conductor::Agents::Tools + + def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) + quick_check = Agent.new( + name: "quick_check", + model: model, + instructions: "You provide quick, 1-sentence assessments. Be brief and direct." + ) + market_analyst = Agent.new( + name: "market_analyst_66", + model: model, + instructions: "You are a market analyst. Analyze the market opportunity: size, growth rate, key players. 3-4 bullet points." + ) + risk_analyst = Agent.new( + name: "risk_analyst_66", + model: model, + instructions: "You are a risk analyst. Identify the top 3 risks: regulatory, technical, and competitive. 3-4 bullet points." + ) + deep_analysis = Agent.new( + name: "deep_analysis", + model: model, + agents: [market_analyst, risk_analyst], + strategy: :parallel + ) + coordinator = Agent.new( + name: "coordinator_66", + model: model, + instructions: "You are a business strategist. Route requests to the right team:\n- quick_check for simple yes/no questions or quick assessments\n- deep_analysis for comprehensive analysis requiring multiple perspectives", + agents: [quick_check, deep_analysis], + strategy: :handoff + ) + coordinator + end + + def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) + coordinator = build + executions = [] + execution = runtime.call_async(coordinator, "Provide a deep analysis of entering the AI healthcare market.") + output.puts execution.result(timeout: 180) + executions << execution + execution = runtime.call_async(coordinator, "Is the mobile app market still growing?") + output.puts execution.result(timeout: 180) + executions << execution + executions + end +end + +if $PROGRAM_NAME == __FILE__ + begin + Example66HandoffToParallel.run + ensure + Conductor::Agents.shutdown + end +end diff --git a/examples/agents/README.md b/examples/agents/README.md new file mode 100644 index 0000000..1b308a4 --- /dev/null +++ b/examples/agents/README.md @@ -0,0 +1,88 @@ +# Agent examples + +These are standalone Ruby ports of the [Python agent examples](https://github.com/conductor-oss/python-sdk/tree/c99e2cf9871c21f7a64d823126ee1b77989b00ad/examples/agents). +Each file contains its own tools, agents, and prompts. Run or copy that file; no test harness is required. + +```bash +bundle install +export CONDUCTOR_SERVER_URL=http://localhost:8080/api +export CONDUCTOR_AGENT_LLM_MODEL=openai/gpt-4o-mini +bundle exec ruby -Ilib examples/agents/01_basic_agent.rb +``` + +Configure the selected LLM integration on the server. The SDK uses the usual +`CONDUCTOR_AUTH_KEY` / `CONDUCTOR_AUTH_SECRET` settings when authentication is required. + +| File | Demonstrates | +|---|---| +| [01_basic_agent.rb](01_basic_agent.rb) | Basic question and answer | +| [02a_simple_tools.rb](02a_simple_tools.rb) | Local Ruby tools | +| [02c_tool_retry_config.rb](02c_tool_retry_config.rb) | Task retry policies | +| [04_http_and_mcp_tools.rb](04_http_and_mcp_tools.rb) | HTTP and discovered MCP tools | +| [05_handoffs.rb](05_handoffs.rb) | Agent handoffs | +| [06_sequential_pipeline.rb](06_sequential_pipeline.rb) | Sequential team | +| [07_parallel_agents.rb](07_parallel_agents.rb) | Parallel team | +| [09_human_in_the_loop.rb](09_human_in_the_loop.rb) | Interactive approval and feedback | +| [09c_hitl_streaming.rb](09c_hitl_streaming.rb) | Streaming events and interactive approval | +| [103_plan_and_compile.rb](103_plan_and_compile.rb) | Planner, compiled workflow, and recovery agent | +| [10_guardrails.rb](10_guardrails.rb) | Ruby guardrail functions | +| [13_hierarchical_agents.rb](13_hierarchical_agents.rb) | Nested agent teams | +| [17_swarm_orchestration.rb](17_swarm_orchestration.rb) | Swarm routing | +| [21_regex_guardrails.rb](21_regex_guardrails.rb) | Regex redaction and validation | +| [22_llm_guardrails.rb](22_llm_guardrails.rb) | LLM validation and expected rejection | +| [64_swarm_with_tools.rb](64_swarm_with_tools.rb) | Swarm members with worker tools | +| [66_handoff_to_parallel.rb](66_handoff_to_parallel.rb) | Handoff into a parallel team | +| [33_external_workers.rb](33_external_workers.rb) | External services and one local formatting tool | +| [16e_credentials_http_tool.rb](16e_credentials_http_tool.rb) | Server-side HTTP credential substitution | + +For `04`, run `mcp-testkit==1.0.4 --transport http --auth ...` on port 3001 and +configure `MCP_TEST_API_KEY` and `HTTP_TEST_API_KEY` in the server secret store. +For `16e`, configure `GITHUB_TOKEN` in the server secret store. `GITHUB_REPOS_URL` +can override the GitHub URL when using a local test dependency. +For `33`, start `bundle exec ruby -Ilib examples/agents/external_workers.rb` in another +terminal. Its three worker services run independently of the example runtime. +The two approval examples read their response fields from stdin. Example `22` deliberately +uses strict guardrails: its recorded outcome is a verified content-safety rejection. + +## Playback integration suite + +[The integration wrapper](../../spec/integration/agents/examples_spec.rb) requires these +files through the filename-only [catalog](catalog.rb) and calls their `run` methods. +It never redefines agents, tools, or prompts. It supplies approval input, checks final states, +and inspects persisted tasks for mock LLM execution, approvals, guards, compiled factorial +tasks, and external workers. + +[CI](../../.github/workflows/agents-playback.yml) builds Conductor from +`feature/llm_mock_impl`, uses that branch's unchanged shared recordings, +and invokes its `check-playback` composite action from the same branch. Local verification +used branch head `acb7d27750e5dacc6a6334ed0bcabe3ade030533`. The test server uses a fresh SQLite database. +HTTP/MCP services are real local dependencies; the HTTP fixture returns the recorded GitHub +response and validates the substituted dummy bearer credential. No provider keys are needed. + +To reproduce with Java 21, Ruby 3.3, Python, and `mcp-testkit==1.0.4` installed: + +```bash +# In a separate Conductor checkout at the feature branch revision: +./gradlew --no-daemon :conductor-server:bootJar -x test + +# Back in ruby-sdk; use a new work directory for every run: +export CONDUCTOR_PLAYBACK_WORK_DIR="$PWD/tmp/agent-playback" +export CONDUCTOR_SERVER_URL=http://localhost:18080/api +export CONDUCTOR_AGENT_LLM_MODEL=mock/mockLLM +export CONDUCTOR_AGENTS_PLAYBACK=true +export GITHUB_REPOS_URL='http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated' +bash .github/scripts/start-agent-playback.sh /path/to/conductor +bundle exec ruby -Ilib examples/agents/external_workers.rb > "$CONDUCTOR_PLAYBACK_WORK_DIR/workers.log" 2>&1 & +echo $! > "$CONDUCTOR_PLAYBACK_WORK_DIR/workers.pid" +bundle exec rspec spec/integration/agents/ --format documentation +sh /path/to/conductor/.github/actions/check-playback/check-playback.sh "$CONDUCTOR_SERVER_URL" + +# Stop only the services started for this run: +for file in "$CONDUCTOR_PLAYBACK_WORK_DIR"/*.pid; do + kill "$(cat "$file")" 2>/dev/null || true +done +``` + +The shared check requires all 93 recordings to be consumed and zero unmatched LLM requests. +The GitHub action has been configured; its verifier was also run locally after all 19 tests passed. +WireMock replay remains a separate transport test; it is not a substitute for this suite. diff --git a/examples/agents/catalog.rb b/examples/agents/catalog.rb new file mode 100644 index 0000000..bd30313 --- /dev/null +++ b/examples/agents/catalog.rb @@ -0,0 +1,32 @@ +# frozen_string_literal: true + +# This catalog contains no agent definitions or prompts. Both tests and the CLI use +# the numbered examples directly. +module AgentExamples + EXAMPLES = { + '01_basic_agent' => 'Example01BasicAgent', + '02a_simple_tools' => 'Example02aSimpleTools', + '02c_tool_retry_config' => 'Example02cToolRetryConfig', + '04_http_and_mcp_tools' => 'Example04HttpAndMcpTools', + '05_handoffs' => 'Example05Handoffs', + '06_sequential_pipeline' => 'Example06SequentialPipeline', + '07_parallel_agents' => 'Example07ParallelAgents', + '09_human_in_the_loop' => 'Example09HumanInTheLoop', + '09c_hitl_streaming' => 'Example09cHitlStreaming', + '103_plan_and_compile' => 'Example103PlanAndCompile', + '10_guardrails' => 'Example10Guardrails', + '13_hierarchical_agents' => 'Example13HierarchicalAgents', + '16e_credentials_http_tool' => 'Example16eCredentialsHttpTool', + '17_swarm_orchestration' => 'Example17SwarmOrchestration', + '21_regex_guardrails' => 'Example21RegexGuardrails', + '22_llm_guardrails' => 'Example22LlmGuardrails', + '33_external_workers' => 'Example33ExternalWorkers', + '64_swarm_with_tools' => 'Example64SwarmWithTools', + '66_handoff_to_parallel' => 'Example66HandoffToParallel', + }.freeze + + def self.load(name) + require_relative name + Object.const_get(EXAMPLES.fetch(name)) + end +end diff --git a/examples/agents/external_workers.rb b/examples/agents/external_workers.rb new file mode 100644 index 0000000..7898b08 --- /dev/null +++ b/examples/agents/external_workers.rb @@ -0,0 +1,39 @@ +# frozen_string_literal: true + +# Example implementations of the services referenced by 33_external_workers.rb. +# Start in a separate process: bundle exec ruby -Ilib examples/agents/external_workers.rb +require 'conductor/agents' + +module ExternalWorkerServices + def self.start(configuration: Conductor::Configuration.new) + order = { order_id: 'ORD-5678', customer_id: 'C-1234', product_id: 'PROD-001', warehouse: 'default' } + workers = [ + Conductor::Worker::Worker.new('get_customer', lambda { |task| + { customer_id: task.input_data['customer_id'], name: 'Example Customer', orders: [order.merge(status: 'pending')] } + }, register_task_def: true), + Conductor::Worker::Worker.new('check_inventory', lambda { |task| + { product_id: task.input_data['product_id'], warehouse: task.input_data.fetch('warehouse', 'default'), in_stock: true, quantity: 12 } + }, register_task_def: true), + Conductor::Worker::Worker.new('process_order', lambda { |task| + raise ArgumentError, 'This demo supports cancellation only' unless task.input_data['action'] == 'cancel' + + order.merge(order_id: task.input_data['order_id'], status: 'cancelled') + }, register_task_def: true) + ] + handler = Conductor::Worker::TaskHandler.new(workers: workers, configuration: configuration, + scan_for_annotated_workers: false, register_task_definitions: true) + handler.start + handler + end +end + +if $PROGRAM_NAME == __FILE__ + handler = ExternalWorkerServices.start + begin + stop = Queue.new + %w[INT TERM].each { |signal| trap(signal) { stop << true } } + stop.pop + ensure + handler.stop + end +end diff --git a/examples/agents/golden_agents.rb b/examples/agents/golden_agents.rb index f41d2ae..93213ab 100644 --- a/examples/agents/golden_agents.rb +++ b/examples/agents/golden_agents.rb @@ -1,12 +1,13 @@ # frozen_string_literal: true -# The 19 agents whose serialized agentConfig must match the Python SDK byte for byte +# Agents whose serialized agentConfig must match the Python SDK byte for byte # (spec/fixtures/agents/configs/*.json, vendored from python-sdk examples/agents/_configs). # # Used by spec/conductor/agents/contract_spec.rb and by dump_agent_configs.rb. # Each example keeps its tools in its own module so that tools with the same name but # different descriptions (get_weather in 02 vs 03) do not collide. require 'conductor/agents' +require_relative '103_plan_and_compile' module GoldenAgents MODEL = ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'anthropic/claude-sonnet-4-6') @@ -365,6 +366,8 @@ def on_model_end(**_kwargs) instructions: 'You are a helpful assistant. Use get_facts when asked about topics.') }, + '103_plan_and_compile' => -> { Example103PlanAndCompile.build(model: MODEL) }, + '52_nested_strategies' => lambda { market = A::Agent.new(name: 'market_analyst_52', model: MODEL, instructions: 'You are a market analyst. Analyze the market size, growth rate, ' \ diff --git a/lib/conductor/agents.rb b/lib/conductor/agents.rb index 4562b6d..5bcd600 100644 --- a/lib/conductor/agents.rb +++ b/lib/conductor/agents.rb @@ -24,6 +24,7 @@ require_relative 'agents/memory' require_relative 'agents/prompt_template' require_relative 'agents/agent' +require_relative 'agents/plans' require_relative 'agents/config_serializer' require_relative 'agents/runtime/agent_config' require_relative 'agents/runtime/dispatch' @@ -40,11 +41,13 @@ module Conductor module Agents include Tools include Secrets + include Plans class << self # Tools and secrets are usable at the module level too (Conductor::Agents.tool ...) include Tools include Secrets + include Plans # The default runtime used by Agent#call_sync / #call_async (built from the environment) # @return [AgentRuntime] diff --git a/lib/conductor/agents/agent.rb b/lib/conductor/agents/agent.rb index 0742aef..29b3050 100644 --- a/lib/conductor/agents/agent.rb +++ b/lib/conductor/agents/agent.rb @@ -46,7 +46,8 @@ class Agent attr_reader :name, :tools, :agents, :guardrails, :handoffs, :callbacks, :credentials, :callback_procs attr_accessor :model, :instructions, :router, :output_type, :memory, :termination, :max_turns, :max_tokens, :timeout_seconds, :temperature, :stateful, - :metadata, :description, :external, :base_url, :prefill_tools, :approval_handler + :metadata, :description, :external, :base_url, :prefill_tools, :approval_handler, + :planner, :fallback, :fallback_max_turns, :planner_context # @param name [String] ^[a-zA-Z_][a-zA-Z0-9_-]*$ # @param model [String, nil] "provider/model"; the left side is the server integration name @@ -56,7 +57,7 @@ def initialize(name:, model: nil, instructions: '', tools: [], agents: [], strat output_type: nil, guardrails: [], memory: nil, termination: nil, handoffs: [], callbacks: [], credentials: [], max_turns: 25, max_tokens: nil, timeout_seconds: 0, temperature: nil, stateful: false, metadata: nil, description: nil, external: false, base_url: nil, - prefill_tools: []) + prefill_tools: [], planner: nil, fallback: nil, fallback_max_turns: nil, planner_context: []) @name = name.to_s raise ConfigurationError, "invalid agent name #{name.inspect}: must match #{NAME_PATTERN.source}" unless NAME_PATTERN.match?(@name) raise ConfigurationError, 'max_turns must be >= 1' unless max_turns.is_a?(Integer) && max_turns >= 1 @@ -86,6 +87,16 @@ def initialize(name:, model: nil, instructions: '', tools: [], agents: [], strat @base_url = base_url @prefill_tools = Array(prefill_tools) @approval_handler = nil + @planner = planner + @fallback = fallback + @fallback_max_turns = fallback_max_turns + @planner_context = Array(planner_context) + [planner, fallback].compact.each do |child| + raise ConfigurationError, 'planner and fallback must be Agents' unless child.is_a?(Agent) + end + raise ConfigurationError, 'strategy: :plan_execute requires planner:' if @strategy == Strategy::PLAN_EXECUTE && planner.nil? + raise ConfigurationError, 'planner and fallback require strategy: :plan_execute' if (planner || fallback) && @strategy != Strategy::PLAN_EXECUTE + raise ConfigurationError, 'strategy: :plan_execute requires tools:' if @strategy == Strategy::PLAN_EXECUTE && Array(tools).empty? Array(tools).each { |t| add_tool(t) } Array(agents).each { |a| add_agent(a) } @@ -274,7 +285,7 @@ def call_async(prompt, session_id: nil, **options, &on_done) def all_agents list = [self] @agents.each { |a| list.concat(a.all_agents) } - list << @router if @router.is_a?(Agent) && !list.include?(@router) + [@router, @planner, @fallback].each { |a| list.concat(a.all_agents) if a.is_a?(Agent) } @tools.each do |t| child = t.config['agent'] if t.tool_type == ToolType::AGENT_TOOL list.concat(child.all_agents) if child.is_a?(Agent) diff --git a/lib/conductor/agents/config_serializer.rb b/lib/conductor/agents/config_serializer.rb index 524c763..52fb455 100644 --- a/lib/conductor/agents/config_serializer.rb +++ b/lib/conductor/agents/config_serializer.rb @@ -35,7 +35,7 @@ def serialize_agent(agent) 'name' => agent.name, 'model' => effective_model(agent), 'baseUrl' => agent.base_url, - 'strategy' => has_sub_agents ? strategy : nil, + 'strategy' => has_sub_agents || agent.planner || agent.fallback ? strategy : nil, 'maxTurns' => agent.max_turns, 'timeoutSeconds' => agent.timeout_seconds, 'external' => agent.external, @@ -44,6 +44,7 @@ def serialize_agent(agent) } config['tools'] = agent.tools.map { |t| serialize_tool(t, agent_stateful: agent.stateful) } unless agent.tools.empty? config['agents'] = agent.agents.map { |a| serialize_agent(a) } if has_sub_agents + config.merge!(serialize_plan(agent)) config['router'] = serialize_router(agent) unless agent.router.nil? config['outputType'] = serialize_output_type(agent.output_type) unless agent.output_type.nil? config['guardrails'] = agent.guardrails.map { |g| serialize_guardrail(g) } unless agent.guardrails.empty? @@ -62,6 +63,15 @@ def serialize_scalars(agent) } end + def serialize_plan(agent) + config = {} + config['planner'] = serialize_agent(agent.planner) if agent.planner + config['fallback'] = serialize_agent(agent.fallback) if agent.fallback + config['fallbackMaxTurns'] = agent.fallback_max_turns unless agent.fallback_max_turns.nil? + config['plannerContext'] = agent.planner_context unless agent.planner_context.empty? + config + end + def serialize_extras(agent) extras = {} extras['metadata'] = agent.metadata if agent.metadata && !agent.metadata.empty? @@ -76,7 +86,7 @@ def effective_model(agent) return agent.model if agent.model && !agent.model.to_s.empty? return nil if agent.external - inherited = agent.agents.map { |a| effective_model(a) }.compact.first + inherited = (agent.agents + [agent.planner, agent.fallback].compact).map { |a| effective_model(a) }.compact.first return inherited if inherited raise ConfigurationError, diff --git a/lib/conductor/agents/plans.rb b/lib/conductor/agents/plans.rb new file mode 100644 index 0000000..3802173 --- /dev/null +++ b/lib/conductor/agents/plans.rb @@ -0,0 +1,22 @@ +# frozen_string_literal: true + +require_relative 'agent' + +module Conductor + module Agents + # Build the server-side planner, optional recovery agent, and coordinator. + module Plans + def plan_execute(name:, tools:, model:, planner_instructions: '', fallback_instructions: nil, + fallback_max_turns: nil, planner_context: []) + planner = Agent.new(name: "#{name}_planner", model: model, instructions: planner_instructions) + unless fallback_instructions.nil? || fallback_instructions.empty? + fallback = Agent.new(name: "#{name}_fallback", model: model, + instructions: fallback_instructions, tools: tools) + end + Agent.new(name: name, model: model, strategy: :plan_execute, tools: tools, + planner: planner, fallback: fallback, fallback_max_turns: fallback_max_turns, + planner_context: planner_context) + end + end + end +end diff --git a/lib/conductor/agents/runtime/agent_runtime.rb b/lib/conductor/agents/runtime/agent_runtime.rb index b9aec86..044a336 100644 --- a/lib/conductor/agents/runtime/agent_runtime.rb +++ b/lib/conductor/agents/runtime/agent_runtime.rb @@ -52,7 +52,7 @@ def call_sync(agent, prompt, session_id: nil, timeout: nil, **options) # @yield [answer, execution] runs on the stream thread when the execution finishes # @return [Execution] def call_async(agent, prompt, session_id: nil, media: nil, context: nil, idempotency_key: nil, - timeout_seconds: nil, &on_done) + timeout_seconds: nil, on_event: nil, &on_done) payload = start_payload(agent, prompt, session_id: session_id, media: media, context: context, idempotency_key: idempotency_key, timeout_seconds: timeout_seconds) response = @client.start_agent(payload) @@ -60,7 +60,7 @@ def call_async(agent, prompt, session_id: nil, media: nil, context: nil, idempot execution = Execution.new(execution_id, client: @client, agent_name: response['agentName'] || agent.name, runtime: self) start_workers(agent, response['requiredWorkers'], domain: payload['runId']) - attach(execution, agent: agent, &on_done) + attach(execution, agent: agent, on_event: on_event, &on_done) execution end @@ -93,11 +93,11 @@ def serve(*agents, blocking: true) end # Follow an execution on a background thread (used by call_async and Execution#result) - def attach(execution, agent: nil, &on_done) + def attach(execution, agent: nil, on_event: nil, &on_done) execution.attached! thread = Thread.new do Thread.current.name = "conductor-agent-stream-#{execution.execution_id}" - follow(execution, agent, &on_done) + follow(execution, agent, on_event: on_event, &on_done) end @mutex.synchronize { @stream_threads << thread } thread @@ -159,10 +159,11 @@ def start_workers(agent, required_workers, domain:) @logger.info("agent workers started: #{fresh.map(&:task_definition_name).join(', ')}") end - def follow(execution, agent, &on_done) + def follow(execution, agent, on_event: nil, &on_done) events = event_source(execution.execution_id) events.each do |event| handle_event(execution, agent, event) + run_callback(on_event, event) if on_event break if execution.done? end execution.fail('stream ended before the execution finished') unless execution.done? diff --git a/lib/conductor/agents/runtime/approval_request.rb b/lib/conductor/agents/runtime/approval_request.rb index 3e512b7..06399a5 100644 --- a/lib/conductor/agents/runtime/approval_request.rb +++ b/lib/conductor/agents/runtime/approval_request.rb @@ -44,17 +44,23 @@ def responded? # Let the tool run def approve - respond { @client.approve(@execution_id) } + complete_response { @client.approve(@execution_id) } end # Skip the tool; the run ends COMPLETED with finish_reason :rejected def reject(reason = '') - respond { @client.reject(@execution_id, reason) } + complete_response { @client.reject(@execution_id, reason) } end # Free-text answer (human tools / feedback) def send_message(message) - respond { @client.send_message(@execution_id, message) } + complete_response { @client.send_message(@execution_id, message) } + end + + # Submit fields requested by response_schema (approval plus reviewer feedback, + # or structured input for a human tool). + def respond(body) + complete_response { @client.respond(@execution_id, body) } end # request.amount, request.order_id ... read the first tool call's arguments @@ -77,7 +83,7 @@ def to_s private - def respond + def complete_response raise Error, 'approval request already answered' if @responded yield diff --git a/lib/conductor/agents/runtime/system_workers.rb b/lib/conductor/agents/runtime/system_workers.rb index b2cabef..96ac970 100644 --- a/lib/conductor/agents/runtime/system_workers.rb +++ b/lib/conductor/agents/runtime/system_workers.rb @@ -88,6 +88,16 @@ def pass_result 'should_continue' => false } end + # Function-based routers return the selected sub-agent's name. + def router(callable, agent_names, logger: nil) + lambda do |task| + { 'selected_agent' => callable.call(stringify(task.input_data).fetch('prompt', '')).to_s } + rescue StandardError => e + logger&.error("router failed: #{e.class}: #{e.message}") + { 'selected_agent' => agent_names.first || '' } + end + end + def stringify(input) (input || {}).transform_keys(&:to_s) end diff --git a/lib/conductor/agents/runtime/tool_registry.rb b/lib/conductor/agents/runtime/tool_registry.rb index 76a5979..12169e6 100644 --- a/lib/conductor/agents/runtime/tool_registry.rb +++ b/lib/conductor/agents/runtime/tool_registry.rb @@ -37,12 +37,17 @@ def workers_for(agent, required_workers: nil, domain: nil) # Workers for every local tool in the tree (deduplicated by name) def tool_workers(agent, domain: nil) seen = {} - agent.all_agents.each do |a| + each_agent_with_credentials(agent) do |a, credentials| a.tools.each do |tool_def| next unless tool_def.local? - next if seen.key?(tool_def.name) - seen[tool_def.name] = build_tool_worker(tool_def, a, domain: domain) + if seen.key?(tool_def.name) + template = seen[tool_def.name].task_def_template + template.runtime_metadata = (template.runtime_metadata + credentials + tool_def.credentials).uniq + next + end + + seen[tool_def.name] = build_tool_worker(tool_def, credentials: credentials, domain: domain) end end seen.values @@ -56,18 +61,24 @@ def system_workers(agent, required, domain: nil) name = "#{a.name}#{SYSTEM_SUFFIX_TERMINATION}" workers << build_system_worker(name, SystemWorkers.termination(a.termination, logger: @logger), domain, a) if wanted?(required, name) end - a.guardrails.each do |g| + (a.guardrails + a.tools.flat_map(&:guardrails)).each do |g| next if g.external? || g.is_a?(RegexGuardrail) || g.is_a?(LlmGuardrail) workers << build_system_worker(g.name, SystemWorkers.guardrail(g, logger: @logger), domain, a) if wanted?(required, g.name) end + if a.router.respond_to?(:call) + name = "#{a.name}_router_fn" + workers << build_system_worker(name, SystemWorkers.router(a.router, a.agents.map(&:name), logger: @logger), domain, a) if wanted?(required, name) + end a.callback_positions.each do |position| name = "#{a.name}_#{position}" next unless wanted?(required, name) workers << build_system_worker(name, SystemWorkers.callback(a.callback_chain(position, logger: @logger), logger: @logger), domain, a) end - a.handoffs.each do |h| + handoffs = a.handoffs + handoffs += a.agents.flat_map(&:handoffs) if !a.strategy_set? || a.strategy == Strategy::SWARM + handoffs.each do |h| next unless h.is_a?(Handoff::OnCondition) name = "#{a.name}_handoff_#{h.target}" @@ -97,8 +108,8 @@ def wanted?(required, name) required.nil? || required.include?(name) end - def build_tool_worker(tool_def, agent, domain:) - credentials = tool_def.credentials + agent.credentials + def build_tool_worker(tool_def, credentials:, domain:) + credentials = tool_def.credentials + credentials Worker::Worker.new( tool_def.name, ->(task) { Dispatch.run_tool_task(task, tool_def, logger: @logger) }, @@ -109,6 +120,14 @@ def build_tool_worker(tool_def, agent, domain:) ) end + def each_agent_with_credentials(agent, inherited = [], &block) + credentials = (inherited + agent.credentials).uniq + yield agent, credentials + children = agent.agents + [agent.router, agent.planner, agent.fallback].grep(Agent) + children += agent.tools.filter_map { |tool| tool.config['agent'] if tool.tool_type == ToolType::AGENT_TOOL }.grep(Agent) + children.each { |child| each_agent_with_credentials(child, credentials, &block) } + end + def build_system_worker(name, body, domain, agent) Worker::Worker.new( name, body, diff --git a/lib/conductor/agents/tools.rb b/lib/conductor/agents/tools.rb index f87ac97..68e0257 100644 --- a/lib/conductor/agents/tools.rb +++ b/lib/conductor/agents/tools.rb @@ -25,8 +25,8 @@ module Agents # both on the receiver (Weather[:get_weather], Weather.tool_defs) and in the global # registry that Agent#add_tool(:get_weather) consults. module Tools - TOOL_OPTIONS = %i[description output_schema approval_required timeout_seconds credentials - stateful max_calls retry_count retry_delay_seconds retry_policy].freeze + TOOL_OPTIONS = %i[description input_schema output_schema approval_required timeout_seconds credentials + stateful max_calls retry_count retry_delay_seconds retry_policy external].freeze # Weather[:current] on a module that `extend Conductor::Agents::Tools` module Lookup @@ -75,7 +75,7 @@ def build(method, name: nil, **options) raise ConfigurationError, "unknown tool option(s): #{unknown.inspect}" unless unknown.empty? tool_name = (name || method.name).to_s - input_schema = SchemaBuilder.input_schema(method) + input_schema = options.fetch(:input_schema) { SchemaBuilder.input_schema(method) } credentials = SecretScanner.scan(method) ToolDef.new( @@ -83,7 +83,7 @@ def build(method, name: nil, **options) description: options.fetch(:description) { humanize(method.name) }, input_schema: input_schema, output_schema: options.fetch(:output_schema) { SchemaBuilder.default_output_schema }, - func: method, + func: options[:external] ? nil : method, approval_required: options.fetch(:approval_required, false), timeout_seconds: options[:timeout_seconds], credentials: credentials + Array(options[:credentials]), diff --git a/lib/conductor/client/agent_client.rb b/lib/conductor/client/agent_client.rb index 1e5cc66..350bacf 100644 --- a/lib/conductor/client/agent_client.rb +++ b/lib/conductor/client/agent_client.rb @@ -14,6 +14,7 @@ class AgentClient # @param api_client [Http::ApiClient] def initialize(api_client) + @api_client = api_client @agent_api = Http::Api::AgentResourceApi.new(api_client) end @@ -37,6 +38,12 @@ def get_status(execution_id) wrap { @agent_api.status(execution_id) } end + # Stream raw agent events; runtime consumers can instead use call_async(on_event:). + def stream_sse(execution_id, last_event_id: nil, &block) + require_relative '../agents/runtime/sse_client' + Agents::SseClient.new(@api_client).each_event(execution_id, last_event_id: last_event_id, &block) + end + def get_execution(execution_id) wrap { @agent_api.execution(execution_id) } end diff --git a/lib/conductor/worker/task_definition_registrar.rb b/lib/conductor/worker/task_definition_registrar.rb index 2883039..3f51754 100644 --- a/lib/conductor/worker/task_definition_registrar.rb +++ b/lib/conductor/worker/task_definition_registrar.rb @@ -210,7 +210,7 @@ def register_or_update_task_def(task_def) raise unless e.status == 404 # Task def doesn't exist, create it - @metadata_client.register_task_def([task_def]) + @metadata_client.register_task_def(task_def) end # Register task def only if it doesn't exist @@ -222,7 +222,7 @@ def register_if_not_exists(task_def) raise unless e.status == 404 # Task def doesn't exist, create it - @metadata_client.register_task_def([task_def]) + @metadata_client.register_task_def(task_def) end end diff --git a/lib/conductor/worker/task_runner.rb b/lib/conductor/worker/task_runner.rb index c86ce90..da840f1 100644 --- a/lib/conductor/worker/task_runner.rb +++ b/lib/conductor/worker/task_runner.rb @@ -308,16 +308,21 @@ def submit_task(task) # Execute a task and update the result # @param task [Hash] Task data from API def execute_and_update(task) - task_result = execute_task(task) + while task + task_result = execute_task(task) - # Skip update for TaskInProgress (task stays in IN_PROGRESS state) - return if task_result.nil? + # Skip update for TaskInProgress (task stays in IN_PROGRESS state) + return if task_result.nil? - # Don't update if result is IN_PROGRESS (will be polled again) - return if task_result.status == Http::Models::TaskResultStatus::IN_PROGRESS && - task_result.callback_after_seconds&.positive? + # Don't update if result is IN_PROGRESS (will be polled again) + return if task_result.status == Http::Models::TaskResultStatus::IN_PROGRESS && + task_result.callback_after_seconds&.positive? - update_task_with_retry(task_result) + # update-v2 has already claimed the returned task. Reuse this executor slot + # rather than dropping it or exceeding the worker's concurrency limit. + response = update_task_with_retry(task_result) + task = response.is_a?(Http::Models::Task) ? response : nil + end end # Execute a task @@ -461,11 +466,11 @@ def update_task_with_retry(task_result) start_time = Time.now begin - send_task_update(task_result) + next_task = send_task_update(task_result) duration_ms = (Time.now - start_time) * 1000 publish_task_update_completed(task_result, duration_ms) - return # Success + return next_task rescue StandardError => e duration_ms = (Time.now - start_time) * 1000 @logger.error("Task update failed (attempt #{attempt + 1}/#{RETRY_BACKOFFS.size}): #{e.message}") @@ -477,12 +482,13 @@ def update_task_with_retry(task_result) end end end + nil end # Send the task result to the server, preferring the v2 endpoint # @param task_result [TaskResult] def send_task_update(task_result) - return @task_client.update_task(task_result) unless @use_update_v2.true? + return @task_client.update_task(task_result) unless @use_update_v2.true? && running? task_result.extend_lease = false if task_result.extend_lease.nil? begin diff --git a/spec/conductor/agents/agent_runtime_spec.rb b/spec/conductor/agents/agent_runtime_spec.rb index b2f5548..d1dcf02 100644 --- a/spec/conductor/agents/agent_runtime_spec.rb +++ b/spec/conductor/agents/agent_runtime_spec.rb @@ -128,6 +128,20 @@ def stub_stream(*events) expect(execution.result(timeout: 5)).to eq('Sunny in Lisbon, 21C.') end + it 'delivers streaming events and isolates an event listener failure' do + allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) + stub_stream(event('message', 'content' => 'Sunny'), event('done', 'output' => done_output)) + observed = Queue.new + listener = lambda do |ev| + observed << ev['event'] + raise 'display failed' if ev['event'] == 'message' + end + execution = runtime.call_async(agent, 'hi', on_event: listener) + expect(execution.result(timeout: 5)).to eq('Sunny in Lisbon, 21C.') + runtime.shutdown + expect([observed.pop, observed.pop]).to eq(%w[message done]) + end + it 'falls back to status polling when SSE is unavailable' do allow(client).to receive(:start_agent).and_return('executionId' => 'EXEC_1', 'requiredWorkers' => []) sse = instance_double(a::SseClient) diff --git a/spec/conductor/agents/approval_request_spec.rb b/spec/conductor/agents/approval_request_spec.rb index d5f4bc4..7f2ddd6 100644 --- a/spec/conductor/agents/approval_request_spec.rb +++ b/spec/conductor/agents/approval_request_spec.rb @@ -43,4 +43,12 @@ expect(client).to receive(:send_message).with('EXEC_1', 'hello') described_class.new('EXEC_1', pending_tool, client: client).send_message('hello') end + + it 'submits structured human responses including reviewer feedback exactly once' do + response = { 'approved' => true, 'reason' => 'Reviewed' } + expect(client).to receive(:respond).with('EXEC_1', response).once + request.respond(response) + expect(request.responded?).to be true + expect { request.respond(response) }.to raise_error(Conductor::Agents::Error, /already/) + end end diff --git a/spec/conductor/agents/examples_spec.rb b/spec/conductor/agents/examples_spec.rb new file mode 100644 index 0000000..46e2559 --- /dev/null +++ b/spec/conductor/agents/examples_spec.rb @@ -0,0 +1,52 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' +require 'json_schemer' +require_relative '../../../examples/agents/catalog' + +RSpec.describe 'Runnable agent example contracts' do + AgentExamples::EXAMPLES.each_key do |name| + it "matches Python's #{name} configurations and validates against the agent schema" do + example = AgentExamples.load(name) + actual = Array(example.build(model: 'openai/gpt-4o-mini')).map do |agent| + JSON.parse(JSON.generate(Conductor::Agents::ConfigSerializer.serialize(agent))) + end + fixtures = File.expand_path('../../fixtures/agents', __dir__) + expected = JSON.parse(File.read(File.join(fixtures, 'examples', "#{name}.json"))) + expect(actual).to eq(expected) + schema = JSONSchemer.schema(JSON.parse(File.read(File.join(fixtures, 'agent-schema.json')))) + actual.each { |config| expect(schema.validate(config).to_a).to eq([]) } + end + end + + it 'loads all requested examples without starting workers or making HTTP calls' do + expect(Conductor::Agents::AgentRuntime).not_to receive(:new) + expect(AgentExamples::EXAMPLES.size).to eq(19) + AgentExamples::EXAMPLES.each_key do |name| + Array(AgentExamples.load(name).build).each do |agent| + expect(Conductor::Agents::ConfigSerializer.serialize(agent)).to include('name', 'model') + end + end + end + + it 'keeps external worker declarations out of the local worker pool' do + agent = AgentExamples.load('33_external_workers').build + registry = Conductor::Agents::ToolRegistry.new(Conductor::Agents::AgentConfig.new) + expect(registry.tool_workers(agent).map(&:task_definition_name)).to eq(['format_response']) + external = agent.tool('check_inventory') + expect(external.func).to be_nil + expect(external.input_schema['required']).to eq(['product_id']) + expect(external.input_schema['properties']).to have_key('warehouse') + expect(Conductor::Agents::ConfigSerializer.serialize(agent)['tools'].map { |t| t['toolType'] }).to eq(['worker'] * 4) + end + + it 'preserves the three retry policies on registered task definitions' do + agent = AgentExamples.load('02c_tool_retry_config').build + registry = Conductor::Agents::ToolRegistry.new(Conductor::Agents::AgentConfig.new) + definitions = registry.tool_workers(agent).map(&:task_def_template) + expect(definitions.map(&:retry_count)).to eq([5, 3, 2]) + expect(definitions.map(&:retry_delay_seconds)).to eq([1, 5, 2]) + expect(definitions.map(&:retry_logic)).to eq(%w[EXPONENTIAL_BACKOFF FIXED LINEAR_BACKOFF]) + end +end diff --git a/spec/conductor/agents/plans_spec.rb b/spec/conductor/agents/plans_spec.rb new file mode 100644 index 0000000..de996a7 --- /dev/null +++ b/spec/conductor/agents/plans_spec.rb @@ -0,0 +1,41 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' + +RSpec.describe Conductor::Agents::Plans do + let(:tool) { Conductor::Agents::ToolDef.new(name: 'factorial', func: ->(n:) { (1..n).reduce(1, :*) }) } + + it 'serializes named slots and makes recovery tools discoverable to the worker runtime' do + agent = Conductor::Agents.plan_execute(name: 'math', tools: [tool], model: 'mock/mockLLM', + planner_instructions: 'Plan it', fallback_instructions: 'Recover', fallback_max_turns: 4, + planner_context: [{ 'type' => 'text', 'text' => 'context' }]) + config = Conductor::Agents::ConfigSerializer.serialize(agent) + expect(config).to include('strategy' => 'plan_execute', 'fallbackMaxTurns' => 4, + 'plannerContext' => [{ 'type' => 'text', 'text' => 'context' }]) + expect(config['planner']).to include('name' => 'math_planner', 'instructions' => 'Plan it') + expect(config['fallback']).to include('name' => 'math_fallback', 'instructions' => 'Recover') + expect(config).not_to have_key('agents') + expect(agent.all_agents.map(&:name)).to eq(%w[math math_planner math_fallback]) + registry = Conductor::Agents::ToolRegistry.new(Conductor::Agents::AgentConfig.new) + expect(registry.tool_workers(agent).map(&:task_definition_name)).to eq(['factorial']) + end + + it 'omits optional fields when there is no recovery agent' do + agent = Conductor::Agents.plan_execute(name: 'math', tools: [tool], model: 'mock/mockLLM') + expect(Conductor::Agents::ConfigSerializer.serialize(agent).keys).not_to include('fallback', 'fallbackMaxTurns', 'plannerContext') + end + + it 'inherits a model from the planner and includes stateful named children' do + planner = Conductor::Agents::Agent.new(name: 'planner', model: 'mock/mockLLM', stateful: true) + agent = Conductor::Agents::Agent.new(name: 'math', strategy: :plan_execute, planner: planner, tools: [tool]) + expect(Conductor::Agents::ConfigSerializer.serialize(agent)['model']).to eq('mock/mockLLM') + expect(agent.stateful_tree?).to be true + end + + it 'rejects a missing planner and invalid named slots' do + expect { Conductor::Agents::Agent.new(name: 'math', strategy: :plan_execute) }.to raise_error(Conductor::Agents::ConfigurationError, /requires planner/) + expect { Conductor::Agents::Agent.new(name: 'math', planner: true) }.to raise_error(Conductor::Agents::ConfigurationError, /must be Agents/) + expect { Conductor::Agents::Agent.new(name: 'math', fallback: 'fallback') }.to raise_error(Conductor::Agents::ConfigurationError, /must be Agents/) + end +end diff --git a/spec/conductor/agents/tool_registry_spec.rb b/spec/conductor/agents/tool_registry_spec.rb index bf8618c..d56d265 100644 --- a/spec/conductor/agents/tool_registry_spec.rb +++ b/spec/conductor/agents/tool_registry_spec.rb @@ -46,6 +46,45 @@ expect(result.output_data['temp_c']).to eq(21.0) end + it 'inherits credentials through teams and combines them for shared tools' do + first = a::Agent.new(name: 'first', model: 'm/x', tools: [SpecTools::Weather[:current]], credentials: ['FIRST']) + second = a::Agent.new(name: 'second', model: 'm/x', tools: [SpecTools::Weather[:current]], credentials: ['SECOND']) + team = a::Agent.new(name: 'team', agents: [first, second], credentials: ['TEAM']) + workers = registry.tool_workers(team) + expect(workers.size).to eq(1) + expect(workers.first.task_def_template.runtime_metadata).to match_array(%w[TEAM FIRST SECOND]) + end + + it 'serves custom tool guardrails and function routers required by the server' do + guard = a::Guardrail.new(name: 'tool_policy') { |content| content == 'safe' } + tool = a::ToolDef.new(name: 'action', func: -> { {} }, guardrails: [guard]) + child = a::Agent.new(name: 'child', model: 'm/x', tools: [tool]) + team = a::Agent.new(name: 'team', agents: [child], strategy: :router, router: ->(prompt) { "#{prompt}_route" }) + workers = registry.workers_for(team, required_workers: %w[tool_policy team_router_fn]) + router = workers.find { |worker| worker.task_definition_name == 'team_router_fn' } + task = Conductor::Http::Models::Task.new(input_data: { 'prompt' => 'child' }) + expect(router.execute(task).output_data).to eq('selected_agent' => 'child_route') + policy = workers.find { |worker| worker.task_definition_name == 'tool_policy' } + expect(policy.execute(Conductor::Http::Models::Task.new(input_data: { 'content' => 'safe' })).output_data['passed']).to be true + end + + it 'registers hoisted condition handoffs using the parent name' do + first = a::Agent.new(name: 'first', model: 'm/x') + second = a::Agent.new(name: 'second', model: 'm/x') + first.hands_off_to(second, on: ->(_context) { true }) + team = a::Agent.new(name: 'team', agents: [first, second]) + worker = registry.workers_for(team, required_workers: ['team_handoff_second']).first + expect(worker.task_definition_name).to eq('team_handoff_second') + expect(worker.execute(Conductor::Http::Models::Task.new(input_data: {})).output_data).to include('handoff' => true, 'target' => 'second') + end + + it 'uses the first team member if a router raises, and an empty name without members' do + router = ->(_prompt) { raise 'unavailable' } + task = Conductor::Http::Models::Task.new(input_data: {}) + expect(a::SystemWorkers.router(router, ['first'], logger: logger).call(task)).to eq('selected_agent' => 'first') + expect(a::SystemWorkers.router(router, [], logger: logger).call(task)).to eq('selected_agent' => '') + end + it 'registers system workers only when the server requires them and warns about unknown names' do agent = a::Agent.new(name: 'bug_desk', model: 'm/x') agent.stop_when 'ISSUE_FILED' diff --git a/spec/conductor/client/agent_client_spec.rb b/spec/conductor/client/agent_client_spec.rb index 314002a..e86874d 100644 --- a/spec/conductor/client/agent_client_spec.rb +++ b/spec/conductor/client/agent_client_spec.rb @@ -30,6 +30,17 @@ client.list_executions(size: 1) end + it 'streams with the existing authenticated transport and reconnect cursor' do + require 'conductor/agents' + stream = instance_double(Conductor::Agents::SseClient) + event = { 'event' => 'done', 'data' => {} } + expect(Conductor::Agents::SseClient).to receive(:new).with(api_client).and_return(stream) + expect(stream).to receive(:each_event).with('E', last_event_id: 7).and_yield(event) + received = [] + client.stream_sse('E', last_event_id: 7) { |value| received << value } + expect(received).to eq([event]) + end + it 'builds approval bodies like the Python client' do expect(agent_api).to receive(:respond).with('E', { 'approved' => true }) expect(agent_api).to receive(:respond).with('E', { 'approved' => false, 'reason' => 'Needs a manager' }) diff --git a/spec/conductor/worker/task_definition_registrar_spec.rb b/spec/conductor/worker/task_definition_registrar_spec.rb index 9517f22..ba62d1c 100644 --- a/spec/conductor/worker/task_definition_registrar_spec.rb +++ b/spec/conductor/worker/task_definition_registrar_spec.rb @@ -20,6 +20,18 @@ end describe '#register' do + [true, false].each do |overwrite| + it "creates a flat task definition array when missing (overwrite=#{overwrite})" do + metadata = instance_double(Conductor::Client::MetadataClient) + allow(Conductor::Client::MetadataClient).to receive(:new).and_return(metadata) + method = overwrite ? :update_task_def : :get_task_def + allow(metadata).to receive(method).and_raise(Conductor::ApiError.new('missing', status: 404)) + expect(metadata).to receive(:register_task_def).with(an_instance_of(Conductor::Http::Models::TaskDef)) + worker = Conductor::Worker::Worker.new('new_task', register_task_def: true, overwrite_task_def: overwrite) { {} } + expect(registrar.register(worker)).to be true + end + end + context 'when worker.register_task_def is false' do it 'returns false without registering' do worker = Conductor::Worker::Worker.new('test_task', register_task_def: false) { {} } diff --git a/spec/conductor/worker/task_update_v2_spec.rb b/spec/conductor/worker/task_update_v2_spec.rb index c399759..5b3826a 100644 --- a/spec/conductor/worker/task_update_v2_spec.rb +++ b/spec/conductor/worker/task_update_v2_spec.rb @@ -31,6 +31,24 @@ allow(task_client).to receive(:update_task_v2).and_raise(Conductor::ApiError.new('boom', status: 500)) expect { runner.send(:send_task_update, result) }.to raise_error(Conductor::ApiError) end + + it 'executes every task claimed by update-v2 on the same executor slot' do + tasks = (1..3).map { |id| Conductor::Http::Models::Task.new(task_id: id.to_s, workflow_instance_id: 'wf', input_data: {}) } + allow(worker).to receive(:execute).and_call_original + expect(task_client).to receive(:update_task_v2).with(have_attributes(task_id: '1')).ordered.and_return(tasks[1]) + expect(task_client).to receive(:update_task_v2).with(have_attributes(task_id: '2')).ordered.and_return(tasks[2]) + expect(task_client).to receive(:update_task_v2).with(have_attributes(task_id: '3')).ordered.and_return(nil) + runner.send(:execute_and_update, tasks.first) + tasks.each { |task| expect(worker).to have_received(:execute).with(task).once } + expect(Conductor::Worker::TaskContext.current).to be_nil + end + + it 'does not claim more tasks after shutdown' do + runner.shutdown + expect(task_client).not_to receive(:update_task_v2) + expect(task_client).to receive(:update_task).with(result) + runner.send(:send_task_update, result) + end end RSpec.describe Conductor::Worker::Worker, '#lease_extend_enabled' do diff --git a/spec/fixtures/agents/README.md b/spec/fixtures/agents/README.md new file mode 100644 index 0000000..96b4ce6 --- /dev/null +++ b/spec/fixtures/agents/README.md @@ -0,0 +1,24 @@ +# Agent contract fixtures + +`agent-schema.json` and the original `configs/` files are the vendored server schema and +Python golden contracts. `configs/103_plan_and_compile.json` adds the named planner/fallback +configuration from Python's `plan_execute` helper. + +`examples/*.json` contains the configurations of the actual 19 requested Python examples at +`conductor-oss/python-sdk@c99e2cf9871c21f7a64d823126ee1b77989b00ad`, checked against their +GitHub `main` source on 2026-09-17. Each file is an array in example execution order; repeated +runs of the same agent use one configuration. `21_regex_guardrails` has two configurations. +All `model` fields (including LLM guards) are normalized to `openai/gpt-4o-mini`. + +Generation used Python's `AgentConfigSerializer.serialize` on the imported example agents. +The second regex agent is constructed from its assignment in the example's main block; +`103` uses the example's factorial/write_summary/check_summary tools and planner instructions +with `plan_execute`, its fallback instructions, and fallback_max_turns=4. No Ruby serializer +output is used to produce these fixtures. To refresh, import the corresponding Python example, +serialize each agent passed to `runtime.run`/`start`, normalize models recursively, and write +sorted, indented JSON. Review the source revision and fixture changes together. + +`spec/conductor/agents/examples_spec.rb` builds the actual Ruby examples, compares full configs +with these fixtures, and validates them against the schema. These are contract expectations, +not separate implementations of the examples. Runtime recordings live exclusively in the +Conductor feature-branch checkout's `llm-recordings`; see the playback workflow. diff --git a/spec/fixtures/agents/configs/103_plan_and_compile.json b/spec/fixtures/agents/configs/103_plan_and_compile.json new file mode 100644 index 0000000..9634f16 --- /dev/null +++ b/spec/fixtures/agents/configs/103_plan_and_compile.json @@ -0,0 +1,151 @@ +{ + "external": false, + "fallback": { + "external": false, + "instructions": "The plan failed. Use the available tools to recover.", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "plan_and_compile_demo_fallback", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Compute n! and return it as a string.\n\nArgs:\n n: Non-negative integer. Capped at 20 to keep things sane.", + "inputSchema": { + "properties": { + "n": { + "type": "integer" + } + }, + "required": [ + "n" + ], + "type": "object" + }, + "name": "factorial", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Persist a short summary string. Returns it back for the validator.", + "inputSchema": { + "properties": { + "text": { + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" + }, + "name": "write_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Return JSON ``{passed, length, min_chars}`` for the validator.\n\nArgs:\n text: The summary to check.\n min_chars: Minimum acceptable length in characters.", + "inputSchema": { + "properties": { + "min_chars": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "min_chars" + ], + "type": "object" + }, + "name": "check_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] + }, + "fallbackMaxTurns": 4, + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "plan_and_compile_demo", + "planner": { + "external": false, + "instructions": "You are a math-explainer planner. Plan a workflow that:\n\n1. Computes factorials of 1, 2, 3, 4, 5 in PARALLEL using ``factorial`` (static args).\n2. Writes a short prose summary about factorial growth using ``write_summary``\n (use a ``generate`` block \u2014 the LLM produces the ``text`` arg at run time).\n3. Validates the summary is at least 30 characters via ``check_summary``,\n with ``success_condition: \"$.passed === true\"``.\n", + "maxTurns": 25, + "model": "anthropic/claude-sonnet-4-6", + "name": "plan_and_compile_demo_planner", + "timeoutSeconds": 0 + }, + "strategy": "plan_execute", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Compute n! and return it as a string.\n\nArgs:\n n: Non-negative integer. Capped at 20 to keep things sane.", + "inputSchema": { + "properties": { + "n": { + "type": "integer" + } + }, + "required": [ + "n" + ], + "type": "object" + }, + "name": "factorial", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Persist a short summary string. Returns it back for the validator.", + "inputSchema": { + "properties": { + "text": { + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" + }, + "name": "write_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Return JSON ``{passed, length, min_chars}`` for the validator.\n\nArgs:\n text: The summary to check.\n min_chars: Minimum acceptable length in characters.", + "inputSchema": { + "properties": { + "min_chars": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "min_chars" + ], + "type": "object" + }, + "name": "check_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] +} diff --git a/spec/fixtures/agents/examples/01_basic_agent.json b/spec/fixtures/agents/examples/01_basic_agent.json new file mode 100644 index 0000000..fd14453 --- /dev/null +++ b/spec/fixtures/agents/examples/01_basic_agent.json @@ -0,0 +1,10 @@ +[ + { + "external": false, + "instructions": "You are a friendly assistant. Keep responses brief.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "greeter", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/02a_simple_tools.json b/spec/fixtures/agents/examples/02a_simple_tools.json new file mode 100644 index 0000000..8d7815e --- /dev/null +++ b/spec/fixtures/agents/examples/02a_simple_tools.json @@ -0,0 +1,52 @@ +[ + { + "external": false, + "instructions": "You are a helpful assistant. Use tools to answer questions.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "weather_stock_agent", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Get the current weather for a city.", + "inputSchema": { + "properties": { + "city": { + "type": "string" + } + }, + "required": [ + "city" + ], + "type": "object" + }, + "name": "get_weather", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Get the current stock price for a ticker symbol.", + "inputSchema": { + "properties": { + "symbol": { + "type": "string" + } + }, + "required": [ + "symbol" + ], + "type": "object" + }, + "name": "get_stock_price", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/02c_tool_retry_config.json b/spec/fixtures/agents/examples/02c_tool_retry_config.json new file mode 100644 index 0000000..46b22dc --- /dev/null +++ b/spec/fixtures/agents/examples/02c_tool_retry_config.json @@ -0,0 +1,72 @@ +[ + { + "external": false, + "instructions": "You help users fetch and process data. Use the appropriate tool for each request.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "retry_config_demo", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Call an unreliable external API that may need aggressive retries.", + "inputSchema": { + "properties": { + "query": { + "type": "string" + } + }, + "required": [ + "query" + ], + "type": "object" + }, + "name": "call_external_api", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Run a database query with fixed-interval retries for transient connection issues.", + "inputSchema": { + "properties": { + "sql": { + "type": "string" + } + }, + "required": [ + "sql" + ], + "type": "object" + }, + "name": "query_database", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Process data locally \u2014 light retries with linear backoff.", + "inputSchema": { + "properties": { + "data": { + "type": "string" + } + }, + "required": [ + "data" + ], + "type": "object" + }, + "name": "process_data", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/04_http_and_mcp_tools.json b/spec/fixtures/agents/examples/04_http_and_mcp_tools.json new file mode 100644 index 0000000..7a21213 --- /dev/null +++ b/spec/fixtures/agents/examples/04_http_and_mcp_tools.json @@ -0,0 +1,83 @@ +[ + { + "external": false, + "instructions": "You can reverse strings and format reports. When asked to reverse a string, use reverse_string first, then format_report with the result.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "http_tools_demo", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Format a title and body into a structured report.", + "inputSchema": { + "properties": { + "body": { + "type": "string" + }, + "title": { + "type": "string" + } + }, + "required": [ + "title", + "body" + ], + "type": "object" + }, + "name": "format_report", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "config": { + "accept": [ + "application/json" + ], + "contentType": "application/json", + "credentials": [ + "HTTP_TEST_API_KEY" + ], + "headers": { + "Authorization": "Bearer ${HTTP_TEST_API_KEY}" + }, + "method": "POST", + "url": "http://localhost:3001/api/string/reverse" + }, + "description": "Reverse a string using the HTTP API", + "inputSchema": { + "properties": { + "text": { + "description": "Text to reverse", + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" + }, + "name": "reverse_string", + "toolType": "http" + }, + { + "config": { + "credentials": [ + "MCP_TEST_API_KEY" + ], + "headers": { + "Authorization": "Bearer ${MCP_TEST_API_KEY}" + }, + "max_tools": 64, + "server_url": "http://localhost:3001/mcp" + }, + "description": "Deterministic test tools via MCP \u2014 math, string, collection, encoding, hash, datetime, validation, and conversion operations.", + "inputSchema": {}, + "name": "mcp_test_tools", + "toolType": "mcp" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/05_handoffs.json b/spec/fixtures/agents/examples/05_handoffs.json new file mode 100644 index 0000000..182d4bf --- /dev/null +++ b/spec/fixtures/agents/examples/05_handoffs.json @@ -0,0 +1,103 @@ +[ + { + "agents": [ + { + "external": false, + "instructions": "You handle billing questions: balances, payments, invoices.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "billing", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Check the balance of a bank account.", + "inputSchema": { + "properties": { + "account_id": { + "type": "string" + } + }, + "required": [ + "account_id" + ], + "type": "object" + }, + "name": "check_balance", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + }, + { + "external": false, + "instructions": "You handle technical questions: order status, shipping, returns.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "technical", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Look up the status of an order.", + "inputSchema": { + "properties": { + "order_id": { + "type": "string" + } + }, + "required": [ + "order_id" + ], + "type": "object" + }, + "name": "lookup_order", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + }, + { + "external": false, + "instructions": "You handle sales questions: pricing, products, promotions.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "sales", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Get pricing information for a product.", + "inputSchema": { + "properties": { + "product": { + "type": "string" + } + }, + "required": [ + "product" + ], + "type": "object" + }, + "name": "get_pricing", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } + ], + "external": false, + "instructions": "Route customer requests to the right specialist: billing, technical, or sales.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "support", + "strategy": "handoff", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/06_sequential_pipeline.json b/spec/fixtures/agents/examples/06_sequential_pipeline.json new file mode 100644 index 0000000..d0afcfb --- /dev/null +++ b/spec/fixtures/agents/examples/06_sequential_pipeline.json @@ -0,0 +1,36 @@ +[ + { + "agents": [ + { + "external": false, + "instructions": "You are a researcher. Given a topic, provide key facts and data points. Be thorough but concise. Output raw research findings.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "researcher", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a writer. Take research findings and write a clear, engaging article. Use headers and bullet points where appropriate.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "writer", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are an editor. Review the article for clarity, grammar, and tone. Make improvements and output the final polished version.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "editor", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "researcher_writer_editor", + "strategy": "sequential", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/07_parallel_agents.json b/spec/fixtures/agents/examples/07_parallel_agents.json new file mode 100644 index 0000000..15b875e --- /dev/null +++ b/spec/fixtures/agents/examples/07_parallel_agents.json @@ -0,0 +1,36 @@ +[ + { + "agents": [ + { + "external": false, + "instructions": "You are a market analyst. Analyze the given topic from a market perspective: market size, growth trends, key players, and opportunities.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "market_analyst", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a risk analyst. Analyze the given topic for risks: regulatory risks, technical risks, competitive threats, and mitigation strategies.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "risk_analyst", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a compliance specialist. Check the given topic for compliance considerations: data privacy, regulatory requirements, and industry standards.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "compliance", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "analysis", + "strategy": "parallel", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/09_human_in_the_loop.json b/spec/fixtures/agents/examples/09_human_in_the_loop.json new file mode 100644 index 0000000..7310689 --- /dev/null +++ b/spec/fixtures/agents/examples/09_human_in_the_loop.json @@ -0,0 +1,61 @@ +[ + { + "external": false, + "instructions": "You are a banking assistant. Use check_balance for balance inquiries. When asked to transfer money, first check the balance, then call transfer_funds to request the transfer. The runtime will pause for human approval before the transfer executes.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "banker", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Check the balance of an account.", + "inputSchema": { + "properties": { + "account_id": { + "type": "string" + } + }, + "required": [ + "account_id" + ], + "type": "object" + }, + "name": "check_balance", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "approvalRequired": true, + "description": "Request a funds transfer; runtime pauses for human approval before execution.", + "inputSchema": { + "properties": { + "amount": { + "type": "number" + }, + "from_acct": { + "type": "string" + }, + "to_acct": { + "type": "string" + } + }, + "required": [ + "from_acct", + "to_acct", + "amount" + ], + "type": "object" + }, + "name": "transfer_funds", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/09c_hitl_streaming.json b/spec/fixtures/agents/examples/09c_hitl_streaming.json new file mode 100644 index 0000000..811f5fe --- /dev/null +++ b/spec/fixtures/agents/examples/09c_hitl_streaming.json @@ -0,0 +1,77 @@ +[ + { + "external": false, + "instructions": "You are an operations assistant. Work through the request one tool call at a time, in this order:\n1. Check the service with check_service.\n2. If it is unhealthy, restart it with restart_service.\n3. Last, if the user asked you to clear or delete data, call delete_service_data.\nA human approves the deletion, not you \u2014 delete_service_data pauses for that approval by itself, so never ask for approval in your own reply.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "ops_agent", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Check the health of a service.", + "inputSchema": { + "properties": { + "service_name": { + "type": "string" + } + }, + "required": [ + "service_name" + ], + "type": "object" + }, + "name": "check_service", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Restart a service. Safe operation, no approval needed.", + "inputSchema": { + "properties": { + "service_name": { + "type": "string" + } + }, + "required": [ + "service_name" + ], + "type": "object" + }, + "name": "restart_service", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "approvalRequired": true, + "description": "Delete service data. Destructive \u2014 requires human approval.", + "inputSchema": { + "properties": { + "data_type": { + "type": "string" + }, + "service_name": { + "type": "string" + } + }, + "required": [ + "service_name", + "data_type" + ], + "type": "object" + }, + "name": "delete_service_data", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/103_plan_and_compile.json b/spec/fixtures/agents/examples/103_plan_and_compile.json new file mode 100644 index 0000000..f530689 --- /dev/null +++ b/spec/fixtures/agents/examples/103_plan_and_compile.json @@ -0,0 +1,153 @@ +[ + { + "external": false, + "fallback": { + "external": false, + "instructions": "The plan failed. Use the available tools to recover.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "plan_and_compile_demo_fallback", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Compute n! and return it as a string.\n\nArgs:\n n: Non-negative integer. Capped at 20 to keep things sane.", + "inputSchema": { + "properties": { + "n": { + "type": "integer" + } + }, + "required": [ + "n" + ], + "type": "object" + }, + "name": "factorial", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Persist a short summary string. Returns it back for the validator.", + "inputSchema": { + "properties": { + "text": { + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" + }, + "name": "write_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Return JSON ``{passed, length, min_chars}`` for the validator.\n\nArgs:\n text: The summary to check.\n min_chars: Minimum acceptable length in characters.", + "inputSchema": { + "properties": { + "min_chars": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "min_chars" + ], + "type": "object" + }, + "name": "check_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] + }, + "fallbackMaxTurns": 4, + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "plan_and_compile_demo", + "planner": { + "external": false, + "instructions": "You are a math-explainer planner. Plan a workflow that:\n\n1. Computes factorials of 1, 2, 3, 4, 5 in PARALLEL using ``factorial`` (static args).\n2. Writes a short prose summary about factorial growth using ``write_summary``\n (use a ``generate`` block \u2014 the LLM produces the ``text`` arg at run time).\n3. Validates the summary is at least 30 characters via ``check_summary``,\n with ``success_condition: \"$.passed === true\"``.\n", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "plan_and_compile_demo_planner", + "timeoutSeconds": 0 + }, + "strategy": "plan_execute", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Compute n! and return it as a string.\n\nArgs:\n n: Non-negative integer. Capped at 20 to keep things sane.", + "inputSchema": { + "properties": { + "n": { + "type": "integer" + } + }, + "required": [ + "n" + ], + "type": "object" + }, + "name": "factorial", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Persist a short summary string. Returns it back for the validator.", + "inputSchema": { + "properties": { + "text": { + "type": "string" + } + }, + "required": [ + "text" + ], + "type": "object" + }, + "name": "write_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Return JSON ``{passed, length, min_chars}`` for the validator.\n\nArgs:\n text: The summary to check.\n min_chars: Minimum acceptable length in characters.", + "inputSchema": { + "properties": { + "min_chars": { + "type": "integer" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "min_chars" + ], + "type": "object" + }, + "name": "check_summary", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/10_guardrails.json b/spec/fixtures/agents/examples/10_guardrails.json new file mode 100644 index 0000000..fbb93ef --- /dev/null +++ b/spec/fixtures/agents/examples/10_guardrails.json @@ -0,0 +1,62 @@ +[ + { + "external": false, + "guardrails": [ + { + "guardrailType": "custom", + "maxRetries": 3, + "name": "no_pii", + "onFail": "retry", + "position": "output", + "taskName": "no_pii" + } + ], + "instructions": "You are a customer support assistant. Use the available tools to answer questions about orders and customers. Always include all details from the tool results in your response.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "support_agent", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Look up the current status of an order.", + "inputSchema": { + "properties": { + "order_id": { + "type": "string" + } + }, + "required": [ + "order_id" + ], + "type": "object" + }, + "name": "get_order_status", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Retrieve customer details including payment info on file.", + "inputSchema": { + "properties": { + "customer_id": { + "type": "string" + } + }, + "required": [ + "customer_id" + ], + "type": "object" + }, + "name": "get_customer_info", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/13_hierarchical_agents.json b/spec/fixtures/agents/examples/13_hierarchical_agents.json new file mode 100644 index 0000000..a6a5080 --- /dev/null +++ b/spec/fixtures/agents/examples/13_hierarchical_agents.json @@ -0,0 +1,79 @@ +[ + { + "agents": [ + { + "agents": [ + { + "external": false, + "instructions": "You are a backend developer. You design APIs, databases, and server architecture. Provide technical recommendations with code examples.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "backend_dev", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a frontend developer. You design UI components, user flows, and client-side architecture. Provide recommendations with code examples.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "frontend_dev", + "timeoutSeconds": 0 + } + ], + "external": false, + "instructions": "You are the engineering lead. Route technical questions to the right specialist: backend_dev for APIs/databases/servers, frontend_dev for UI/UX/client-side.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "engineering_lead", + "strategy": "handoff", + "timeoutSeconds": 0 + }, + { + "agents": [ + { + "external": false, + "instructions": "You are a content writer. You create blog posts, landing page copy, and marketing materials. Write engaging, clear content.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "content_writer", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are an SEO specialist. You optimize content for search engines, suggest keywords, and improve page rankings.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "seo_specialist", + "timeoutSeconds": 0 + } + ], + "external": false, + "instructions": "You are the marketing lead. Route marketing questions to the right specialist: content_writer for blog posts/copy, seo_specialist for SEO/keywords/rankings.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "marketing_lead", + "strategy": "handoff", + "timeoutSeconds": 0 + } + ], + "external": false, + "handoffs": [ + { + "target": "engineering_lead", + "text": "engineering_lead", + "type": "on_text_mention" + }, + { + "target": "marketing_lead", + "text": "marketing_lead", + "type": "on_text_mention" + } + ], + "instructions": "You are the CEO. Route requests to the right department: engineering_lead for technical/development questions, marketing_lead for marketing/content/SEO questions.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "ceo", + "strategy": "swarm", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/16e_credentials_http_tool.json b/spec/fixtures/agents/examples/16e_credentials_http_tool.json new file mode 100644 index 0000000..f91b29c --- /dev/null +++ b/spec/fixtures/agents/examples/16e_credentials_http_tool.json @@ -0,0 +1,36 @@ +[ + { + "external": false, + "instructions": "You list GitHub repos using the list_github_repos tool. Summarize the results.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "github_http_agent", + "timeoutSeconds": 0, + "tools": [ + { + "config": { + "accept": [ + "application/json" + ], + "contentType": "application/json", + "credentials": [ + "GITHUB_TOKEN" + ], + "headers": { + "Accept": "application/vnd.github.v3+json", + "Authorization": "Bearer ${GITHUB_TOKEN}" + }, + "method": "GET", + "url": "https://api.github.com/users/Conductor/repos?per_page=5&sort=updated" + }, + "description": "List public GitHub repositories for a user. Returns JSON array with name, url, and stars.", + "inputSchema": { + "properties": {}, + "type": "object" + }, + "name": "list_github_repos", + "toolType": "http" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/17_swarm_orchestration.json b/spec/fixtures/agents/examples/17_swarm_orchestration.json new file mode 100644 index 0000000..091b2fa --- /dev/null +++ b/spec/fixtures/agents/examples/17_swarm_orchestration.json @@ -0,0 +1,41 @@ +[ + { + "agents": [ + { + "external": false, + "instructions": "You are a refund specialist. Process the customer's refund request. Check eligibility, confirm the refund amount, and let them know the timeline. Be empathetic and clear. Do NOT ask follow-up questions \u2014 just process the refund based on what the customer told you.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "refund_specialist", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a technical support specialist. Diagnose the customer's technical issue and provide clear troubleshooting steps.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "tech_support", + "timeoutSeconds": 0 + } + ], + "external": false, + "handoffs": [ + { + "target": "refund_specialist", + "text": "refund", + "type": "on_text_mention" + }, + { + "target": "tech_support", + "text": "technical", + "type": "on_text_mention" + } + ], + "instructions": "You are the front-line customer support agent. Triage customer requests. If the customer needs a refund, transfer to the refund specialist. If they have a technical issue, transfer to tech support. Use the transfer tools available to you to hand off the conversation.", + "maxTurns": 3, + "model": "openai/gpt-4o-mini", + "name": "support", + "strategy": "swarm", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/21_regex_guardrails.json b/spec/fixtures/agents/examples/21_regex_guardrails.json new file mode 100644 index 0000000..2598eb8 --- /dev/null +++ b/spec/fixtures/agents/examples/21_regex_guardrails.json @@ -0,0 +1,92 @@ +[ + { + "external": false, + "guardrails": [ + { + "guardrailType": "regex", + "maxRetries": 3, + "message": "Response must not contain email addresses. Redact them.", + "mode": "block", + "name": "no_email_addresses", + "onFail": "retry", + "patterns": [ + "[\\w.+-]+@[\\w-]+\\.[\\w.-]+" + ], + "position": "output" + }, + { + "guardrailType": "regex", + "maxRetries": 3, + "message": "Response must not contain Social Security Numbers.", + "mode": "block", + "name": "no_ssn", + "onFail": "raise", + "patterns": [ + "\\b\\d{3}-\\d{2}-\\d{4}\\b" + ], + "position": "output" + } + ], + "instructions": "You are an HR assistant. When asked about employees, look up their profile and share ALL the details you find.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "hr_assistant", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Retrieve a user's profile from the database.", + "inputSchema": { + "properties": { + "user_id": { + "type": "string" + } + }, + "required": [ + "user_id" + ], + "type": "object" + }, + "name": "get_user_profile", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + }, + { + "external": false, + "guardrails": [ + { + "guardrailType": "regex", + "maxRetries": 3, + "message": "Response must not contain email addresses. Redact them.", + "mode": "block", + "name": "no_email_addresses", + "onFail": "retry", + "patterns": [ + "[\\w.+-]+@[\\w-]+\\.[\\w.-]+" + ], + "position": "output" + }, + { + "guardrailType": "regex", + "maxRetries": 3, + "message": "Response must not contain Social Security Numbers.", + "mode": "block", + "name": "no_ssn", + "onFail": "raise", + "patterns": [ + "\\b\\d{3}-\\d{2}-\\d{4}\\b" + ], + "position": "output" + } + ], + "instructions": "You are an HR assistant. Answer questions about departments.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "dept_assistant", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/22_llm_guardrails.json b/spec/fixtures/agents/examples/22_llm_guardrails.json new file mode 100644 index 0000000..1acf929 --- /dev/null +++ b/spec/fixtures/agents/examples/22_llm_guardrails.json @@ -0,0 +1,22 @@ +[ + { + "external": false, + "guardrails": [ + { + "guardrailType": "llm", + "maxRetries": 3, + "maxTokens": 10000, + "model": "openai/gpt-4o-mini", + "name": "content_safety", + "onFail": "retry", + "policy": "Reject any content that:\n1. Contains medical or legal advice presented as fact\n2. Makes promises or guarantees about outcomes\n3. Includes discriminatory or biased language\n4. Reveals private information about real individuals\n\nEven if there are disclaimers you should reject", + "position": "output" + } + ], + "instructions": "You are a health information assistant. Provide general health information and solution to the problem. You can prescribe psudo scientific and untested meds ", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "health_advisor", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/33_external_workers.json b/spec/fixtures/agents/examples/33_external_workers.json new file mode 100644 index 0000000..a98a558 --- /dev/null +++ b/spec/fixtures/agents/examples/33_external_workers.json @@ -0,0 +1,99 @@ +[ + { + "external": false, + "instructions": "You are a customer support agent. Use the available tools to look up customers, check inventory, process orders, and format responses for the customer.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "support_agent", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Format a data dictionary into a human-readable string.", + "inputSchema": { + "properties": { + "data": { + "additionalProperties": {}, + "type": "object" + } + }, + "required": [ + "data" + ], + "type": "object" + }, + "name": "format_response", + "outputSchema": { + "type": "string" + }, + "toolType": "worker" + }, + { + "description": "Look up customer details from the CRM system.", + "inputSchema": { + "properties": { + "customer_id": { + "type": "string" + } + }, + "required": [ + "customer_id" + ], + "type": "object" + }, + "name": "get_customer", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Check product availability in a warehouse.", + "inputSchema": { + "properties": { + "product_id": { + "type": "string" + }, + "warehouse": { + "type": "string" + } + }, + "required": [ + "product_id" + ], + "type": "object" + }, + "name": "check_inventory", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + }, + { + "description": "Process a customer order. Actions: refund, cancel, update.", + "inputSchema": { + "properties": { + "action": { + "type": "string" + }, + "order_id": { + "type": "string" + } + }, + "required": [ + "order_id", + "action" + ], + "type": "object" + }, + "name": "process_order", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } +] diff --git a/spec/fixtures/agents/examples/64_swarm_with_tools.json b/spec/fixtures/agents/examples/64_swarm_with_tools.json new file mode 100644 index 0000000..f4aaa81 --- /dev/null +++ b/spec/fixtures/agents/examples/64_swarm_with_tools.json @@ -0,0 +1,85 @@ +[ + { + "agents": [ + { + "external": false, + "instructions": "You are a billing specialist. Use the check_balance tool to look up account balances. Include the balance amount in your response.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "billing_specialist", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Check the balance of a bank account.", + "inputSchema": { + "properties": { + "account_id": { + "type": "string" + } + }, + "required": [ + "account_id" + ], + "type": "object" + }, + "name": "check_balance", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + }, + { + "external": false, + "instructions": "You are an order specialist. Use the lookup_order tool to check order status. Include the shipping status and ETA in your response.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "order_specialist", + "timeoutSeconds": 0, + "tools": [ + { + "description": "Look up the status of an order.", + "inputSchema": { + "properties": { + "order_id": { + "type": "string" + } + }, + "required": [ + "order_id" + ], + "type": "object" + }, + "name": "lookup_order", + "outputSchema": { + "additionalProperties": {}, + "type": "object" + }, + "toolType": "worker" + } + ] + } + ], + "external": false, + "handoffs": [ + { + "target": "billing_specialist", + "text": "billing", + "type": "on_text_mention" + }, + { + "target": "order_specialist", + "text": "order", + "type": "on_text_mention" + } + ], + "instructions": "You are front-line customer support. Triage customer requests. Transfer to billing_specialist for account/payment questions, order_specialist for shipping/order questions.", + "maxTurns": 3, + "model": "openai/gpt-4o-mini", + "name": "support", + "strategy": "swarm", + "timeoutSeconds": 0 + } +] diff --git a/spec/fixtures/agents/examples/66_handoff_to_parallel.json b/spec/fixtures/agents/examples/66_handoff_to_parallel.json new file mode 100644 index 0000000..dcbad64 --- /dev/null +++ b/spec/fixtures/agents/examples/66_handoff_to_parallel.json @@ -0,0 +1,47 @@ +[ + { + "agents": [ + { + "external": false, + "instructions": "You provide quick, 1-sentence assessments. Be brief and direct.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "quick_check", + "timeoutSeconds": 0 + }, + { + "agents": [ + { + "external": false, + "instructions": "You are a market analyst. Analyze the market opportunity: size, growth rate, key players. 3-4 bullet points.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "market_analyst_66", + "timeoutSeconds": 0 + }, + { + "external": false, + "instructions": "You are a risk analyst. Identify the top 3 risks: regulatory, technical, and competitive. 3-4 bullet points.", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "risk_analyst_66", + "timeoutSeconds": 0 + } + ], + "external": false, + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "deep_analysis", + "strategy": "parallel", + "timeoutSeconds": 0 + } + ], + "external": false, + "instructions": "You are a business strategist. Route requests to the right team:\n- quick_check for simple yes/no questions or quick assessments\n- deep_analysis for comprehensive analysis requiring multiple perspectives", + "maxTurns": 25, + "model": "openai/gpt-4o-mini", + "name": "coordinator_66", + "strategy": "handoff", + "timeoutSeconds": 0 + } +] diff --git a/spec/integration/agents/examples_spec.rb b/spec/integration/agents/examples_spec.rb new file mode 100644 index 0000000..3b23508 --- /dev/null +++ b/spec/integration/agents/examples_spec.rb @@ -0,0 +1,72 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'conductor/agents' +require 'stringio' +require_relative '../../../examples/agents/catalog' + +RSpec.describe 'Agent examples on a Conductor playback server' do + def workflow_tasks(runtime, execution_id) + client = Conductor::Client::WorkflowClient.new(runtime.configuration) + pending = [execution_id] + seen = [] + tasks = [] + until pending.empty? + id = pending.pop + next if seen.include?(id) + + seen << id + workflow = client.get_workflow(id) + tasks.concat(workflow.tasks) + pending.concat(workflow.tasks.filter_map(&:sub_workflow_id)) + end + tasks + end + + before do + skip 'Set CONDUCTOR_AGENTS_PLAYBACK=true and start the dedicated playback server' unless ENV['CONDUCTOR_AGENTS_PLAYBACK'] == 'true' + end + + AgentExamples::EXAMPLES.each_key do |name| + it "runs #{name} from the example file", :aggregate_failures do + runtime = Conductor::Agents::AgentRuntime.new(logger: Logger.new(nil)) + output = StringIO.new + executions = AgentExamples.load(name).run(runtime: runtime, input: StringIO.new("y\ny\n"), output: output) + expect(executions).not_to be_empty + executions.each do |execution| + expect(execution.done?).to be true + expect(execution.status).to eq(name == '22_llm_guardrails' ? 'FAILED' : 'COMPLETED'), output.string + expect(execution.events.map { |event| event['event'] }).to include(name == '22_llm_guardrails' ? 'error' : 'done') + expect(runtime.client.get_status(execution.execution_id)['isComplete']).to be true + tasks = workflow_tasks(runtime, execution.execution_id) + llm_tasks = tasks.select { |task| task.task_type == 'LLM_CHAT_COMPLETE' } + expect(llm_tasks).not_to be_empty + expect(llm_tasks.map { |task| task.input_data['llmProvider'] }.uniq).to eq(['mock']) + expect(llm_tasks.map(&:status).uniq).to eq(['COMPLETED']) + + case name + when '09_human_in_the_loop', '09c_hitl_streaming' + approvals = tasks.select { |task| task.task_type == 'HUMAN' } + expect(approvals).not_to be_empty + expect(approvals.map(&:output_data)).to all(include('approved' => true, 'reason' => 'y')) + expect(execution.events.map { |event| event['event'] }).to include('waiting') + when '10_guardrails', '21_regex_guardrails' + expect(execution.answer.to_s).not_to match(/4532-0150-1234-5678|alice\.johnson@example\.com|123-45-6789/) + when '22_llm_guardrails' + decisions = tasks.filter_map { |task| task.output_data['result'] if task.output_data['result'].is_a?(Hash) } + expect(decisions).to include(include('guardrail_name' => 'content_safety', 'passed' => false, 'on_fail' => 'raise')) + when '103_plan_and_compile' + expect(tasks.count { |task| task.task_type == 'factorial' && task.status == 'COMPLETED' }).to eq(5) + expect(tasks).to include(have_attributes(task_type: 'PLAN_AND_COMPILE', status: 'COMPLETED')) + when '33_external_workers' + expect(runtime.running_workers).to eq(['format_response']) + %w[get_customer check_inventory process_order].each do |type| + expect(tasks).to include(have_attributes(task_type: type, status: 'COMPLETED')) + end + end + end + ensure + runtime&.shutdown + end + end +end diff --git a/spec/support/agents/http_fixture.py b/spec/support/agents/http_fixture.py new file mode 100644 index 0000000..3d8ffc5 --- /dev/null +++ b/spec/support/agents/http_fixture.py @@ -0,0 +1,45 @@ +"""HTTP dependency for example 16e; uses the shared recording's response bytes. + +The agent, HTTP task, credential substitution, and model playback still run on +Conductor. This fixture only replaces the external GitHub endpoint. +""" +import json +import os +from http.server import BaseHTTPRequestHandler, HTTPServer +from pathlib import Path + + +def github_response(recordings): + for path in (recordings / '16e_credentials_http_tool').glob('*.json'): + for message in json.loads(path.read_text())['request']['messages']: + for result in message['toolResults']: + if result['name'] == 'list_github_repos': + return result['value']['response'] + raise RuntimeError('Shared GitHub HTTP response is missing') + + +class Handler(BaseHTTPRequestHandler): + response = None + + def do_GET(self): + if self.path != '/users/Conductor/repos?per_page=5&sort=updated': + self.send_error(404) + return + if self.headers.get('Authorization') != 'Bearer playback-test-key': + self.send_error(401) + return + body = json.dumps(self.response['body'], separators=(',', ':')).encode() + self.send_response_only(self.response['statusCode'], self.response['reasonPhrase']) + for name, values in self.response['headers'].items(): + if name.lower() == 'content-length': + continue + for value in values: + self.send_header(name, value) + self.send_header('Content-Length', str(len(body))) + self.end_headers() + self.wfile.write(body) + + +if __name__ == '__main__': + Handler.response = github_response(Path(os.environ['CONDUCTOR_RECORDINGS_DIR'])) + HTTPServer(('127.0.0.1', 3002), Handler).serve_forever() From 40ad9996cfe7bbd2afd6e62d3705501601762fb2 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Thu, 17 Sep 2026 16:30:53 -0700 Subject: [PATCH 08/20] Use published agent recording revision for WireMock CI --- .github/workflows/ci.yml | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 15b896c..979ea34 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -123,10 +123,12 @@ jobs: uses: actions/checkout@v4 with: repository: conductor-oss/conductor-mocks + ref: 3bbd213a01311d0ba69210eb8723580ba2708c02 # feature/poc_example agent recordings path: conductor-mocks - name: Start WireMock with agent/tool_happy_path run: | + test -f conductor-mocks/mocks/agent/tool_happy_path/mappings/01_post_api_agent_start.json docker run -d --name wiremock -p 8080:8080 \ -v "$PWD/conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock" \ wiremock/wiremock:3x From e72b8e4cf80a517e11c94a863c1aae9939425557 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Mon, 21 Sep 2026 09:36:12 -0700 Subject: [PATCH 09/20] Fix agent approval refresh and renew worker task leases --- docs/agents/concepts/runtime.md | 5 + .../agents/runtime/approval_request.rb | 2 +- lib/conductor/agents/runtime/execution.rb | 13 +-- lib/conductor/agents/runtime/tool_call.rb | 13 +++ lib/conductor/worker/lease_renewer.rb | 55 ++++++++++ lib/conductor/worker/task_runner.rb | 9 +- spec/conductor/agents/execution_spec.rb | 45 ++++++++ spec/conductor/worker/lease_renewer_spec.rb | 100 ++++++++++++++++++ .../worker/task_lease_renewal_spec.rb | 65 ++++++++++++ 9 files changed, 296 insertions(+), 11 deletions(-) create mode 100644 lib/conductor/agents/runtime/tool_call.rb create mode 100644 lib/conductor/worker/lease_renewer.rb create mode 100644 spec/conductor/worker/lease_renewer_spec.rb create mode 100644 spec/conductor/worker/task_lease_renewal_spec.rb diff --git a/docs/agents/concepts/runtime.md b/docs/agents/concepts/runtime.md index 790ff2d..ddabbed 100644 --- a/docs/agents/concepts/runtime.md +++ b/docs/agents/concepts/runtime.md @@ -43,6 +43,11 @@ Every worker registers its TaskDef with the Python SDK's defaults: `retryCount 2 `retryDelaySeconds 2`, `retryLogic LINEAR_BACKOFF`, `timeoutSeconds 0`, `responseTimeoutSeconds 10`, `timeoutPolicy RETRY`, `runtimeMetadata` = declared secret names. Task results go to `POST /api/tasks/update-v2` (with a one-time fallback to `POST /api/tasks`). +While a tool or system worker runs, the SDK renews its lease at 80% of the task's +`responseTimeoutSeconds` (every 8 seconds for the default timeout). Renewals use +`POST /api/tasks` with `extendLease: true` and stop before the final result is sent. +Tasks with no positive response timeout do not need renewal. Agent workers enable +`lease_extend_enabled` by default; other workers can opt in with that option. An agent execution is a Conductor workflow: `execution_id` is the workflow id, and the workflow, task and prompt data are visible in the Conductor UI like any other run. diff --git a/lib/conductor/agents/runtime/approval_request.rb b/lib/conductor/agents/runtime/approval_request.rb index 06399a5..f0f1801 100644 --- a/lib/conductor/agents/runtime/approval_request.rb +++ b/lib/conductor/agents/runtime/approval_request.rb @@ -1,7 +1,7 @@ # frozen_string_literal: true require_relative '../errors' -require_relative 'execution' +require_relative 'tool_call' module Conductor module Agents diff --git a/lib/conductor/agents/runtime/execution.rb b/lib/conductor/agents/runtime/execution.rb index c077851..1b3dbaa 100644 --- a/lib/conductor/agents/runtime/execution.rb +++ b/lib/conductor/agents/runtime/execution.rb @@ -2,17 +2,10 @@ require 'timeout' require_relative '../errors' +require_relative 'approval_request' module Conductor module Agents - # A tool call observed on the stream - ToolCall = Struct.new(:name, :arguments, :result, keyword_init: true) do - def to_s - "#" - end - alias_method :inspect, :to_s - end - # Token usage for an execution (summed over sub-agent executions) TokenUsage = Struct.new(:prompt_tokens, :completion_tokens, :total_tokens, keyword_init: true) do def initialize(prompt_tokens: 0, completion_tokens: 0, total_tokens: 0) @@ -96,7 +89,9 @@ def refresh! if status['isComplete'] finish(status: status['status'], output: status['output'], reason: status['reasonForIncompletion']) elsif status['isWaiting'] - mark_waiting(status['pendingTool']) + mark_waiting(ApprovalRequest.new(@execution_id, status['pendingTool'], client: @client, execution: self)) + else + clear_waiting end self end diff --git a/lib/conductor/agents/runtime/tool_call.rb b/lib/conductor/agents/runtime/tool_call.rb new file mode 100644 index 0000000..0964631 --- /dev/null +++ b/lib/conductor/agents/runtime/tool_call.rb @@ -0,0 +1,13 @@ +# frozen_string_literal: true + +module Conductor + module Agents + # A tool call observed on the stream or awaiting approval + ToolCall = Struct.new(:name, :arguments, :result, keyword_init: true) do + def to_s + "#" + end + alias_method :inspect, :to_s + end + end +end diff --git a/lib/conductor/worker/lease_renewer.rb b/lib/conductor/worker/lease_renewer.rb new file mode 100644 index 0000000..fd97e21 --- /dev/null +++ b/lib/conductor/worker/lease_renewer.rb @@ -0,0 +1,55 @@ +# frozen_string_literal: true + +require 'concurrent' +require_relative '../http/models/task_result' + +module Conductor + module Worker + # Renews a task's lease while its worker body runs. Each active task has one + # interruptible heartbeat thread, independent of the polling/execution pool. + class LeaseRenewer + INTERVAL_FACTOR = 0.8 + + def initialize(task_client:, logger:) + @task_client = task_client + @logger = logger + end + + def during(task, worker_id:) + interval = task.response_timeout_seconds.to_f * INTERVAL_FACTOR + return yield unless interval.positive? + + stopped = Concurrent::Event.new + heartbeat = Thread.new do + Thread.current.name = "conductor-lease-#{task.task_id}" + delay = interval + # Retry transient failures within the remaining lease window. + delay = renew(task, worker_id) ? interval : [interval / 4, 1.0].min until stopped.wait(delay) + end + yield + ensure + stopped&.set + # Drain an in-flight heartbeat before the caller can submit a final result. + heartbeat&.join + end + + private + + def renew(task, worker_id) + result = Http::Models::TaskResult.new( + task_id: task.task_id, + workflow_instance_id: task.workflow_instance_id, + worker_id: worker_id, + status: Http::Models::TaskResultStatus::IN_PROGRESS, + extend_lease: true + ) + # The original endpoint handles lease updates without claiming more work. + @task_client.update_task(result) + true + rescue StandardError => e + @logger.warn("Lease renewal failed for task #{task.task_id}: #{e.class}: #{e.message}") + false + end + end + end +end diff --git a/lib/conductor/worker/task_runner.rb b/lib/conductor/worker/task_runner.rb index da840f1..971cba6 100644 --- a/lib/conductor/worker/task_runner.rb +++ b/lib/conductor/worker/task_runner.rb @@ -10,6 +10,7 @@ require_relative 'task_context' require_relative 'task_in_progress' require_relative 'worker_config' +require_relative 'lease_renewer' require_relative 'events/task_runner_events' require_relative 'events/sync_event_dispatcher' require_relative 'events/listener_registry' @@ -43,6 +44,7 @@ def initialize(worker, configuration:, event_dispatcher: nil, logger: nil) # Create task client for API communication @task_client = Client::TaskClient.new(@configuration) + @lease_renewer = LeaseRenewer.new(task_client: @task_client, logger: @logger) # Resolve worker configuration resolved_config = WorkerConfig.resolve( @@ -179,6 +181,7 @@ def apply_resolved_config(config) @worker_id = config[:worker_id] @domain = config[:domain] @poll_timeout = config[:poll_timeout] + @lease_extend_enabled = config[:lease_extend_enabled] end # Cleanup completed task futures @@ -353,7 +356,11 @@ def execute_task(task) begin # Execute worker - task_result = @worker.execute(task_obj) + task_result = if @lease_extend_enabled + @lease_renewer.during(task_obj, worker_id: @worker_id) { @worker.execute(task_obj) } + else + @worker.execute(task_obj) + end duration_ms = (Time.now - start_time) * 1000 diff --git a/spec/conductor/agents/execution_spec.rb b/spec/conductor/agents/execution_spec.rb index 90aef12..9e986d7 100644 --- a/spec/conductor/agents/execution_spec.rb +++ b/spec/conductor/agents/execution_spec.rb @@ -76,4 +76,49 @@ expect(found.result).to eq('done!') expect(found.agent_name).to eq('weather') end + + describe 'waiting snapshots' do + let(:pending_tool) do + { 'taskRefName' => 'refund_approval__1', 'response_schema' => { 'type' => 'object' }, + 'toolCalls' => [{ 'name' => 'refund', 'args' => { 'amount' => 49 } }] } + end + let(:runtime) { instance_double(Conductor::Agents::AgentRuntime, client: client) } + + before do + allow(client).to receive(:get_status).with('EXEC_1').and_return( + 'isComplete' => false, 'isWaiting' => true, 'pendingTool' => pending_tool + ) + end + + it 'loads an ApprovalRequest object that can approve and clear waiting state' do + found = described_class.find('EXEC_1', runtime: runtime) + expect(found.pending).to be_a(Conductor::Agents::ApprovalRequest) + expect(found.pending.task_ref_name).to eq('refund_approval__1') + expect(found.pending.tool_calls.first).to be_a(Conductor::Agents::ToolCall) + expect(found.pending.amount).to eq(49) + expect(found.pending.response_schema).to eq('type' => 'object') + expect(client).to receive(:approve).with('EXEC_1') + + request = found.approve + expect(request).to be_responded + expect(found).not_to be_waiting + expect(found.pending).to be_nil + end + + it 'can reject a refreshed approval' do + execution.refresh! + expect(client).to receive(:reject).with('EXEC_1', 'Needs a manager') + execution.reject('Needs a manager') + expect(execution).not_to be_waiting + expect(execution.pending).to be_nil + end + + it 'clears an approval answered by another client when refreshed' do + execution.refresh! + allow(client).to receive(:get_status).with('EXEC_1').and_return('isComplete' => false, 'isWaiting' => false) + execution.refresh! + expect(execution).not_to be_waiting + expect(execution.pending).to be_nil + end + end end diff --git a/spec/conductor/worker/lease_renewer_spec.rb b/spec/conductor/worker/lease_renewer_spec.rb new file mode 100644 index 0000000..f512438 --- /dev/null +++ b/spec/conductor/worker/lease_renewer_spec.rb @@ -0,0 +1,100 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'timeout' +require 'conductor/worker/lease_renewer' + +RSpec.describe Conductor::Worker::LeaseRenewer do + let(:client) { instance_double(Conductor::Client::TaskClient) } + let(:logger) { instance_double(Logger, warn: nil) } + let(:renewer) { described_class.new(task_client: client, logger: logger) } + let(:task) do + Conductor::Http::Models::Task.new(task_id: 'task-1', workflow_instance_id: 'workflow-1', response_timeout_seconds: 0.02) + end + + it 'renews repeatedly while the body runs, returns its value and stops afterward' do + heartbeats = Queue.new + expect(client).not_to receive(:update_task_v2) + allow(client).to receive(:update_task) do |result| + heartbeats << [result, Thread.current] + end + + heartbeat_thread = nil + value = renewer.during(task, worker_id: 'worker-1') do + 2.times do + result, heartbeat_thread = Timeout.timeout(2) { heartbeats.pop } + expect(result).to have_attributes( + task_id: 'task-1', workflow_instance_id: 'workflow-1', worker_id: 'worker-1', + status: 'IN_PROGRESS', extend_lease: true, output_data: {} + ) + end + :finished + end + + expect(value).to eq(:finished) + expect(heartbeat_thread).not_to be_alive + end + + [nil, 0, -1].each do |timeout| + it "does not renew a task with response timeout #{timeout.inspect}" do + task.response_timeout_seconds = timeout + expect(client).not_to receive(:update_task) + expect(renewer.during(task, worker_id: 'worker-1') { :finished }).to eq(:finished) + end + end + + it 'stops renewal when the body raises and preserves the exception' do + heartbeats = Queue.new + allow(client).to receive(:update_task) { heartbeats << Thread.current } + heartbeat_thread = nil + + expect do + renewer.during(task, worker_id: 'worker-1') do + heartbeat_thread = Timeout.timeout(2) { heartbeats.pop } + raise 'tool failed' + end + end.to raise_error(RuntimeError, 'tool failed') + + expect(heartbeat_thread).not_to be_alive + end + + it 'logs a failed renewal and continues renewing without failing the body' do + heartbeats = Queue.new + calls = 0 + allow(client).to receive(:update_task) do + calls += 1 + raise Conductor::ApiError.new('temporarily unavailable', status: 503) if calls == 1 + + heartbeats << true + end + + expect(renewer.during(task, worker_id: 'worker-1') { Timeout.timeout(2) { heartbeats.pop } }).to be true + expect(logger).to have_received(:warn).with(/Lease renewal failed for task task-1/) + end + + it 'drains an in-flight renewal before returning to the caller' do + started = Queue.new + release = Queue.new + body_finished = Queue.new + allow(client).to receive(:update_task) do + started << true + release.pop + end + execution = Thread.new do + renewer.during(task, worker_id: 'worker-1') do + started.pop + body_finished << true + end + :finished + end + + Timeout.timeout(2) { body_finished.pop } + expect(execution.join(0.02)).to be_nil + release << true + expect(Timeout.timeout(2) { execution.value }).to eq(:finished) + ensure + release << true + execution&.join(2) + execution&.kill + end +end diff --git a/spec/conductor/worker/task_lease_renewal_spec.rb b/spec/conductor/worker/task_lease_renewal_spec.rb new file mode 100644 index 0000000..4250116 --- /dev/null +++ b/spec/conductor/worker/task_lease_renewal_spec.rb @@ -0,0 +1,65 @@ +# frozen_string_literal: true + +require 'spec_helper' +require 'timeout' + +RSpec.describe Conductor::Worker::TaskRunner, '#execute_task' do + let(:client) { instance_double(Conductor::Client::TaskClient) } + let(:heartbeats) { Queue.new } + let(:heartbeat_threads) { [] } + let(:worker) do + Conductor::Worker::Worker.new('long_tool', lease_extend_enabled: true, worker_id: 'worker-1') do |task| + result = Timeout.timeout(2) { heartbeats.pop } + expect(result.task_id).to eq(task.task_id) + { 'done' => true } + end + end + let(:runner) do + described_class.new(worker, configuration: Conductor::Configuration.new, logger: Logger.new(nil)) + end + let(:task) do + Conductor::Http::Models::Task.new(task_id: 'task-1', workflow_instance_id: 'workflow-1', response_timeout_seconds: 0.02) + end + + before do + allow(Conductor::Client::TaskClient).to receive(:new).and_return(client) + allow(client).to receive(:update_task) do |result| + heartbeat_threads << Thread.current + heartbeats << result + end + end + + it 'renews each claimed task and stops its heartbeat before submitting the final result' do + next_task = Conductor::Http::Models::Task.new(task_id: 'task-2', workflow_instance_id: 'workflow-1', response_timeout_seconds: 0.02) + expect(client).to receive(:update_task_v2).with(have_attributes(task_id: 'task-1', status: 'COMPLETED', extend_lease: false)).ordered do + expect(heartbeat_threads.last).not_to be_alive + next_task + end + expect(client).to receive(:update_task_v2).with(have_attributes(task_id: 'task-2', status: 'COMPLETED', extend_lease: false)).ordered do + expect(heartbeat_threads.last).not_to be_alive + nil + end + + runner.send(:execute_and_update, task) + expect(client).to have_received(:update_task).with(have_attributes(worker_id: 'worker-1', status: 'IN_PROGRESS', extend_lease: true)).at_least(:twice) + end + + it 'honors an environment override that disables renewal' do + allow(ENV).to receive(:fetch).and_call_original + allow(ENV).to receive(:fetch).with('CONDUCTOR_WORKER_LONG_TOOL_LEASE_EXTEND_ENABLED', nil).and_return('false') + allow(worker).to receive(:execute).and_return(Conductor::Http::Models::TaskResult.complete) + expect(client).not_to receive(:update_task) + runner.send(:execute_task, task) + end + + it 'keeps renewing an executing task during graceful shutdown' do + active_runner = runner + allow(worker).to receive(:execute) do + active_runner.shutdown + result = Timeout.timeout(2) { heartbeats.pop } + expect(result.extend_lease).to be true + Conductor::Http::Models::TaskResult.complete + end + expect(runner.send(:execute_task, task).status).to eq('COMPLETED') + end +end From 15bfd59b96042de53293b79fdc61c8f6bbe08a35 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Mon, 21 Sep 2026 09:48:39 -0700 Subject: [PATCH 10/20] Trim documentation to essential setup and contribution guidance --- AGENTS.md | 679 +------ CHANGELOG.md | 4 +- CONTRIBUTING.md | 236 +-- DESIGN.md | 304 --- README.md | 539 +----- docs/METRICS_AND_INTERCEPTORS.md | 827 -------- docs/agents/README.md | 55 - docs/agents/concepts/runtime.md | 74 - docs/agents/concepts/secrets.md | 72 - docs/agents/concepts/streaming-hitl.md | 93 - docs/agents/concepts/teams.md | 92 - docs/agents/concepts/tools.md | 115 -- docs/design/AGENTS_IMPLEMENTATION_PLAN.md | 661 ------- docs/design/AGENTS_PARITY_AUDIT.md | 71 - docs/design/AGENTS_PARITY_ONEPAGER.md | 587 ------ docs/design/AGENT_SECRETS.md | 134 -- docs/design/AGENT_STREAMING.md | 153 -- docs/design/AGENT_TEAMS.md | 75 - docs/design/AGENT_TESTING.md | 109 -- docs/design/AGENT_TOOLS_DSL.md | 106 -- docs/design/EVENT_INTERCEPTOR_SYSTEM.md | 907 --------- docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md | 554 ------ docs/design/WORKER_DESIGN.md | 1781 ------------------ docs/design/WORKFLOW_DSL.md | 561 ------ examples/agents/README.md | 88 - harness/README.md | 50 - harness/manifests/README.md | 132 -- 27 files changed, 66 insertions(+), 8993 deletions(-) delete mode 100644 DESIGN.md delete mode 100644 docs/METRICS_AND_INTERCEPTORS.md delete mode 100644 docs/agents/README.md delete mode 100644 docs/agents/concepts/runtime.md delete mode 100644 docs/agents/concepts/secrets.md delete mode 100644 docs/agents/concepts/streaming-hitl.md delete mode 100644 docs/agents/concepts/teams.md delete mode 100644 docs/agents/concepts/tools.md delete mode 100644 docs/design/AGENTS_IMPLEMENTATION_PLAN.md delete mode 100644 docs/design/AGENTS_PARITY_AUDIT.md delete mode 100644 docs/design/AGENTS_PARITY_ONEPAGER.md delete mode 100644 docs/design/AGENT_SECRETS.md delete mode 100644 docs/design/AGENT_STREAMING.md delete mode 100644 docs/design/AGENT_TEAMS.md delete mode 100644 docs/design/AGENT_TESTING.md delete mode 100644 docs/design/AGENT_TOOLS_DSL.md delete mode 100644 docs/design/EVENT_INTERCEPTOR_SYSTEM.md delete mode 100644 docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md delete mode 100644 docs/design/WORKER_DESIGN.md delete mode 100644 docs/design/WORKFLOW_DSL.md delete mode 100644 examples/agents/README.md delete mode 100644 harness/README.md delete mode 100644 harness/manifests/README.md diff --git a/AGENTS.md b/AGENTS.md index c0d7135..20fbe72 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,663 +1,16 @@ -# Conductor Ruby SDK - AI Agent Guide - -This document provides an overview of the Conductor Ruby SDK codebase for AI coding agents. - -## Project Overview - -This is the official Ruby SDK for [Conductor OSS](https://github.com/conductor-oss/conductor), a durable workflow orchestration engine. The SDK provides: - -- **Workflow DSL** - Ruby-idiomatic block-based workflow definition -- **Worker Framework** - Multi-threaded task execution with events and metrics -- **Full API Coverage** - 17 Resource APIs, 9 high-level clients -- **LLM/AI Tasks** - Chat completion, embeddings, image/audio generation -- **Agents** - `Conductor::Agents`: agent definitions serialized to the server's `agentConfig`, tool workers, SSE streaming (`lib/conductor/agents/`) - -## Key Design Documents - -| Document | Description | -|----------|-------------| -| [DESIGN.md](DESIGN.md) | High-level architecture and design principles | -| [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) | Worker infrastructure design (polling, events, concurrency) | -| [docs/design/WORKFLOW_DSL.md](docs/design/WORKFLOW_DSL.md) | Workflow DSL design and API reference | -| [docs/design/AGENTS_IMPLEMENTATION_PLAN.md](docs/design/AGENTS_IMPLEMENTATION_PLAN.md) | Agents: verified wire contract, decisions, work breakdown | -| [docs/agents/README.md](docs/agents/README.md) | Agents user guide (tools, streaming/approval, teams, secrets, runtime) | -| [README.md](README.md) | User-facing documentation with examples | -| [CONTRIBUTING.md](CONTRIBUTING.md) | Development workflow and guidelines | - ---- - -## Development Requirements - -**IMPORTANT: All changes MUST follow these requirements:** - -### 1. Linting - -After making any changes, always run RuboCop to ensure code style compliance: - -```bash -# Check for linting issues -bundle exec rubocop - -# Auto-fix safe issues -bundle exec rubocop -a - -# Auto-fix all issues (including unsafe) -bundle exec rubocop -A -``` - -**All code must pass RuboCop checks before being committed.** - -### 2. Testing - -Tests MUST be run after every change: - -```bash -# Run all unit tests (REQUIRED after every change) -bundle exec rspec spec/conductor/ - -# Run specific test file -bundle exec rspec spec/conductor/workflow/dsl/workflow_builder_spec.rb - -# Run with coverage report -bundle exec rspec spec/conductor/ --format documentation -``` - -### 3. Code Coverage - -**Any change MUST increase (or at minimum maintain) code coverage.** - -- New features MUST include comprehensive tests -- Bug fixes MUST include regression tests -- Refactoring MUST NOT decrease coverage - -Check coverage: -```bash -# Run tests with coverage (if SimpleCov is configured) -COVERAGE=true bundle exec rspec spec/conductor/ -``` - -### 4. Build Verification - -Before committing, verify the gem builds correctly: - -```bash -# Verify library loads without errors -bundle exec ruby -Ilib -e "require 'conductor'; puts 'OK: ' + Conductor::VERSION" - -# Verify syntax of all Ruby files -find lib -name "*.rb" -exec ruby -c {} \; - -# Run full test suite -bundle exec rspec -``` - -### Complete Pre-Commit Checklist - -```bash -# 1. Run linter and fix issues -bundle exec rubocop -a - -# 2. Run all tests -bundle exec rspec spec/conductor/ - -# 3. Verify library loads -bundle exec ruby -Ilib -e "require 'conductor'; puts Conductor::VERSION" - -# 4. Check for any remaining lint issues -bundle exec rubocop -``` - ---- - -## Architecture Overview - -``` -┌─────────────────────────────────────────────────────────────┐ -│ User Code │ -│ Conductor.workflow :name do ... end │ -│ class MyWorker; include WorkerModule; end │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Workflow DSL (lib/conductor/workflow/dsl/) │ -│ • WorkflowBuilder - Core DSL engine with task methods │ -│ • WorkflowDefinition - Wrapper with .register/.execute │ -│ • TaskRef, OutputRef, InputRef - Reference types │ -│ • ParallelBuilder, SwitchBuilder - Control flow helpers │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Worker Framework (lib/conductor/worker/) │ -│ • TaskRunner - Polling and execution │ -│ • TaskHandler - Worker orchestration │ -│ • WorkerModule - Mixin for class-based workers │ -│ • Events - Task lifecycle hooks │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ High-Level Clients (lib/conductor/client/) │ -│ • WorkflowClient, TaskClient, MetadataClient │ -│ • WorkflowExecutor - Synchronous execution │ -│ • OrkesClients - Factory for all clients │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Resource APIs (lib/conductor/http/api/) │ -│ • WorkflowResourceApi, TaskResourceApi, etc. │ -│ • Direct mapping to Conductor REST endpoints │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ HTTP Transport (lib/conductor/http/) │ -│ • ApiClient - Auth, serialization, dispatch │ -│ • RestClient - Faraday-based HTTP client │ -│ • Models - 50+ request/response models │ -└─────────────────────────────────────────────────────────────┘ -``` - ---- - -## Worker Framework (Detailed) - -The worker framework provides multi-threaded task execution with a comprehensive event system for interceptors and metrics. - -### Component Hierarchy - -``` -┌─────────────────────────────────────────────────────────────────────┐ -│ User Code │ -│ (Worker classes, Worker.define blocks) │ -└─────────────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────────┐ -│ TaskHandler │ -│ • Discovers workers (registry + auto-scan) │ -│ • Resolves configuration (3-tier hierarchy) │ -│ • Creates one Thread per worker type │ -│ • Manages lifecycle (start/stop/join) │ -│ • Aggregates events/metrics │ -└─────────────────────────────────────────────────────────────────────┘ - │ - ┌─────────────┼─────────────┐ - ▼ ▼ ▼ -┌─────────────────────┐ ┌─────────────────────┐ ┌─────────────────────┐ -│ TaskRunner │ │ TaskRunner │ │ TaskRunner │ -│ (Worker A) │ │ (Worker B) │ │ (Worker C) │ -│ │ │ │ │ │ -│ • ThreadPoolExecutor│ │ • ThreadPoolExecutor│ │ • ThreadPoolExecutor│ -│ • Batch polling │ │ • Batch polling │ │ • Batch polling │ -│ • Adaptive backoff │ │ • Adaptive backoff │ │ • Adaptive backoff │ -│ • Event publishing │ │ • Event publishing │ │ • Event publishing │ -└─────────────────────┘ └─────────────────────┘ └─────────────────────┘ -``` - -### Multi-Threading Model - -```ruby -# Each TaskRunner has its own ThreadPoolExecutor -class TaskRunner - def initialize(worker, configuration:, event_dispatcher:) - @worker = worker - @executor = Concurrent::ThreadPoolExecutor.new( - min_threads: 1, - max_threads: worker.thread_count, # Configurable per worker - max_queue: 0, # Synchronous handoff - fallback_policy: :caller_runs - ) - end - - def run - while @running - # 1. Check capacity - available_slots = @worker.thread_count - @running_tasks.size - - # 2. Batch poll for tasks - tasks = batch_poll(available_slots) - - # 3. Submit each task to thread pool - tasks.each do |task| - future = @executor.post { execute_and_update(task) } - @running_tasks << future - end - - # 4. Adaptive backoff for empty polls - apply_backoff if tasks.empty? - end - end -end -``` - -### Event System (Interceptors) - -The event system allows hooking into task lifecycle for logging, metrics, and custom behavior: - -```ruby -# Event Types -module Conductor::Worker::Events - PollStarted # Fired before polling - PollCompleted # Fired after successful poll - PollFailure # Fired on poll error - TaskExecutionStarted # Fired before task execution - TaskExecutionCompleted # Fired after successful execution - TaskExecutionFailure # Fired on execution error - TaskUpdateFailure # CRITICAL: Fired when result update fails -end - -# Custom Event Listener (Interceptor) -class MyInterceptor - def on_poll_started(event) - puts "Polling for #{event.task_type}..." - end - - def on_task_execution_started(event) - puts "Starting task #{event.task_id}" - end - - def on_task_execution_completed(event) - puts "Task #{event.task_id} completed in #{event.duration_ms}ms" - end - - def on_task_execution_failure(event) - puts "Task #{event.task_id} FAILED: #{event.cause.message}" - # Send to error tracking service - ErrorTracker.capture(event.cause, context: { task_id: event.task_id }) - end -end - -# Register listener -handler = TaskHandler.new( - configuration: config, - event_listeners: [MyInterceptor.new] -) -``` - -### Event Dispatcher (Thread-Safe) - -```ruby -class SyncEventDispatcher - def initialize - @listeners = Hash.new { |h, k| h[k] = [] } - @mutex = Mutex.new - end - - def register(event_type, listener) - @mutex.synchronize do - @listeners[event_type] << listener - end - end - - def publish(event) - listeners = @mutex.synchronize { @listeners[event.class].dup } - listeners.each do |listener| - begin - listener.call(event) - rescue StandardError => e - # Listener failure is isolated - never breaks the worker - warn "[Conductor] Event listener error: #{e.message}" - end - end - end -end -``` - -### Metrics Collection - -`MetricsCollector.create` returns a collector that emits the canonical -(harmonized) metric surface: - -```ruby -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) -``` - -See [docs/METRICS_AND_INTERCEPTORS.md](docs/METRICS_AND_INTERCEPTORS.md) for -the full metrics catalog and label reference. - -### Worker Configuration (3-Tier Hierarchy) - -Configuration is resolved in order of priority: - -1. **Worker-specific environment variable** (highest priority) - ```bash - CONDUCTOR_WORKER_PROCESS_ORDER_POLL_INTERVAL=200 - ``` - -2. **Global worker environment variable** - ```bash - CONDUCTOR_WORKER_ALL_POLL_INTERVAL=100 - ``` - -3. **Code-level default** (lowest priority) - ```ruby - worker_task 'process_order', poll_interval: 100 - ``` - -### Configuration Properties - -| Property | Type | Default | Description | -|----------|------|---------|-------------| -| `poll_interval` | Integer | 100 | Polling interval in milliseconds | -| `thread_count` | Integer | 1 | Max concurrent tasks per worker | -| `domain` | String | nil | Task domain for isolation | -| `worker_id` | String | auto | Unique worker identifier | -| `poll_timeout` | Integer | 100 | Server-side long poll timeout (ms) | -| `register_task_def` | Boolean | false | Auto-register task definition | -| `paused` | Boolean | false | Pause worker (stop polling) | - -### Task Context (Thread-Local) - -```ruby -# Access execution context from anywhere in worker code -def execute(task) - ctx = Conductor::Worker::TaskContext.current - - ctx.add_log("Processing task #{ctx.task_id}") - ctx.add_log("Retry count: #{ctx.retry_count}") - - # Long-running task - set callback - if will_take_long? - ctx.set_callback_after(60) # Check back in 60 seconds - return TaskInProgress.new(output: { status: 'processing' }) - end - - { result: 'success' } -end -``` - ---- - -## Directory Structure - -``` -lib/conductor/ -├── version.rb # VERSION constant -├── configuration.rb # Configuration class -├── exceptions.rb # Exception hierarchy -├── client/ # High-level client facades -│ ├── workflow_client.rb -│ ├── task_client.rb -│ ├── metadata_client.rb -│ └── ... -├── http/ -│ ├── api/ # Resource API classes (17) -│ │ ├── workflow_resource_api.rb -│ │ ├── task_resource_api.rb -│ │ └── ... -│ ├── models/ # HTTP models (50+) -│ │ ├── workflow_def.rb -│ │ ├── workflow_task.rb -│ │ ├── task.rb -│ │ └── ... -│ ├── api_client.rb # Auth + serialization -│ └── rest_client.rb # Faraday HTTP client -├── orkes/ # Orkes Cloud specific -│ ├── orkes_clients.rb # Main factory -│ └── models/ -├── worker/ # Worker framework -│ ├── task_runner.rb # Polling loop + ThreadPoolExecutor -│ ├── task_handler.rb # Worker management -│ ├── worker.rb # Worker module -│ ├── worker_config.rb # Configuration resolver -│ ├── worker_registry.rb # Global worker registry -│ ├── task_context.rb # Thread-local context -│ ├── task_in_progress.rb # Long-running task signal -│ ├── events/ # Event system -│ │ ├── conductor_event.rb # Base event class -│ │ ├── task_runner_events.rb # All event types -│ │ ├── sync_event_dispatcher.rb # Thread-safe dispatcher -│ │ ├── listeners.rb # Listener protocol -│ │ └── listener_registry.rb # Registration helper -│ └── telemetry/ # Metrics -│ ├── metrics_collector.rb # MetricsCollector class + NullBackend -│ └── prometheus_backend.rb # PrometheusBackend + MetricsServer -└── workflow/ - ├── dsl/ # Workflow DSL - │ ├── workflow_builder.rb # Core DSL engine (~1000 lines) - │ ├── workflow_definition.rb # Wrapper class - │ ├── task_ref.rb # Task reference - │ ├── output_ref.rb # Output reference (task[:field]) - │ ├── input_ref.rb # Input reference (wf[:param]) - │ ├── parallel_builder.rb # parallel do...end - │ └── switch_builder.rb # decide do...end - ├── llm/ # LLM helper classes - │ ├── chat_message.rb - │ ├── tool_call.rb - │ ├── tool_spec.rb - │ └── embedding_model.rb - ├── task_type.rb # Task type constants - ├── timeout_policy.rb - └── workflow_executor.rb -lib/conductor/agents.rb # require 'conductor/agents' entry point + default runtime -lib/conductor/agents/ -├── agent.rb # Agent definition + sugar (add_tool, hands_off_to, on_approval, >>) -├── tool_def.rb # ToolDef, ToolType, server-side tool factories -├── tools.rb # `tool def` DSL, describe, requires_approval, registries -├── tools/schema_builder.rb # keyword defaults -> JSON schema (AST) -├── tools/secret_scanner.rb # secret('X') literals -> credentials -├── tools/ruby_llm_adapter.rb # RubyLLM::Tool -> ToolDef -├── guardrail.rb, termination.rb, handoff.rb, callback_handler.rb, memory.rb, prompt_template.rb -├── config_serializer.rb # Agent tree -> agentConfig (Python-identical) -└── runtime/ - ├── agent_runtime.rb # call_sync / call_async / deploy / serve - ├── sse_client.rb # GET /agent/stream SSE with reconnect - ├── status_poller.rb # polling fallback - ├── execution.rb # Execution, ToolCall, TokenUsage, FinishReason - ├── approval_request.rb # waiting -> approve / reject - ├── tool_registry.rb # ToolDef -> Worker with Python TaskDef defaults - ├── dispatch.rb # task input -> kwargs -> result - ├── system_workers.rb # termination / guardrail / callback / handoff bodies - ├── secrets.rb # secret(), secrets_env() - └── agent_config.rb # CONDUCTOR_AGENT_* settings -``` - ---- - -## Key Files to Understand - -### Workflow DSL (Most Important) - -1. **`lib/conductor/workflow/dsl/workflow_builder.rb`** (~1000 lines) - - Core DSL engine with all task methods - - `simple`, `http`, `wait`, `terminate`, `sub_workflow` - - `parallel`, `decide`, `loop_over`, `when_true/when_false` - - LLM tasks: `llm_chat`, `llm_embed`, `generate_image`, etc. - - Value resolution: `OutputRef`, `InputRef` → expression strings - -2. **`lib/conductor/workflow/dsl/workflow_definition.rb`** - - Wrapper class returned by `Conductor.workflow` - - Provides `.register()`, `.execute()`, `.call()` methods - - Delegates to WorkflowExecutor for execution - -3. **`lib/conductor/workflow/dsl/task_ref.rb`** - - Stores task metadata during DSL evaluation - - Converts to `WorkflowTask` model for serialization - - Supports `[]` operator for output references - -### Worker Framework - -1. **`lib/conductor/worker/task_runner.rb`** - - Main polling loop with adaptive backoff - - ThreadPoolExecutor for concurrent task execution - - Event publishing for lifecycle hooks - -2. **`lib/conductor/worker/worker.rb`** - - `WorkerModule` mixin for class-based workers - - `worker_task` class method for registration - - `Conductor::Worker.define` for block-based workers - -3. **`lib/conductor/worker/events/`** - - Event classes for task lifecycle - - SyncEventDispatcher for thread-safe event publishing - - ListenerRegistry for listener management - -### HTTP Layer - -1. **`lib/conductor/http/api_client.rb`** - - Token management with TTL-based refresh - - Request serialization, response deserialization - - Retry logic for 401/403 errors - -2. **`lib/conductor/http/models/workflow_def.rb`** - - WorkflowDef model with all workflow properties - - Used for registration and serialization - ---- - -## Common Patterns - -### Creating a Workflow - -```ruby -workflow = Conductor.workflow :my_workflow, version: 1, executor: executor do - user = simple :get_user, user_id: wf[:user_id] - simple :send_email, email: user[:email] - output result: user[:name] -end - -workflow.register(overwrite: true) -result = workflow.execute(input: { user_id: 123 }) -``` - -### Task Reference Flow - -``` -DSL Method Call TaskRef Created WorkflowTask Generated -───────────────────────────────────────────────────────────────────────── -simple :foo, x: wf[:y] → TaskRef(ref: 'foo_ref') → WorkflowTask( - task_name: 'foo' name: 'foo' - inputs: {...} type: 'SIMPLE' - input_parameters: {...} - ) -``` - -### Output References - -```ruby -task[:field] # → OutputRef → "${task_ref.output.field}" -task[:nested][:path] # → OutputRef → "${task_ref.output.nested.path}" -wf[:param] # → InputRef → "${workflow.input.param}" -wf.var(:counter) # → InputRef → "${workflow.variables.counter}" -``` - ---- - -## Testing - -### Test Structure - -``` -spec/ -├── conductor/ -│ ├── workflow/ -│ │ ├── dsl/ -│ │ │ └── workflow_builder_spec.rb # 52 DSL tests -│ │ └── llm_tasks_spec.rb # LLM helper tests -│ ├── client/ -│ ├── http/ -│ ├── worker/ -│ └── agents/ # Agents unit + contract tests (no server) -│ └── contract_spec.rb # 20 golden configs must equal python-sdk + validate against agent-schema.json -├── fixtures/agents/ # vendored agent-schema.json and golden configs -├── agents/ # Replay tests against WireMock (conductor-mocks recordings) -└── integration/ # Requires live server; agents/ wraps the 19 examples in OSS playback -``` - -### Running Tests - -```bash -# Run all unit tests (REQUIRED after every change) -bundle exec rspec spec/conductor/ - -# Run DSL tests specifically -bundle exec rspec spec/conductor/workflow/dsl/ - -# Run worker tests -bundle exec rspec spec/conductor/worker/ - -# Run with documentation format -bundle exec rspec --format documentation - -# Integration tests (requires Conductor server) -CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ - -# Agents replay tests (WireMock serving conductor-mocks/mocks/agent/tool_happy_path on :8080) -CONDUCTOR_AGENTS_REPLAY_URL=http://localhost:8080 bundle exec rspec spec/agents -``` - -No Ruby installed? The suite runs in the official image: - -```bash -docker run --rm -v "$PWD":/app:z -w /app -v ruby-sdk-bundle:/usr/local/bundle:z ruby:3.3 \ - bash -lc "bundle install --quiet && bundle exec rspec spec/conductor/ && bundle exec rubocop" -``` - ---- - -## Making Changes - -### Adding a New Task Type - -1. Add constant to `lib/conductor/workflow/task_type.rb` -2. Add DSL method to `lib/conductor/workflow/dsl/workflow_builder.rb` -3. Handle conversion in `lib/conductor/workflow/dsl/task_ref.rb` -4. Add tests in `spec/conductor/workflow/dsl/workflow_builder_spec.rb` -5. Update examples in `examples/workflow_dsl.rb` -6. Run linter: `bundle exec rubocop -a` -7. Run tests: `bundle exec rspec spec/conductor/` - -### Adding a New API Endpoint - -1. Add method to appropriate Resource API in `lib/conductor/http/api/` -2. Add corresponding method to high-level client in `lib/conductor/client/` -3. Add tests in `spec/conductor/http/api/` and `spec/conductor/client/` -4. Run linter: `bundle exec rubocop -a` -5. Run tests: `bundle exec rspec spec/conductor/` - -### Modifying Worker Behavior - -1. Review `docs/design/WORKER_DESIGN.md` for detailed design -2. Modify `lib/conductor/worker/task_runner.rb` for polling behavior -3. Modify `lib/conductor/worker/worker.rb` for worker definition -4. Add tests in `spec/conductor/worker/` -5. Run linter: `bundle exec rubocop -a` -6. Run tests: `bundle exec rspec spec/conductor/` - -### Changing the agents wire format - -1. The Python SDK (`python-sdk/src/conductor/ai/agents/config_serializer.py`) is the parity source; the server (`conductor/agentspan`) is the contract -2. Edit `lib/conductor/agents/config_serializer.rb`, then `bundle exec rspec spec/conductor/agents/contract_spec.rb` -3. New wire fields need a golden fixture: add the agent to `examples/agents/golden_agents.rb`, generate the Python side with `python-sdk/examples/agents/dump_agent_configs.py`, vendor it into `spec/fixtures/agents/configs/` -4. Runtime behaviour is verified by replay (`spec/agents`); record new scenarios in `conductor-oss/conductor-mocks` - -### Adding Event Listeners / Interceptors - -1. Create class implementing listener methods (`on_poll_started`, `on_task_execution_completed`, etc.) -2. Register with TaskHandler via `event_listeners:` option -3. Add tests in `spec/conductor/worker/events/` - ---- - -## Important Conventions - -- **Snake case** for methods and variables -- **Keyword arguments** for optional parameters -- **Blocks** for control flow (`parallel do`, `decide do`) -- **Symbol-to-string** task names are auto-converted -- **Output references** use `[]` operator (`task[:field]`) -- **Input references** use `wf[:param]` syntax -- **Thread safety** - Use Mutex for shared state in event system - ---- - -## Dependencies - -**Runtime:** -- `faraday ~> 2.0` - HTTP client -- `faraday-net_http_persistent ~> 2.0` - Connection pooling -- `faraday-retry ~> 2.0` - Automatic retries -- `concurrent-ruby ~> 1.2` - Thread pool executor - -**Development:** -- `rspec ~> 3.0` - Testing -- `webmock ~> 3.0` - HTTP mocking -- `rubocop ~> 1.0` - Linting +# Working in this repository + +- Follow [CONTRIBUTING.md](CONTRIBUTING.md) for setup and validation. +- After changes, run `bundle exec rspec spec/conductor/` and `bundle exec rubocop`. +- Before committing, verify the library loads, check Ruby syntax, and run `bundle exec rspec`. +- Add tests for new behavior and bug fixes; do not reduce coverage. +- Use keyword arguments for options and mutexes for shared worker state. +- Keep documentation brief; prefer runnable examples over separate guides or design reports. + +Code: `lib/conductor/client/` (clients), `http/` (transport/models), +`workflow/dsl/` (workflow definitions), `worker/` (execution), `agents/` (agents). +Unit tests mirror these paths under `spec/conductor/`. + +For agent wire changes, use the Python SDK serializer as the parity source and +Conductor's `agentspan` module as the server contract. Update Python-derived golden +fixtures and run `bundle exec rspec spec/conductor/agents/contract_spec.rb`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 684fe9f..26a9f27 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,7 +23,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Added -- Canonical (harmonized) metrics as the sole metric surface -- [details](docs/METRICS_AND_INTERCEPTORS.md#detailed-technical-notes----unreleased) +- Canonical (harmonized) metrics as the sole metric surface - Bounded `uri` label on `http_api_client_request_seconds`: uses path templates (e.g. `/workflow/{workflowId}`) instead of fully-resolved paths, preventing metric cardinality explosion - `WorkflowStatusProbe` in harness: opt-in probe (via `HARNESS_PROBE_RATE_PER_SEC`) that exercises UUID-bearing endpoints to validate template URI metrics @@ -78,7 +78,7 @@ end ### Added -- **Agents** (`require 'conductor/agents'`) - Ruby port of the Python SDK's agents package, same `agentConfig` on the wire -- [guide](docs/agents/README.md) +- **Agents** (`require 'conductor/agents'`) - Ruby port of the Python SDK's agents package, same `agentConfig` on the wire - `tool def` DSL: types from keyword defaults, secrets from `secret('...')` literals, `describe`, `requires_approval`, module scoping, RubyLLM::Tool adapter - `Agent` with `add_tool`, `add_agent`, `hands_off_to`, `redact`, `stop_when`, `stop_after`, `on_approval`, `>>`; guardrails, termination conditions, handoffs, callbacks, memory, prompt templates - `ConfigSerializer` verified against the Python SDK's 19 golden configs and `agent-schema.json` diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 5255a2a..5dcd5e1 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,226 +1,42 @@ -# Contributing to Conductor Ruby SDK +# Contributing -Thank you for your interest in contributing to the Conductor Ruby SDK! This document provides guidelines and instructions for contributing. - -## Code of Conduct - -Please be respectful and constructive in all interactions. We welcome contributors of all experience levels. - -## Getting Started - -### Prerequisites - -- Ruby 2.6+ (Ruby 3.2+ recommended) -- Bundler -- Git - -### Setup - -1. Fork the repository on GitHub -2. Clone your fork: - ```bash - git clone https://github.com/YOUR_USERNAME/ruby-sdk.git - cd ruby-sdk - ``` - -3. Install dependencies: - ```bash - bundle install - ``` - -4. Run tests to verify setup: - ```bash - bundle exec rspec spec/conductor/ - ``` - -## Development Workflow - -### Branch Naming - -- `feature/description` - New features -- `fix/description` - Bug fixes -- `docs/description` - Documentation updates -- `refactor/description` - Code refactoring - -### Making Changes - -1. Create a new branch: - ```bash - git checkout -b feature/my-new-feature - ``` - -2. Make your changes - -3. Run tests: - ```bash - bundle exec rspec spec/conductor/ - ``` - -4. Run linting: - ```bash - bundle exec rubocop - ``` - -5. Commit your changes: - ```bash - git commit -m "Add feature: description of changes" - ``` - -6. Push to your fork: - ```bash - git push origin feature/my-new-feature - ``` - -7. Create a Pull Request - -## Testing - -### Running Tests +Use Ruby 3.0+ and Bundler. From the repository root: ```bash -# All unit tests +bundle install bundle exec rspec spec/conductor/ - -# Specific test file -bundle exec rspec spec/conductor/client/workflow_client_spec.rb - -# With documentation format -bundle exec rspec spec/conductor/ --format documentation - -# Integration tests (requires Conductor server) -CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ -``` - -### Writing Tests - -- Place unit tests in `spec/conductor/` -- Place integration tests in `spec/integration/` -- Use descriptive test names -- Follow existing test patterns - -Example test structure: -```ruby -RSpec.describe Conductor::Client::WorkflowClient do - describe '#start' do - context 'when workflow exists' do - it 'returns a workflow ID' do - # test implementation - end - end - - context 'when workflow does not exist' do - it 'raises an ApiError' do - # test implementation - end - end - end -end -``` - -## Code Style - -We use RuboCop for code style enforcement. Key guidelines: - -- Use 2 spaces for indentation -- Use snake_case for methods and variables -- Use CamelCase for classes and modules -- Add documentation comments for public methods -- Keep methods small and focused - -Run RuboCop to check your code: -```bash bundle exec rubocop - -# Auto-fix safe issues -bundle exec rubocop -a +bundle exec ruby -Ilib -rconductor -e 'puts Conductor::VERSION' +gem build conductor_ruby.gemspec ``` -## Documentation - -- Update README.md for user-facing changes -- Add YARD documentation for new public methods -- Update CHANGELOG.md for notable changes -- Include examples for new features +Add regression tests for fixes and tests for new behavior. Keep coverage at least +at its current level. Put unit tests in `spec/conductor/` and server tests in +`spec/integration/`. Update runnable examples when the public API changes. -### YARD Documentation Example +For local OSS integration tests (requires Docker and Ruby): -```ruby -# Starts a workflow execution -# -# @param name [String] The workflow name -# @param input [Hash] The workflow input (default: {}) -# @param version [Integer, nil] The workflow version (optional) -# @return [String] The workflow ID -# @raise [ApiError] If the workflow doesn't exist -# -# @example Start a simple workflow -# workflow_id = client.start('my_workflow', input: { key: 'value' }) -# -def start(name, input: {}, version: nil) - # implementation -end +```bash +scripts/run-integration-oss.sh ``` -## Pull Request Guidelines +Agent contract tests use the schema and Python fixtures in +[spec/fixtures/agents](spec/fixtures/agents/README.md). The Python serializer is +the parity source; the server's `agentConfig` is the wire contract. Agent playback +setup lives in [.github/workflows/agents-playback.yml](.github/workflows/agents-playback.yml). -### Before Submitting +Without local Ruby, run unit tests and lint in Docker: -- [ ] Tests pass locally -- [ ] RuboCop passes (or issues are intentional) -- [ ] Documentation is updated -- [ ] CHANGELOG.md is updated (if applicable) -- [ ] Commits are clean and well-described - -### PR Description - -Include: -- Summary of changes -- Motivation/context -- How to test -- Screenshots (if UI-related) - -### Review Process - -1. Automated CI checks run -2. Maintainers review code -3. Address feedback -4. Merge when approved - -## Release Process - -Releases are managed by maintainers: - -1. Update version in `lib/conductor/version.rb` -2. Update CHANGELOG.md -3. Create GitHub release with tag -4. CI automatically publishes to RubyGems - -## Project Structure - -``` -lib/ -├── conductor.rb # Main entry point -├── conductor/ -│ ├── version.rb # Version constant -│ ├── configuration.rb # Configuration class -│ ├── exceptions.rb # Exception classes -│ ├── client/ # High-level clients -│ ├── http/ -│ │ ├── api/ # Resource API classes -│ │ ├── models/ # Model classes -│ │ ├── api_client.rb # HTTP client wrapper -│ │ └── rest_client.rb # Faraday client -│ ├── orkes/ # Orkes-specific code -│ ├── worker/ # Worker framework -│ └── workflow/ # Workflow DSL +```bash +docker run --rm -v "$PWD":/app:z -w /app \ + -v ruby-sdk-bundle:/usr/local/bundle:z ruby:3.3 \ + bash -lc 'bundle install --quiet && bundle exec rspec spec/conductor/ && bundle exec rubocop' ``` -## Getting Help - -- Open an issue for bugs or feature requests -- Join [Conductor Slack](https://join.slack.com/t/orkes-conductor/shared_invite/zt-2vdbx239s-Eacdyqya9giNLHfrCavfaA) -- Check existing issues and PRs - -## License +The load harness runs with `bundle exec ruby harness/main.rb` using the same +server/auth environment variables as the SDK. It continuously starts workflows; +`HARNESS_WORKFLOWS_PER_SEC` defaults to 2. Kubernetes templates are in +[harness/manifests](harness/manifests/). -By contributing, you agree that your contributions will be licensed under the Apache 2.0 License. +For releases, update `lib/conductor/version.rb` and `CHANGELOG.md`, then create a +GitHub release with a tag; CI publishes the gem. diff --git a/DESIGN.md b/DESIGN.md deleted file mode 100644 index 86c4ca1..0000000 --- a/DESIGN.md +++ /dev/null @@ -1,304 +0,0 @@ -# Conductor Ruby SDK - Design Document - -## Overview - -This SDK is a Ruby port of the [Conductor Python SDK](https://github.com/conductor-sdk/conductor-python), designed to provide a Ruby-native interface to Conductor OSS while maintaining architectural parity with the Python implementation. - -## Design Principles - -1. **Follow Python SDK architecture** - Maintain the same layers, patterns, and behaviors -2. **Ruby-idiomatic API** - Use Ruby conventions (snake_case, blocks, symbols) while keeping the same capabilities -3. **Feature parity** - Support all features from the Python SDK -4. **Clean DSL** - Block-based workflow definition with method chaining for output references - -## Architecture Layers - -``` -┌─────────────────────────────────────────────────────────────┐ -│ User Code │ -│ Conductor.workflow :name do ... end │ -│ class MyWorker; include WorkerModule; end │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Workflow DSL │ -│ • Conductor.workflow - Entry point for workflow definition │ -│ • WorkflowBuilder - Core DSL engine with task methods │ -│ • WorkflowDefinition - Wrapper with .register/.execute │ -│ • OutputRef/InputRef - Reference expression builders │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Worker Framework │ -│ • WorkerModule - Mixin for class-based workers │ -│ • Worker.define - Block-based worker definition │ -│ • TaskRunner - Polling and execution │ -│ • TaskHandler - Worker orchestration │ -│ • Event system - Lifecycle hooks │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ High-Level Clients (Facades) │ -│ • WorkflowClient, TaskClient, MetadataClient │ -│ • SchedulerClient, PromptClient, SecretClient │ -│ • IntegrationClient, AuthorizationClient, SchemaClient │ -│ • WorkflowExecutor - Synchronous workflow execution │ -│ • OrkesClients - Factory for all clients │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Resource APIs (17 classes) │ -│ • WorkflowResourceApi, TaskResourceApi │ -│ • MetadataResourceApi, SchedulerResourceApi │ -│ • EventResourceApi, WorkflowBulkResourceApi │ -│ • PromptResourceApi, SecretResourceApi, etc. │ -│ • Direct mapping to Conductor REST API │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ HTTP Transport │ -│ • ApiClient - Auth, serialization, dispatch │ -│ • RestClient - Faraday with HTTP/2, connection pooling │ -│ • BaseModel - SWAGGER_TYPES pattern for serialization │ -│ • 50+ Model classes for request/response types │ -└─────────────────────────────────────────────────────────────┘ - │ -┌─────────────────────────────▼───────────────────────────────┐ -│ Configuration │ -│ • Server URL, auth settings │ -│ • Token management (class-level cache, TTL refresh) │ -│ • Environment variable support │ -└─────────────────────────────────────────────────────────────┘ -``` - -## Workflow DSL Design - -The SDK provides a clean, Ruby-idiomatic DSL for building workflows: - -### Entry Point - -```ruby -workflow = Conductor.workflow :order_processing, version: 1, executor: executor do - # Workflow definition using DSL methods - user = simple :get_user, user_id: wf[:user_id] - simple :send_email, email: user[:email] - output result: user[:name] -end - -workflow.register(overwrite: true) -result = workflow.execute(input: { user_id: 123 }) -``` - -### DSL Components - -| Component | Purpose | -|-----------|---------| -| `WorkflowBuilder` | Core DSL engine with all task methods | -| `WorkflowDefinition` | Wrapper returned by `Conductor.workflow` | -| `TaskRef` | Stores task metadata, converts to WorkflowTask | -| `OutputRef` | Enables `task[:field]` syntax | -| `InputRef` | Enables `wf[:param]` syntax | -| `ParallelBuilder` | Handles `parallel do...end` blocks | -| `SwitchBuilder` | Handles `decide do...end` blocks | - -### Reference Resolution - -The DSL automatically converts references to Conductor expression strings: - -```ruby -task[:field] # → "${task_ref.output.field}" -task[:nested][:path] # → "${task_ref.output.nested.path}" -wf[:param] # → "${workflow.input.param}" -wf.var(:counter) # → "${workflow.variables.counter}" -``` - -### Task Types (25+) - -**Basic Tasks:** -- `simple` - Worker task execution -- `http` - HTTP request -- `javascript` - Inline JavaScript execution -- `jq` - JSON JQ transformation -- `set` - Set workflow variables -- `human` - Human/manual task - -**Control Flow:** -- `parallel do...end` - Fork/Join execution -- `decide expr do...end` - Switch/case branching -- `when_true/when_false` - Conditional shortcuts -- `loop_over`, `loop_while`, `loop_times` - Iteration -- `sub_workflow` - Call another workflow -- `inline_workflow` - Define sub-workflow inline - -**System Tasks:** -- `wait` - Wait for duration or time -- `terminate` - End workflow -- `event` - Publish event -- `wait_for_webhook` - Wait for external callback -- `http_poll` - Poll HTTP endpoint -- `dynamic`, `dynamic_fork` - Runtime task resolution -- `kafka_publish` - Publish to Kafka -- `start_workflow` - Fire-and-forget workflow start - -**LLM/AI Tasks:** -- `llm_chat` - Chat completion -- `llm_complete` - Text completion -- `llm_embed` - Generate embeddings -- `llm_index`, `llm_search` - Vector operations -- `llm_store_embeddings`, `llm_search_embeddings` - Vector DB operations -- `generate_image`, `generate_audio` - Media generation -- `list_mcp_tools`, `call_mcp_tool` - MCP integration -- `get_document` - Document retrieval - -## Worker Framework Design - -See [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) for detailed design. - -### Worker Definition Patterns - -```ruby -# Pattern 1: Class-based (recommended for complex workers) -class OrderProcessor - include Conductor::Worker::WorkerModule - worker_task 'process_order', poll_interval: 200, thread_count: 5 - - def execute(task) - order_id = get_input(task, 'order_id') - # Process... - { processed: true, order_id: order_id } - end -end - -# Pattern 2: Block-based (quick workers) -Conductor::Worker.define('send_email', poll_interval: 100) do |task| - EmailService.deliver(task.input_data['to'], task.input_data['subject']) - { sent: true } -end -``` - -### Concurrency Model - -**Python SDK:** Uses `multiprocessing.Process` (one OS process per worker) to avoid GIL. - -**Ruby SDK:** Uses `Thread` + `concurrent-ruby`'s `ThreadPoolExecutor`. Ruby's GVL releases during I/O operations, making threads sufficient for typical worker workloads. - -## Auth Flow (Matches Python SDK) - -| Scenario | Behavior | -|----------|----------| -| No auth configured | Skip auth. No `X-Authorization` header. | -| Initial token fetch | POST `/api/token` → cache token (class-level). | -| `/token` returns 404 | **Disable auth.** Set `auth_configured? = false`. Log "OSS mode". | -| Token TTL expired | Proactively refresh before next request. | -| Server 401/403 with `EXPIRED_TOKEN` | Force-refresh, retry request once. | -| Server 401/403 with `INVALID_TOKEN` | Force-refresh, retry request once. | -| Token refresh fails | Exponential backoff: 2^n seconds, max 5 attempts. | - -## File Structure - -``` -lib/conductor/ -├── version.rb # VERSION constant -├── configuration.rb # Configuration class -├── exceptions.rb # Exception hierarchy -├── client/ # High-level client facades (9) -│ ├── workflow_client.rb -│ ├── task_client.rb -│ ├── metadata_client.rb -│ └── ... -├── http/ -│ ├── api/ # Resource API classes (17) -│ │ ├── workflow_resource_api.rb -│ │ ├── task_resource_api.rb -│ │ └── ... -│ ├── models/ # HTTP models (50+) -│ │ ├── workflow_def.rb -│ │ ├── workflow_task.rb -│ │ └── ... -│ ├── api_client.rb # Auth + serialization -│ └── rest_client.rb # Faraday HTTP client -├── orkes/ # Orkes Cloud specific -│ ├── orkes_clients.rb # Main factory -│ └── models/ -├── worker/ # Worker framework -│ ├── task_runner.rb # Polling loop -│ ├── task_handler.rb # Worker management -│ ├── worker.rb # Worker module -│ └── events/ # Event system -└── workflow/ - ├── dsl/ # Workflow DSL - │ ├── workflow_builder.rb # Core DSL engine - │ ├── workflow_definition.rb # Wrapper class - │ ├── task_ref.rb # Task reference - │ ├── output_ref.rb # Output reference - │ ├── input_ref.rb # Input reference - │ ├── parallel_builder.rb # parallel do...end - │ └── switch_builder.rb # decide do...end - ├── llm/ # LLM helper classes - │ ├── chat_message.rb - │ ├── tool_call.rb - │ ├── tool_spec.rb - │ └── embedding_model.rb - ├── task_type.rb # Task type constants - ├── timeout_policy.rb - └── workflow_executor.rb -lib/conductor/agents.rb # require 'conductor/agents' -lib/conductor/agents/ # Agents (see docs/agents/README.md) -├── agent.rb, tool_def.rb, tools.rb # definition layer + `tool def` DSL -├── guardrail.rb, termination.rb, handoff.rb, callback_handler.rb, memory.rb, prompt_template.rb -├── config_serializer.rb # -> agentConfig, identical to the Python SDK -└── runtime/ # AgentRuntime, SseClient, Execution, ApprovalRequest, ToolRegistry, Dispatch, Secrets -``` - -## Agents - -`Conductor::Agents` ports the Python SDK's `conductor.ai.agents` package. An `Agent` tree is -serialized by `ConfigSerializer` to the same `agentConfig` JSON Python sends; the server compiles -it into a workflow and runs the LLM loop. `AgentRuntime` starts the execution -(`POST /api/agent/start`), registers a Conductor worker for every task the server lists in -`requiredWorkers` (the user's `tool def` tools plus `_termination`, custom guardrails and -callbacks), and follows the run over SSE (`GET /api/agent/stream/{id}`, polling fallback). -Secrets travel as `TaskDef.runtimeMetadata` names and come back as `Task.runtimeMetadata` -values, read inside tools with `secret('NAME')`. The transport for `/api/agent/*` is -`AgentResourceApi` / `AgentClient` like every other resource. Decisions and the verified wire -contract are in `docs/design/AGENTS_IMPLEMENTATION_PLAN.md`. - -## Dependencies - -```ruby -# Core HTTP -gem 'faraday', '~> 2.0' -gem 'faraday-net_http_persistent', '~> 2.0' # Connection pooling -gem 'faraday-retry', '~> 2.0' # Automatic retries - -# Concurrency -gem 'concurrent-ruby', '~> 1.2' # ThreadPoolExecutor - -# JSON -gem 'json', '>= 2.0' -``` - -## Testing Strategy - -1. **Unit tests** - Mock HTTP, test logic in isolation (~400 tests) -2. **Integration tests** - Against Conductor server (~137 tests) -3. **Example-driven** - All examples in `examples/` directory - -```bash -# Run all tests -bundle exec rspec - -# Run DSL tests -bundle exec rspec spec/conductor/workflow/dsl/ - -# Integration tests (requires server) -CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ -``` - -## Related Documents - -- [AGENTS.md](AGENTS.md) - Guide for AI coding agents -- [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) - Worker infrastructure design -- [docs/design/WORKFLOW_DSL.md](docs/design/WORKFLOW_DSL.md) - Workflow DSL design -- [README.md](README.md) - User documentation -- [CONTRIBUTING.md](CONTRIBUTING.md) - Development guidelines diff --git a/README.md b/README.md index dfd0348..16db78b 100644 --- a/README.md +++ b/README.md @@ -1,540 +1,45 @@ # Conductor Ruby SDK -Official Ruby SDK for [Conductor OSS](https://github.com/conductor-oss/conductor) - a durable workflow orchestration engine. +Ruby clients, workflow definitions, workers, and agents for [Conductor](https://github.com/conductor-oss/conductor). +Requires Ruby 3.0 or newer. -[![Gem Version](https://badge.fury.io/rb/conductor_ruby.svg)](https://badge.fury.io/rb/conductor_ruby) -[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE) - -## Features - -- **Full Feature Parity** with Python SDK -- **Ruby-Idiomatic Workflow DSL** - Clean block-based syntax with 25+ task types -- **Worker Framework** - Multi-threaded task execution with class-based and block-based workers -- **LLM/AI Tasks** - Chat completion, embeddings, RAG, image/audio generation -- **Agents** - Define agents with Ruby methods as tools, run them on the server, stream the answer (`require 'conductor/agents'`) -- **Orkes Cloud Support** - Authentication, secrets, integrations, prompts -- **Comprehensive Testing** - 400+ unit tests, 110 integration tests - -## Installation - -Add to your Gemfile: +## Install ```ruby gem 'conductor_ruby' ``` -Or install directly: +Set `CONDUCTOR_SERVER_URL` to your server's API URL (for example, +`http://localhost:8080/api`). For authenticated servers, also set +`CONDUCTOR_AUTH_KEY` and `CONDUCTOR_AUTH_SECRET`. -```bash -gem install conductor_ruby -``` - -## Quick Start +## Start a workflow -### Hello World +This starts an existing workflow definition on your server: ```ruby require 'conductor' -# Configuration (reads CONDUCTOR_SERVER_URL from environment) -config = Conductor::Configuration.new - -# Create clients -clients = Conductor::Orkes::OrkesClients.new(config) -executor = clients.get_workflow_executor - -# Define a worker -class GreetWorker - include Conductor::Worker::WorkerModule - worker_task 'greet' - - def execute(task) - name = get_input(task, 'name', 'World') - { 'result' => "Hello, #{name}!" } - end -end - -# Build workflow using new DSL -workflow = Conductor.workflow :greetings, version: 1, executor: executor do - greet = simple :greet, name: wf[:name] - output result: greet[:result] -end - -# Register and execute -workflow.register(overwrite: true) - -# Start workers -runner = Conductor::Worker::TaskRunner.new(config) -runner.register_worker(GreetWorker.new) -runner.start - -# Execute workflow -result = workflow.execute(input: { 'name' => 'Ruby' }, wait_for_seconds: 30) -puts "Result: #{result.output['result']}" # => "Hello, Ruby!" - -runner.stop -``` - -## Workflow DSL - -The SDK provides a clean, Ruby-idiomatic DSL for building workflows: - -```ruby -workflow = Conductor.workflow :order_processing, version: 1, executor: executor do - # Access workflow inputs with wf[:param] - user = simple :get_user, user_id: wf[:user_id] - - # Reference task outputs with task[:field] - order = simple :validate_order, email: user[:email] - - # HTTP calls - http :call_api, url: 'https://api.example.com', method: :post, body: { id: order[:id] } - - # Parallel execution - parallel do - simple :ship_order, order_id: order[:id] - simple :send_confirmation, email: user[:email] - end - - # Conditional branching - decide order[:region] do - on 'US' do - simple :us_shipping - end - on 'EU' do - simple :eu_shipping - end - otherwise do - terminate :failed, 'Unsupported region' - end - end - - # Set workflow output - output tracking: order[:tracking_number], status: 'completed' -end - -# Register and execute -workflow.register(overwrite: true) -result = workflow.execute(input: { user_id: 123 }, wait_for_seconds: 60) -``` - -### Task Methods Reference - -#### Basic Tasks - -```ruby -# Simple task (worker execution) -result = simple :task_name, input1: 'value', input2: wf[:param] - -# Inline code execution -jq :transform, query: '.items | map(.name)', input: previous[:data] -javascript :compute, script: 'return inputs.a + inputs.b', a: 1, b: 2 - -# Set workflow variables -set_variable :save_state, user_id: user[:id], status: 'active' - -# Human/manual task -human :approval, display_name: 'Manager Approval', form_template: 'approval_form' -``` - -#### HTTP Tasks - -```ruby -# HTTP request -http :call_api, - url: 'https://api.example.com/users', - method: :post, - headers: { 'Authorization' => 'Bearer ${workflow.secrets.api_token}' }, - body: { name: wf[:name], email: wf[:email] } - -# HTTP polling (wait for condition) -http_poll :wait_for_ready, - url: 'https://api.example.com/status/${workflow.input.job_id}', - method: :get, - termination_condition: '$.status == "ready"', - polling_interval: 5, - polling_strategy: :fixed -``` - -#### Control Flow - -```ruby -# Parallel execution (fork/join) -parallel do - simple :branch_a - simple :branch_b - simple :branch_c -end - -# Conditional branching -decide order[:status] do - on 'pending' do - simple :process_pending - end - on 'approved' do - simple :process_approved - end - otherwise do - simple :handle_unknown - end -end - -# Conditional shortcuts -when_true user[:is_premium] do - simple :apply_discount -end - -when_false order[:validated] do - terminate :failed, 'Order validation failed' -end - -# Loop over items -loop_over users[:list], as: :user do - simple :process_user, user_id: iteration[:user][:id] -end - -# Do-while loop -do_while :retry_loop, condition: '${retry_ref.output.success} == false' do - simple :retry_operation -end -``` - -#### Sub-workflows - -```ruby -# Call another workflow -sub_workflow :process_order, - workflow_name: 'order_processor', - version: 2, - input: { order_id: wf[:order_id] } - -# Start workflow (fire-and-forget) -start_workflow :trigger_notification, - workflow_name: 'send_notifications', - input: { user_id: user[:id] } - -# Inline sub-workflow definition -inline_workflow :nested_process do - simple :step1 - simple :step2 -end -``` - -#### Wait and Events - -```ruby -# Wait for duration -wait :pause, duration: '30s' # or '5m', '1h', '2d' - -# Wait until specific time -wait :scheduled, until: '2024-12-25T00:00:00Z' - -# Wait for external webhook -wait_for_webhook :external_callback, - matches: { 'type' => 'payment', 'order_id' => '${workflow.input.order_id}' } - -# Publish event -event :notify, sink: 'conductor:workflow_events', payload: { status: 'completed' } -``` - -#### Termination - -```ruby -# Complete workflow -terminate :success, 'Processing completed successfully' - -# Fail workflow -terminate :failed, 'Validation error: missing required field' -``` - -#### Dynamic Tasks - -```ruby -# Dynamic task name (resolved at runtime) -dynamic :run_handler, task_to_execute: wf[:handler_name] - -# Dynamic fork (parallel tasks determined at runtime) -dynamic_fork :process_all, - tasks_input: wf[:items], - task_name: 'process_item' -``` - -### LLM/AI Tasks - -```ruby -workflow = Conductor.workflow :ai_assistant, executor: executor do - # Chat completion (messages auto-converted from simple format) - response = llm_chat :chat, - provider: 'openai', - model: 'gpt-4', - messages: [ - { role: :system, message: 'You are a helpful assistant.' }, - { role: :user, message: wf[:question] } - ], - temperature: 0.7 - - # Text completion - llm_text :complete, - provider: 'anthropic', - model: 'claude-3-sonnet', - prompt: 'Summarize: ${workflow.input.text}' - - # Generate embeddings - embeddings = llm_embeddings :embed, - provider: 'openai', - model: 'text-embedding-3-small', - text: wf[:document] - - # Store embeddings in vector DB - llm_store_embeddings :store, - provider: 'pinecone', - index: 'documents', - embeddings: embeddings[:embeddings], - metadata: { doc_id: wf[:doc_id] } - - # Search embeddings - llm_search_embeddings :search, - provider: 'pinecone', - index: 'documents', - query: wf[:search_query], - max_results: 10 - - # Generate image - generate_image :create_image, - provider: 'openai', - model: 'dall-e-3', - prompt: 'A sunset over mountains', - size: '1024x1024' - - # Generate audio (text-to-speech) - generate_audio :speak, - provider: 'openai', - model: 'tts-1', - text: response[:content], - voice: 'nova' - - # MCP (Model Context Protocol) integration - tools = list_mcp_tools :get_tools, server_name: 'my_mcp_server' - - call_mcp_tool :use_tool, - server_name: 'my_mcp_server', - tool_name: 'search_documents', - arguments: { query: wf[:query] } - - output answer: response[:content] -end -``` - -### Agents - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } -end - -agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', instructions: 'Answer weather questions.') -agent.add_tool :get_weather - -puts agent.call_sync('Weather in Lisbon?') -``` - -`tool def` turns a method into a tool (types from the keyword defaults, secrets from -`secret('...')` literals). The server runs the LLM loop; this process runs the tools. Teams, -approval (`requires_approval` + `on_approval`), streaming (`call_async`), guardrails and -termination are covered in [docs/agents/](docs/agents/README.md). - -### Output References - -The DSL uses a clean syntax for referencing outputs: - -```ruby -# Workflow input reference -wf[:user_id] # => '${workflow.input.user_id}' - -# Task output reference -task[:field] # => '${task_ref.output.field}' -task[:nested][:path] # => '${task_ref.output.nested.path}' - -# Loop iteration references (inside loop_over) -iteration[:current_item] # Current item being processed -iteration[:index] # Current index (0-based) -iteration[:user][:name] # If `as: :user` specified +client = Conductor::Client::WorkflowClient.new(Conductor::Configuration.new) +workflow_id = client.start('my_workflow', input: { 'name' => 'Ruby' }) +puts workflow_id ``` ## Examples -The `examples/` directory contains comprehensive examples: - -| Example | Description | -|---------|-------------| -| [`helloworld/`](examples/helloworld/) | Simplest complete example - worker + workflow + execution | -| [`workflow_dsl.rb`](examples/workflow_dsl.rb) | Comprehensive new DSL showcase | -| [`simple_worker.rb`](examples/simple_worker.rb) | Worker patterns: class-based, block-based, error handling | -| [`kitchensink.rb`](examples/kitchensink.rb) | All major task types using new DSL | -| [`dynamic_workflow.rb`](examples/dynamic_workflow.rb) | Create and execute workflows at runtime | -| [`workflow_ops.rb`](examples/workflow_ops.rb) | Lifecycle operations: pause, resume, restart, retry | -| [`agentic_workflows/`](examples/agentic_workflows/) | LLM chat and AI workflow examples | -| [`agents/`](examples/agents/) | Agents: tools (`weather.rb`), approval + streaming (`support_approval.rb`), team + secret (`bug_desk.rb`) | - -Run examples: - -```bash -# Set environment variables -export CONDUCTOR_SERVER_URL=http://localhost:8080/api -# For Orkes Cloud: -# export CONDUCTOR_AUTH_KEY=your_key -# export CONDUCTOR_AUTH_SECRET=your_secret - -# Run hello world -cd examples/helloworld && bundle exec ruby helloworld.rb - -# Run DSL showcase -bundle exec ruby examples/workflow_dsl.rb - -# Run kitchen sink -bundle exec ruby examples/kitchensink.rb -``` - -## Worker Framework - -### Class-Based Workers - -```ruby -class ImageProcessor - include Conductor::Worker::WorkerModule - - worker_task 'process_image', poll_interval: 1, thread_count: 4 - - def execute(task) - url = get_input(task, 'image_url') - # Process image... - - result = Conductor::Http::Models::TaskResult.complete - result.add_output_data('processed_url', processed_url) - result.log('Image processed successfully') - result - end -end -``` - -### Block-Based Workers - -```ruby -worker = Conductor::Worker.define('simple_task') do |task| - input = task.input_data['value'] - { result: input * 2 } # Return hash for automatic TaskResult -end -``` - -### Running Workers - -```ruby -runner = Conductor::Worker::TaskRunner.new(config) -runner.register_worker(ImageProcessor.new) -runner.register_worker(worker) -runner.start(threads: 4) - -# Graceful shutdown -trap('INT') { runner.stop } -sleep while runner.running? -``` - -## Configuration - -### Environment Variables - -```bash -export CONDUCTOR_SERVER_URL=http://localhost:8080/api -export CONDUCTOR_AUTH_KEY=your_key # For Orkes Cloud -export CONDUCTOR_AUTH_SECRET=your_secret # For Orkes Cloud -``` - -### Programmatic - -```ruby -config = Conductor::Configuration.new( - server_api_url: 'https://play.orkes.io/api', - auth_key: 'your_key', - auth_secret: 'your_secret', - auth_token_ttl_min: 45, - verify_ssl: true -) -``` - -## API Coverage - -### Resource APIs (17 classes) - -| API | Description | -|-----|-------------| -| WorkflowResourceApi | Workflow execution and management | -| TaskResourceApi | Task polling and updates | -| MetadataResourceApi | Workflow/task definitions | -| SchedulerResourceApi | Scheduled workflows | -| EventResourceApi | Event handlers | -| WorkflowBulkResourceApi | Bulk operations | -| PromptResourceApi | AI prompt templates | -| SecretResourceApi | Secret management | -| IntegrationResourceApi | External integrations | -| + 8 more | Authorization, Users, Groups, Roles, etc. | +- [Workflow DSL](examples/workflow_dsl.rb) +- [Worker configuration](examples/worker_configuration_example.rb) +- [Task context](examples/task_context_example.rb) and [event listeners](examples/task_listener_example.rb) +- [Agents](examples/agents/): tools, streaming, approvals, guardrails, and teams -### High-Level Clients (9 classes) - -```ruby -clients = Conductor::Orkes::OrkesClients.new(config) - -workflow_client = clients.get_workflow_client -task_client = clients.get_task_client -metadata_client = clients.get_metadata_client -scheduler_client = clients.get_scheduler_client -prompt_client = clients.get_prompt_client -secret_client = clients.get_secret_client -authorization_client = clients.get_authorization_client -workflow_executor = clients.get_workflow_executor -``` - -## Testing +To run an agent example from this checkout, configure an LLM integration on your +Conductor server, then run: ```bash -# Unit tests -bundle exec rspec spec/conductor/ - -# Integration tests (requires Conductor server) -CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ +bundle install +CONDUCTOR_AGENT_LLM_MODEL=openai/gpt-4o-mini \ + bundle exec ruby -Ilib examples/agents/01_basic_agent.rb ``` -## Requirements - -- Ruby 2.6+ (Ruby 3+ recommended) -- Conductor OSS 3.x or Orkes Cloud - -## Dependencies - -- `faraday ~> 2.0` - HTTP client -- `faraday-net_http_persistent ~> 2.0` - Connection pooling -- `faraday-retry ~> 2.0` - Automatic retries -- `concurrent-ruby ~> 1.2` - Thread-safe concurrency - -## Contributing - -1. Fork the repository -2. Create your feature branch (`git checkout -b feature/amazing-feature`) -3. Run tests (`bundle exec rspec`) -4. Commit your changes (`git commit -m 'Add amazing feature'`) -5. Push to the branch (`git push origin feature/amazing-feature`) -6. Open a Pull Request - -## License - -Apache 2.0 - see [LICENSE](LICENSE) for details. - -## Links - -- [Conductor OSS](https://github.com/conductor-oss/conductor) -- [Orkes Cloud](https://orkes.io) -- [Documentation](https://conductor-oss.org) -- [Python SDK](https://github.com/conductor-sdk/conductor-python) -- [Community Slack](https://join.slack.com/t/orkes-conductor/shared_invite/zt-2vdbx239s-Eacdyqya9giNLHfrCavfaA) +See [CONTRIBUTING.md](CONTRIBUTING.md) for development commands and +[CHANGELOG.md](CHANGELOG.md) for release history. Licensed under [Apache 2.0](LICENSE). diff --git a/docs/METRICS_AND_INTERCEPTORS.md b/docs/METRICS_AND_INTERCEPTORS.md deleted file mode 100644 index fd2f289..0000000 --- a/docs/METRICS_AND_INTERCEPTORS.md +++ /dev/null @@ -1,827 +0,0 @@ -# Metrics and Interceptors Guide - -The Conductor Ruby SDK can expose Prometheus metrics for worker polling, task -execution, task result updates, payload sizes, workflow starts, and HTTP API -client latency. It also provides an event-driven interceptor system for custom -logging, error tracking, and observability. - -This document covers the Ruby SDK metrics emitted by `MetricsCollector`. It -does not cover Conductor server metrics or metrics emitted by other SDKs. - -## Table of Contents - -- [Quick Start](#quick-start) -- [Metrics Catalog](#metrics-catalog) -- [Metrics Not Applicable to Ruby](#metrics-not-applicable-to-ruby) -- [Labels](#labels) -- [Prometheus Integration](#prometheus-integration) -- [Custom Metrics Backends](#custom-metrics-backends) -- [Troubleshooting](#troubleshooting) -- [Interceptor System](#interceptor-system) -- [Event Types](#event-types) -- [Creating Custom Interceptors](#creating-custom-interceptors) -- [Advanced Use Cases](#advanced-use-cases) -- [Best Practices](#best-practices) -- [Reference](#reference) - ---- - -### Payload Size Metrics - -Recording `workflow_input_size_bytes` requires JSON-serializing the workflow -input to measure its byte size. This is enabled by default. Override explicitly -via the factory: - -```ruby -# Skip the JSON serialization for large payloads -metrics = MetricsCollector.create(backend: :prometheus, measure_payload_size: false) -``` - -### Collector Lifecycle - -`MetricsCollector` responds to `stop`. Call `stop` to unsubscribe from -process-wide dispatchers (e.g. the `GlobalDispatcher` used for HTTP metrics). -`TaskHandler#stop` calls `stop` on all registered event listeners -automatically. If you manage a collector outside of `TaskHandler`, call `stop` -when the collector is no longer needed. - ---- - -## Quick Start - -```ruby -require 'conductor' - -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) - -# Start metrics HTTP server -metrics_server = Conductor::Worker::Telemetry::MetricsServer.new(port: 9090) -metrics_server.start - -# Create handler with metrics -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [metrics] -) - -handler.start -handler.join - -# Cleanup -metrics_server.stop -``` - ---- - -## Metrics Catalog - -Timing values are seconds. Size values are bytes. Label names use camelCase. -Metric definitions are pre-registered when the Prometheus backend is -created, but specific label-combination time series only appear in -`/metrics` after the corresponding event first records them. - -### Counters - -| Metric | Labels | Description | -|---|---|---| -| `task_poll_total` | `taskType` | Incremented each time the worker issues a poll request. | -| `task_execution_started_total` | `taskType` | Incremented when a polled task is dispatched to the worker function. | -| `task_poll_error_total` | `taskType`, `exception` | Incremented when a poll request fails client-side. | -| `task_execute_error_total` | `taskType`, `exception` | Incremented when the worker function throws. | -| `task_update_error_total` | `taskType`, `exception` | Incremented when updating the task result fails. | -| `task_paused_total` | `taskType` | Incremented when a worker is paused and skips acting on a poll. | -| `thread_uncaught_exceptions_total` | `exception` | Incremented when a worker thread raises an uncaught exception. | -| `workflow_start_error_total` | `workflowType`, `exception` | Incremented when starting a workflow fails client-side. | - -### Time Histograms - -All time histograms use buckets (in seconds): - -```text -0.001, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10 -``` - -| Metric | Labels | Description | -|---|---|---| -| `task_poll_time_seconds` | `taskType`, `status` | Poll request latency. `status` is `SUCCESS` or `FAILURE`. | -| `task_execute_time_seconds` | `taskType`, `status` | Worker function execution duration. `status` is `SUCCESS` or `FAILURE`. | -| `task_update_time_seconds` | `taskType`, `status` | Task-result update latency. `status` is `SUCCESS` or `FAILURE`. | -| `http_api_client_request_seconds` | `method`, `uri`, `status` | HTTP API client request latency. `status` is the HTTP status code as a string, or `"0"` on network failure. | - -Each histogram exposes Prometheus series such as: - -```prometheus -task_execute_time_seconds_bucket{taskType="my_task",status="SUCCESS",le="0.1"} 42.0 -task_execute_time_seconds_count{taskType="my_task",status="SUCCESS"} 50.0 -task_execute_time_seconds_sum{taskType="my_task",status="SUCCESS"} 2.3 -``` - -### Size Histograms - -All size histograms use buckets (in bytes): - -```text -100, 1000, 10000, 100000, 1000000, 10000000 -``` - -| Metric | Labels | Description | -|---|---|---| -| `task_result_size_bytes` | `taskType` | Serialized task result output size. | -| `workflow_input_size_bytes` | `workflowType`, `version` | Serialized workflow input size. `version` is an empty string when the workflow version is absent. | - -### Gauges - -| Metric | Labels | Description | -|---|---|---| -| `active_workers` | `taskType` | Current number of worker threads actively executing tasks. | - ---- - -## Metrics Not Applicable to Ruby - -The cross-SDK canonical catalog defines additional metrics that are not -applicable to the Ruby SDK's runtime model: - -| Canonical metric | Why N/A for Ruby | -|---|---| -| `task_ack_error_total` | Batch-poll response is the ack; there is no separate ack call. | -| `task_ack_failed_total` | Same reason. | -| `task_execution_queue_full_total` | Ruby uses `fallback_policy: :caller_runs` which back-pressures the polling thread instead of rejecting tasks. | -| `worker_restart_total` | Python-only. Its multi-process supervisor restarts child processes. Ruby uses threads, fibers, or ractors. | -| `external_payload_used_total` | Ruby SDK has no external-payload-storage integration. | - -Users cross-referencing the harmonization spec or documentation from other -Conductor SDKs may notice these metrics in other catalogs. Their absence in -the Ruby SDK is intentional. - -### `thread_uncaught_exceptions_total` - -The `thread_uncaught_exceptions_total` counter and its collector handler exist -in the Ruby SDK for API completeness, but the metric is not currently wired to -any runtime event. The harmonization spec defines it as "incremented when a -worker thread terminates with an uncaught exception." In SDKs that implement it -(Java, Go, Rust), it fires only at the thread/goroutine death boundary. Python -and JavaScript also define the metric surface but do not wire it. Future Ruby -SDK versions may connect it to `Thread.report_on_exception` or a similar -mechanism. - -### Ractor Runner Limitations (Work-in-Progress -- Untested) - -> **Warning:** The `RactorTaskRunner` is an experimental work-in-progress. -> Metrics collection from Ractor-based workers is **partially implemented -> and not yet functional**. In practice, **no metrics are delivered** from Ractor workers -> in the current implementation due to the incomplete event bridge described -> below. Do not rely on Ractor worker metrics in production. - -The event communication channel between Ractor workers and the main-thread -`SyncEventDispatcher` is not yet implemented. The `TaskHandler` method -`create_event_receiver_ractor` returns `nil` (with the comment _"Ractor -event communication needs more work"_), and the spawned event-receiver thread -exits immediately without forwarding events. Because `RactorTaskRunner` -receives `nil` as its `event_queue`, the `publish_event` guard -(`return unless @event_queue`) causes **every event to be silently -discarded** -- poll events, execution events, update events, and failure -events alike. As a result, none of the metrics documented in this guide -(counters, histograms, or gauges) will be populated for Ractor-based -workers. - -Additionally: - -- The `active_workers` gauge is **never emitted** by `RactorTaskRunner`. - Unlike `TaskRunner` and `FiberTaskRunner`, which call - `publish_active_workers` when tasks start and complete, the Ractor runner - has no equivalent tracking or publishing logic. -- The `cleanup` method is a no-op stub. -- Ractor shutdown in `TaskHandler#stop` is rudimentary (`ractor.take` with - rescued errors) because Ractors lack a clean shutdown mechanism. - -These limitations will be resolved in a future release when proper -Ractor-to-main-thread event forwarding is implemented. Until then, use -`:thread` isolation (the default) or `:fiber` execution if you need metrics -and observability. - ---- - -## Labels - -| Label | Used by | Values | -|---|---|---| -| `taskType` | Worker metrics | Task definition name. | -| `workflowType` | Workflow metrics | Workflow definition name. | -| `version` | `workflow_input_size_bytes` | Workflow version as a string. Empty string when the version is absent. | -| `status` | Task time metrics | `SUCCESS` or `FAILURE`. For `http_api_client_request_seconds`, the HTTP status code as a string (e.g. `"200"`), or `"0"` on network failure. | -| `exception` | Error counters | Exception class name, such as `Faraday::TimeoutError`. | -| `method` | HTTP metrics | HTTP verb (`GET`, `POST`, etc.). | -| `uri` | HTTP metrics | API-relative path template (e.g. `/tasks/poll/batch/{taskType}`). Dynamic path segments retain `{placeholder}` tokens so label cardinality is bounded. | - ---- - -## Prometheus Integration - -### Setup - -Add the `prometheus-client` gem to your Gemfile: - -```ruby -gem 'prometheus-client', '~> 4.0' -``` - -### Basic Usage - -```ruby -require 'conductor' - -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) - -metrics_server = Conductor::Worker::Telemetry::MetricsServer.new(port: 9090) -metrics_server.start - -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [metrics] -) - -handler.start -handler.join - -metrics_server.stop -``` - -### Metrics Endpoints - -The MetricsServer exposes: - -- `GET /metrics` - Prometheus metrics in text format -- `GET /health` - Health check endpoint (`{"status":"healthy"}`) - -### Kubernetes Integration - -```yaml -# Pod annotations for Prometheus scraping -metadata: - annotations: - prometheus.io/scrape: "true" - prometheus.io/port: "9090" - prometheus.io/path: "/metrics" -``` - -### Custom Prometheus Registry - -```ruby -require 'prometheus/client' - -registry = Prometheus::Client::Registry.new -backend = Conductor::Worker::Telemetry::PrometheusBackend.new(registry: registry) -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: backend) -``` - ---- - -## Custom Metrics Backends - -You can create a custom backend by implementing three methods: - -```ruby -class DatadogBackend - def initialize(statsd_client) - @statsd = statsd_client - end - - def increment(name, labels: {}) - tags = labels.map { |k, v| "#{k}:#{v}" } - @statsd.increment(name, tags: tags) - end - - def observe(name, value, labels: {}) - tags = labels.map { |k, v| "#{k}:#{v}" } - @statsd.histogram(name, value, tags: tags) - end - - def set(name, value, labels: {}) - tags = labels.map { |k, v| "#{k}:#{v}" } - @statsd.gauge(name, value, tags: tags) - end -end - -# Use with MetricsCollector -require 'datadog/statsd' -statsd = Datadog::Statsd.new('localhost', 8125) -metrics = Conductor::Worker::Telemetry::MetricsCollector.create( - backend: DatadogBackend.new(statsd) -) -``` - ---- - -## Troubleshooting - -### Metrics Are Empty - -- Verify that `MetricsCollector.create` is called and the collector is passed - to `TaskHandler` via `event_listeners:`. -- Verify workers have polled or executed tasks. Metrics are created lazily when - the relevant event occurs. -- Confirm the scrape endpoint is reachable at the expected host and port. - -### Missing HTTP or Workflow Metrics - -- `http_api_client_request_seconds` requires the collector to be subscribed to - `GlobalDispatcher` for `HttpApiRequest` events from the HTTP layer. This - happens automatically when `subscribe_global_http: true` (the default). -- `workflow_input_size_bytes` and `workflow_start_error_total` only record when - the corresponding `WorkflowExecutor` events fire. - -### High Cardinality - -- The `uri` label on `http_api_client_request_seconds` uses path templates - (e.g. `/workflow/{workflowId}`) to keep cardinality bounded. If you see - fully-resolved paths in your metrics, verify that HTTP requests are going - through the SDK's `ApiClient` rather than a standalone `RestClient`. -- The `exception` label uses exception class names to keep cardinality bounded. -- Avoid embedding user identifiers or unbounded values in task type, workflow - type, or other label values. - ---- - -## Interceptor System - -The Conductor Ruby SDK provides an event-driven interceptor system that allows -you to: - -- **Monitor performance** - Track polling times, execution durations, error rates -- **Implement custom logging** - Add structured logging for task execution -- **Track errors** - Send failures to error tracking services (Sentry, Bugsnag, etc.) -- **Collect metrics** - Export to Prometheus, Datadog, or custom backends -- **Build alerting** - Monitor SLAs and trigger alerts on violations - -``` -TaskRunner - │ - │ publishes events - ▼ -SyncEventDispatcher ──────► Listener 1 (MetricsCollector) - ──────► Listener 2 (LoggingInterceptor) - ──────► Listener 3 (SentryInterceptor) -``` - -When a worker polls for tasks, executes them, or encounters errors, events are -published to all registered listeners. Listeners can then process these events -independently. - ---- - -## Event Types - -The SDK publishes events during worker execution. Any object that responds to -the corresponding `on_*` method can listen for these events. - -### Poll Events - -| Event | When Published | Key Attributes | -|---|---|---| -| `PollStarted` | Before polling for tasks | `task_type`, `worker_id`, `poll_count` | -| `PollCompleted` | After successful poll | `task_type`, `duration_ms`, `tasks_received` | -| `PollFailure` | When poll fails | `task_type`, `duration_ms`, `cause` | - -### Execution Events - -| Event | When Published | Key Attributes | -|---|---|---| -| `TaskExecutionStarted` | Before task execution | `task_type`, `task_id`, `worker_id`, `workflow_instance_id` | -| `TaskExecutionCompleted` | After successful execution | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `duration_ms`, `output_size_bytes` | -| `TaskExecutionFailure` | When execution fails | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `duration_ms`, `cause`, `is_retryable` | - -### Update Events - -| Event | When Published | Key Attributes | -|---|---|---| -| `TaskUpdateCompleted` | After successful result update | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `duration_ms` | -| `TaskUpdateFailure` | When result update fails after all retries | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `cause`, `retry_count`, `task_result`, `duration_ms` | - -### Worker State Events - -| Event | When Published | Key Attributes | -|---|---|---| -| `TaskPaused` | When a paused worker skips a poll | `task_type` | -| `ThreadUncaughtException` | When a worker thread raises an uncaught exception | `cause`, `task_type` | -| `ActiveWorkersChanged` | When the active worker count changes | `task_type`, `count` | - -### Workflow Events - -| Event | When Published | Key Attributes | -|---|---|---| -| `WorkflowStartError` | When starting a workflow fails client-side | `workflow_type`, `version`, `cause` | -| `WorkflowInputSize` | When a workflow is started | `workflow_type`, `version`, `size_bytes` | - -### HTTP Events - -| Event | When Published | Key Attributes | -|---|---|---| -| `HttpApiRequest` | After every HTTP API client request | `method`, `uri`, `status`, `duration_ms` | - -**Important**: `TaskUpdateFailure` is a critical event indicating that a task -result was lost. You should always handle this event to prevent silent data -loss. - ---- - -## Creating Custom Interceptors - -### Basic Structure - -An interceptor is any object that responds to one or more `on_*` methods: - -```ruby -class MyInterceptor - def on_poll_started(event) - # Called before each poll - end - - def on_poll_completed(event) - # Called after successful poll - end - - def on_poll_failure(event) - # Called when poll fails - end - - def on_task_execution_started(event) - # Called before task execution - end - - def on_task_execution_completed(event) - # Called after successful execution - end - - def on_task_execution_failure(event) - # Called when execution fails - end - - def on_task_update_failure(event) - # Called when result update fails (CRITICAL) - end -end -``` - -### Error Tracking Interceptor (Sentry) - -```ruby -require 'sentry-ruby' - -class SentryInterceptor - def on_task_execution_failure(event) - Sentry.capture_exception(event.cause, extra: { - task_id: event.task_id, - task_type: event.task_type, - workflow_instance_id: event.workflow_instance_id, - duration_ms: event.duration_ms, - is_retryable: event.is_retryable - }) - end - - def on_task_update_failure(event) - Sentry.capture_message( - "CRITICAL: Task result lost after #{event.retry_count} retries", - level: :fatal, - extra: { - task_id: event.task_id, - task_type: event.task_type, - workflow_instance_id: event.workflow_instance_id - } - ) - end -end -``` - -### Structured Logging Interceptor - -```ruby -require 'json' - -class StructuredLoggingInterceptor - def initialize(output: $stdout) - @output = output - end - - def on_task_execution_started(event) - log('task_started', event) - end - - def on_task_execution_completed(event) - log('task_completed', event, duration_ms: event.duration_ms) - end - - def on_task_execution_failure(event) - log('task_failed', event, - duration_ms: event.duration_ms, - error: event.cause.class.name, - message: event.cause.message, - retryable: event.is_retryable) - end - - private - - def log(action, event, extra = {}) - entry = { - timestamp: event.timestamp.iso8601(3), - action: action, - task_type: event.task_type, - task_id: event.task_id, - worker_id: event.worker_id, - workflow_instance_id: event.workflow_instance_id, - **extra - } - @output.puts(entry.to_json) - end -end -``` - ---- - -## Advanced Use Cases - -### SLA Monitor - -Alert when tasks exceed duration thresholds: - -```ruby -class SLAMonitor - def initialize(thresholds:, alerter:) - @thresholds = thresholds # { 'task_type' => max_ms } - @alerter = alerter - end - - def on_task_execution_completed(event) - threshold = @thresholds[event.task_type] - return unless threshold && event.duration_ms > threshold - - @alerter.alert( - type: :sla_violation, - task_type: event.task_type, - task_id: event.task_id, - duration_ms: event.duration_ms, - threshold_ms: threshold - ) - end -end - -# Usage -sla_monitor = SLAMonitor.new( - thresholds: { - 'process_order' => 5000, - 'send_email' => 2000 - }, - alerter: SlackAlerter.new(webhook_url: ENV['SLACK_WEBHOOK']) -) -``` - -### Cost Tracker - -Track compute costs per task type: - -```ruby -class CostTracker - def initialize(cost_per_ms: {}) - @cost_per_ms = cost_per_ms - @costs = Hash.new(0.0) - @mutex = Mutex.new - end - - def on_task_execution_completed(event) - rate = @cost_per_ms[event.task_type] || 0.0001 - cost = rate * event.duration_ms - @mutex.synchronize { @costs[event.task_type] += cost } - end - - def report - @mutex.synchronize { @costs.dup } - end -end -``` - -### Retry Tracker - -Monitor retry patterns: - -```ruby -class RetryTracker - def initialize - @retries = Hash.new { |h, k| h[k] = [] } - @mutex = Mutex.new - end - - def on_task_execution_failure(event) - return unless event.is_retryable - - @mutex.synchronize do - @retries[event.task_type] << { - task_id: event.task_id, - error: event.cause.class.name, - timestamp: event.timestamp - } - end - end - - def retry_rate(task_type, window_seconds: 300) - @mutex.synchronize do - cutoff = Time.now - window_seconds - recent = @retries[task_type].select { |r| r[:timestamp] > cutoff } - recent.size - end - end -end -``` - ---- - -## Best Practices - -### 1. Keep Interceptors Fast - -Interceptors run synchronously in the worker thread. Keep processing fast to -avoid impacting task execution: - -```ruby -# BAD: Slow synchronous HTTP call -def on_task_execution_completed(event) - HTTParty.post('https://analytics.example.com', body: event.to_h.to_json) -end - -# GOOD: Queue for background processing -def on_task_execution_completed(event) - @queue << event.to_h -end -``` - -### 2. Handle Errors in Interceptors - -Errors in interceptors are caught and logged but don't affect other -interceptors: - -```ruby -def on_task_execution_completed(event) - external_service.track(event) -rescue => e - @logger.warn("Failed to track: #{e.message}") -end -``` - -### 3. Always Handle TaskUpdateFailure - -This is a critical event -- task results are lost: - -```ruby -def on_task_update_failure(event) - @logger.fatal("Task result lost: #{event.task_id}") - PagerDuty.trigger(severity: :critical, summary: "Task result lost: #{event.task_id}") - FailedTaskStore.save(event.task_result) -end -``` - -### 4. Use Multiple Interceptors - -Separate concerns into different interceptors: - -```ruby -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [ - Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus), - StructuredLoggingInterceptor.new, - SentryInterceptor.new, - SLAMonitor.new(thresholds: sla_config), - ] -) -``` - -### 5. Test Your Interceptors - -```ruby -RSpec.describe SentryInterceptor do - let(:interceptor) { described_class.new } - - describe '#on_task_execution_failure' do - it 'captures exception in Sentry' do - event = Conductor::Worker::Events::TaskExecutionFailure.new( - task_type: 'my_task', - task_id: 'task-123', - worker_id: 'worker-1', - workflow_instance_id: 'workflow-456', - duration_ms: 100, - cause: StandardError.new('Test error'), - is_retryable: true - ) - - expect(Sentry).to receive(:capture_exception) - interceptor.on_task_execution_failure(event) - end - end -end -``` - ---- - -## Reference - -### Telemetry Classes - -- `Conductor::Worker::Telemetry::MetricsCollector` - Canonical metric collector; `.create` is a convenience factory -- `Conductor::Worker::Telemetry::NullBackend` - No-op metrics backend -- `Conductor::Worker::Telemetry::PrometheusBackend` - Prometheus backend with canonical label schemas -- `Conductor::Worker::Telemetry::MetricsServer` - WEBrick HTTP server for `/metrics` and `/health` endpoints - -### Event Classes - -All events are in the `Conductor::Worker::Events` namespace: - -- `PollStarted`, `PollCompleted`, `PollFailure` -- `TaskExecutionStarted`, `TaskExecutionCompleted`, `TaskExecutionFailure` -- `TaskUpdateCompleted`, `TaskUpdateFailure` -- `TaskPaused`, `ThreadUncaughtException`, `ActiveWorkersChanged` -- `WorkflowStartError`, `WorkflowInputSize` -- `HttpApiRequest` - -### Registration Methods - -Register interceptors via `TaskHandler`: - -```ruby -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [listener1, listener2] -) -``` - -Or register manually with the event dispatcher: - -```ruby -dispatcher = handler.event_dispatcher -dispatcher.register(Conductor::Worker::Events::PollStarted, ->(event) { puts event }) -``` - ---- - -## Detailed Technical Notes -- [Unreleased] - -### Path template `uri` label (metric_uri) - -The `uri` label on `http_api_client_request_seconds` now carries the **path -template** (e.g. `/workflow/{workflowId}`) rather than the fully-resolved -request path (e.g. `/api/workflow/abc-123-def`). This keeps metric label -cardinality bounded regardless of how many unique workflow IDs, task types, or -other dynamic path segments pass through the SDK. - -**Data flow:** - -1. Every API resource method calls `ApiClient#call_api` with a `resource_path` - that contains `{placeholder}` tokens (e.g. `/tasks/poll/batch/{taskType}`). -2. `ApiClient#call_api_no_retry` saves `resource_path` as `metric_uri` *before* - substituting path parameters. -3. The substituted path is concatenated with `server_url` to form the HTTP URL. - `metric_uri` is passed alongside as a keyword argument to - `RestClient#request`. -4. `RestClient#emit_http_event` prefers `metric_uri` when present; it only - falls back to extracting the path from the full URL when `metric_uri` is - `nil` (e.g. for direct `RestClient` calls outside of `ApiClient`). -5. The `HttpApiRequest` event carries the template string as its `uri` field. - `MetricsCollector#on_http_api_request` records it as-is into the histogram. - -This approach mirrors the Python SDK (`metric_uri` parameter), the Java SDK -(`PathTemplateTag` on the OkHttp request), and the Go SDK (`WithPathTemplate` -context value). The base-URL path prefix (e.g. `/api`) is never included -because `metric_uri` is always the raw API-relative resource path. - -### MetricsCollector factory - -`MetricsCollector.create(backend:)` is a convenience constructor that returns -a `MetricsCollector` instance. It emits the harmonized cross-SDK catalog with -`camelCase` domain labels, Prometheus histograms with explicit bucket -boundaries, and the `exception` label derived from the Ruby exception class -name. - -### Event system - -The following event types are used by the collector: - -- `PollStarted`, `PollCompleted`, `PollFailure`, `TaskExecutionStarted`, - `TaskExecutionCompleted`, `TaskExecutionFailure`, `TaskUpdateCompleted`, - `TaskUpdateFailure`, `TaskPaused`, `ActiveWorkersChanged` -- emitted by - `TaskRunner` (and `FiberTaskRunner`). `RactorTaskRunner` constructs these - same events internally but **does not deliver them** to the - `SyncEventDispatcher` because the Ractor-to-main-thread event bridge is - not yet implemented (see - [Ractor Runner Limitations](#ractor-runner-limitations-work-in-progress----untested)). -- `WorkflowStartError`, `WorkflowInputSize` -- emitted by - `WorkflowExecutor`. -- `HttpApiRequest` -- emitted by `RestClient` via the process-wide - `GlobalDispatcher`, but only when at least one `HttpApiRequest` listener - is subscribed (i.e. a `MetricsCollector` is active). With no collector, - `RestClient` skips all timing overhead. -- `ThreadUncaughtException` -- event class and collector handler exist for - API completeness but are not currently emitted by any runner (see - [thread_uncaught_exceptions_total](#thread_uncaught_exceptions_total)). - -All events flow through the `SyncEventDispatcher` -> listener registry -> -collector pattern. The `GlobalDispatcher` singleton provides a secondary -channel so that `RestClient` (which has no direct reference to the task -handler's dispatcher) can still emit HTTP events. diff --git a/docs/agents/README.md b/docs/agents/README.md deleted file mode 100644 index fc87009..0000000 --- a/docs/agents/README.md +++ /dev/null @@ -1,55 +0,0 @@ -# Agents - -Define an agent in Ruby, run it on a Conductor server. Same `agentConfig` on the wire as the -Python SDK; the server compiles and runs the loop, this process runs your tools. - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } -end - -agent = Agent.new( - name: 'weather', - model: 'openai/gpt-4o', - instructions: 'Answer weather questions.' -) -agent.add_tool :get_weather - -puts agent.call_sync('Weather in Lisbon?') -``` - -Run it: `CONDUCTOR_SERVER_URL=http://localhost:8080/api ruby weather.rb`. Server URL and auth -come from the usual `CONDUCTOR_*` variables; nothing else to configure. - -| Guide | What it covers | -|---|---| -| [Tools](concepts/tools.md) | `tool def`, types from keyword defaults, `describe`, `requires_approval`, modules, RubyLLM tools, server-side tools | -| [Calling an agent](concepts/streaming-hitl.md) | `call_sync`, `call_async`, `Execution`, approval with `on_approval` | -| [Teams](concepts/teams.md) | `add_agent`, `hands_off_to`, strategies, `redact`, `stop_when`, `stop_after` | -| [Secrets](concepts/secrets.md) | `secret()`, `secrets_env()`, how names reach the server and values reach the tool | -| [Runtime and deployment](concepts/runtime.md) | `AgentRuntime`, `deploy`, `serve`, `CONDUCTOR_AGENT_*` settings, what runs where | - -Examples: [19 Python ports and playback instructions](../../examples/agents/README.md), plus -`weather.rb`, `support_approval.rb`, and `bug_desk.rb`. Contract tests cover 20 golden agents -and the actual configurations built by all 19 ports. Integration tests execute the example -files themselves against Conductor OSS and the shared LLM recordings. -See the [parity audit](../design/AGENTS_PARITY_AUDIT.md) for scope and evidence. - -## Requirements - -- A Conductor server with the agent runtime enabled (`conductor.integrations.ai.enabled=true`, - the default) and an LLM integration configured. On Orkes the left side of - `model: 'openai/gpt-4o'` is the integration name; on OSS it is the provider key whose API key - the server reads from `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` / ... -- Secrets delivered to tools (`secret('GH_TOKEN')`) need conductor-oss 3.32.0-rc.8 or later - (`runtimeMetadata`). Older servers still work: `secret()` falls back to `ENV`. - -## What is not ported - -Framework agents (OpenAI Agents SDK, LangGraph, Google ADK, Claude Agent SDK), skills, -local code execution and CLI tools, schedules, semantic memory. See -`docs/design/AGENTS_IMPLEMENTATION_PLAN.md` for the full list and the decisions behind the -port. diff --git a/docs/agents/concepts/runtime.md b/docs/agents/concepts/runtime.md deleted file mode 100644 index ddabbed..0000000 --- a/docs/agents/concepts/runtime.md +++ /dev/null @@ -1,74 +0,0 @@ -# Runtime and deployment - -`Agent#call_sync` / `#call_async` use `Conductor::Agents.runtime`, an `AgentRuntime` built -from the environment (`CONDUCTOR_SERVER_URL`, `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET`). - -```ruby -Conductor::Agents.configure( - configuration: Conductor::Configuration.new(server_api_url: 'https://play.orkes.io/api'), - agent_config: Conductor::Agents::AgentConfig.new(worker_thread_count: 4) -) - -runtime = Conductor::Agents::AgentRuntime.new(configuration: config) # or your own instance -runtime.call_sync(agent, 'hi') -runtime.compile(agent) # { "workflowDef", "requiredWorkers" } without registering -runtime.deploy(a, b) # register on the server; returns names -runtime.serve(a, b) # deploy + run the tool workers until Ctrl-C -runtime.shutdown # stop workers and streams -Conductor::Agents.shutdown # same for the default runtime -``` - -## Settings - -| Variable | Default | Meaning | -|---|---|---| -| `CONDUCTOR_AGENT_WORKER_POLL_INTERVAL` | `100` | tool worker poll interval (ms) | -| `CONDUCTOR_AGENT_WORKER_THREADS` | `1` | concurrent tasks per tool worker | -| `CONDUCTOR_AGENT_STREAMING_ENABLED` | `true` | use SSE; `false` polls `/agent/{id}/status` | -| `CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER` | `false` | reserved (Python parity), not used yet | - -Booleans accept `true/1/yes/on` and `false/0/no/off`. - -## What runs where - -| Piece | Where | -|---|---| -| LLM loop, tool routing, guardrail chain, approval pause, MCP discovery | server | -| `tool def` tools, RubyLLM tools | this process (Conductor workers, one per tool name) | -| `_termination` (`stop_when`, `stop_after`, `termination:`) | this process | -| custom guardrails (with a block), callbacks, `on_condition` handoffs | this process | -| regex / LLM guardrails, http / mcp / human / agent tools | server | - -Every worker registers its TaskDef with the Python SDK's defaults: `retryCount 2`, -`retryDelaySeconds 2`, `retryLogic LINEAR_BACKOFF`, `timeoutSeconds 0`, -`responseTimeoutSeconds 10`, `timeoutPolicy RETRY`, `runtimeMetadata` = declared secret names. -Task results go to `POST /api/tasks/update-v2` (with a one-time fallback to `POST /api/tasks`). -While a tool or system worker runs, the SDK renews its lease at 80% of the task's -`responseTimeoutSeconds` (every 8 seconds for the default timeout). Renewals use -`POST /api/tasks` with `extendLease: true` and stop before the final result is sent. -Tasks with no positive response timeout do not need renewal. Agent workers enable -`lease_extend_enabled` by default; other workers can opt in with that option. - -An agent execution is a Conductor workflow: `execution_id` is the workflow id, and the -workflow, task and prompt data are visible in the Conductor UI like any other run. - -## Stateful runs - -`Agent.new(..., stateful: true)` (or a stateful tool) sends a `runId` with the start request. -The server maps every required worker to that task domain and the SDK polls with it, so each -run's tasks reach the process that started it. - -## Testing - -Contract tests (`spec/conductor/agents/contract_spec.rb`) need no server: every serialized -config must equal the Python SDK's golden file and validate against `agent-schema.json`. - -Runtime tests (`spec/agents/`) replay scenarios recorded from a real server with WireMock -(`conductor-oss/conductor-mocks`): - -```bash -docker run -d -p 8080:8080 -v $PWD/../conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock wiremock/wiremock:3x -CONDUCTOR_AGENTS_REPLAY_URL=http://localhost:8080 bundle exec rspec spec/agents -``` - -A request the recorded server never saw fails the run. diff --git a/docs/agents/concepts/secrets.md b/docs/agents/concepts/secrets.md deleted file mode 100644 index d2c00b7..0000000 --- a/docs/agents/concepts/secrets.md +++ /dev/null @@ -1,72 +0,0 @@ -# Secrets - -Three kinds. Three different owners. None of them live in your code. - -| Secret | Held by | You write | -|---|---|---| -| Conductor auth | your env | `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET` (existing `Configuration`; the token is cached per instance) | -| LLM provider key | Conductor server, as an Integration | `model: 'openai/gpt-4o'` | -| Tool credential | Conductor server, in the secret store | `secret('GH_TOKEN')` in the tool body | - -## Tool credentials - -```ruby -tool def create_issue(title: String, body: '') - Github.create_issue(title, body, token: secret('GH_TOKEN')) -end -``` - -`secret('GH_TOKEN')` is both the read and the declaration. At `tool def` the SDK scans the -method body for literal `secret('...')` and `secrets_env('...')` names and puts them on the -tool's contract: `TaskDef.runtimeMetadata` (names) when the worker registers, and -`tool.config.credentials` in the `agentConfig`. The server resolves the values from its secret -store when your worker polls and attaches them to the task (`Task.runtimeMetadata`, wire-only, -never persisted). Inside the tool, `secret()` reads that map through the task context. - -Store the value once on the server: Orkes UI or `SecretClient#put_secret`; on OSS, -`CONDUCTOR_SECRET_GH_TOKEN` in the server's environment. - -```mermaid -sequenceDiagram - participant SDK as Ruby SDK - participant Server as Conductor Server - participant Store as Secret Store - - Note over SDK: at startup - SDK->>Server: register TaskDef create_issue, runtimeMetadata: [GH_TOKEN] - - Note over Server: LLM calls create_issue - Server->>Store: get GH_TOKEN - Store-->>Server: ghp_... - Server->>+SDK: task create_issue, runtimeMetadata: GH_TOKEN=ghp_... - SDK->>SDK: create_issue runs, secret('GH_TOKEN') -> ghp_... - SDK-->>-Server: result (runtimeMetadata dropped) -``` - -## When the name is not a literal - -```ruby -filer.add_tool :create_issue, credentials: ['GH_TOKEN'] # this tool -tool_credentials :create_issue, 'GH_TOKEN' # same, at definition time -filer = Agent.new(..., credentials: ['GH_TOKEN']) # everything under this agent -``` - -## Subprocesses - -Workers are threads in one process, so the SDK never writes `ENV` (that would leak the secret -into every other tool running at the same time). For `system` / `spawn` / `Open3`: - -```ruby -tool def gh_create_issue(title: String) - system(secrets_env('GH_TOKEN'), 'gh', 'issue', 'create', '--title', title) -end -``` - -`secrets_env('GH_TOKEN')` is `{ 'GH_TOKEN' => 'ghp_...' }` for this call and declares the same -way. - -## No secret store on the server? - -`secret('X')` falls back to `ENV['X']`, then raises `CredentialNotFoundError` (the task fails -terminally with a message naming the key). Servers older than conductor-oss 3.32.0-rc.8 do -not send `runtimeMetadata`; the ENV fallback keeps tools working there. diff --git a/docs/agents/concepts/streaming-hitl.md b/docs/agents/concepts/streaming-hitl.md deleted file mode 100644 index 8f33895..0000000 --- a/docs/agents/concepts/streaming-hitl.md +++ /dev/null @@ -1,93 +0,0 @@ -# Calling an agent: sync, async, approval - -## Sync - -```ruby -answer = agent.call_sync('What is your return policy?') -``` - -Blocks until the agent is finished and returns the answer as a String. Pass `timeout:` seconds -to give up (`Timeout::Error`). A failed, cancelled or timed-out execution raises -`Conductor::Agents::Error`. - -## Async with a callback - -```ruby -agent.call_async('What is your return policy?') do |answer, execution| - Mailer.send(customer, answer) -end -``` - -Returns immediately. The block runs on the stream thread when the agent finishes (`answer` is -nil when it failed; check `execution.finish_reason`). Exceptions in the block are logged, never -raised into the stream. - -## Async, poll it yourself - -```ruby -execution = agent.call_async('What is your return policy?') - -execution.done? # false until finished -execution.waiting? # true while a tool waits for approval -execution.partial_text # text streamed so far -execution.result # blocks until done, returns the answer -execution.finish_reason # :stop | :tool_calls | :length | :content_filter | :rejected | :error | :cancelled | :timeout -execution.tool_calls # [#] with .result once known -execution.token_usage # prompt / completion / total, summed over sub-agents -execution.execution_id # also the Conductor workflow id; Execution.find(id) later -execution.pause; execution.resume; execution.cancel; execution.stop; execution.signal('hurry up') -``` - -## Approval - -```ruby -tool def issue_refund(order_id: String, amount: Float) - Billing.refund(order_id, amount) -end -requires_approval :issue_refund - -agent.on_approval do |request| - request.amount < 100 ? request.approve : request.reject('Needs a manager') -end -``` - -When the model calls an approval-required tool the server pauses on a HUMAN task and the SDK -receives a `waiting` event. `request` is an `ApprovalRequest`: the tool's arguments are -methods (`request.amount`, `request.order_id`), plus `tool_name`, `arguments`, `tool_calls` -(the server gates the whole batch of tool calls in a turn with one approval), `approve`, -`reject(reason)` and `send_message(text)` for human-input tools. - -The block runs on the stream thread, for `call_sync` and `call_async` alike. Without an -`on_approval` block the request is parked on `execution.pending`; `execution.approve` / -`execution.reject` answer it, or any other client can call `POST /api/agent/{id}/respond`. - -A rejected tool ends the run as COMPLETED with `finish_reason == :rejected`, `result` nil and -the reason on `execution.output['rejectionReason']`. - -## Sessions - -```ruby -agent.call_sync(question, session_id: 'cust-77') # same conversation across calls -``` - -## Fire and forget - -`call_async` with no block, keep the `execution_id`, walk away. If the agent has `tool def` -tools they run in *your* process, so it has to stay up. Fire-and-forget only works when every -tool is server-side (http / mcp / human) or you have `deploy`ed the agent and run `serve` -somewhere else. - -## Under the hood - -`call_async(prompt)`: - -1. `POST /api/agent/start` with the serialized `agentConfig`; the reply lists `requiredWorkers` -2. start a worker for every required task this process can serve (your tools, plus - `_termination`, custom guardrails, callbacks) with the same TaskDef defaults as Python -3. return an `Execution`; open `GET /api/agent/stream/{id}` (SSE) on a background thread, - reconnecting with `Last-Event-ID`; fall back to polling `GET /api/agent/{id}/status` if SSE is - unavailable or `CONDUCTOR_AGENT_STREAMING_ENABLED=false` -4. `tool_call` / `tool_result` events fill `tool_calls`; `waiting` builds an `ApprovalRequest` -5. `done` sets the result and `finish_reason`, fetches token usage, fires the block - -`call_sync` is `call_async(prompt).result`. diff --git a/docs/agents/concepts/teams.md b/docs/agents/concepts/teams.md deleted file mode 100644 index 5ef6465..0000000 --- a/docs/agents/concepts/teams.md +++ /dev/null @@ -1,92 +0,0 @@ -# Teams - -```ruby -tool def create_issue(title: String, body: '') - Github.create_issue(title, body, token: secret('GH_TOKEN')) -end - -triage = Agent.new(name: 'triage', model: 'openai/gpt-4o-mini', - instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.') - -filer = Agent.new(name: 'filer', model: 'anthropic/claude-sonnet-4-5', - instructions: 'File the bug as a GitHub issue.') -filer.add_tool :create_issue - -triage.hands_off_to filer, on: 'ACTIONABLE' - -team = Agent.new(name: 'bug_desk') -team.add_agent triage -team.add_agent filer - -puts team.call_sync(File.read('report.md')) -``` - -| Line | Does | -|---|---| -| `filer.add_tool :create_issue` | Give filer a tool. | -| `triage.hands_off_to filer, on: 'ACTIONABLE'` | When triage's answer contains ACTIONABLE, filer takes over. Pass a block/proc as `on:` for a custom condition. | -| `team.add_agent triage` | Give the team a member. First one added starts. `team.add_agents a, b` for several. | -| `Agent.new(name: 'bug_desk')` | No model needed: a team inherits the first member's model (the server requires one on every agent). | - -`tools:` and `agents:` still work as keyword arguments in `Agent.new`. - -## Strategies - -```ruby -team.strategy = :sequential # one after another -team.strategy = :parallel # all at once, merged -team.strategy = :swarm # members transfer to each other via handoffs -team.strategy = :router # router: agent picks the member -# also :round_robin, :random, :manual (default :handoff) - -pipeline = researcher >> writer >> editor # sequential, named researcher_writer_editor -``` - -When members declare `hands_off_to` and the team has no explicit strategy, the serializer -makes the team a `swarm` and lists the members' handoffs on the team, which is where the -server reads them. Set a strategy explicitly to keep it. - -## Guardrails and stopping - -```ruby -filer.redact %w[password api_key] # scrub these from output before anyone sees it -filer.stop_when 'ISSUE_FILED' # stop on this text -filer.stop_after messages: 12 # or after this many messages -filer.add_guardrail RegexGuardrail.new('\b\d{3}-\d{2}-\d{4}\b', name: 'no_ssn', on_fail: :raise) -filer.add_guardrail LlmGuardrail.new('openai/gpt-4o-mini', 'No medical advice', on_fail: :retry) -filer.add_guardrail Guardrail.new(name: 'no_pii', on_fail: :retry) { |text| !text.include?('SSN') } -``` - -`stop_when` / `stop_after` build `Termination::TextMention` / `Termination::MaxMessage` and -combine with `|`; any `Termination::*` condition (also `StopMessage`, `TokenUsage`, `&`, `|`) -can be set directly with `termination:`. The server evaluates them through a -`_termination` worker this process runs. Custom guardrails with a block also run here; -regex and LLM guardrails run on the server. - -## Callbacks - -```ruby -class Timing < Conductor::Agents::CallbackHandler - def on_model_start(messages: nil, **) = (@t0 = Time.now; nil) - def on_model_end(llm_result: nil, **) = (puts Time.now - @t0; nil) -end -agent.add_callback Timing.new -agent.callback(:before_tool) { |**kw| log kw; nil } -``` - -Positions: `before_agent after_agent before_model after_model before_tool after_tool`. Each -becomes a `_` task the server schedules and this process serves. Return a -non-empty Hash to override; nil to continue. - -## Underneath - -Same Python objects, same `agentConfig`. Sugar only. - -| Sugar | Python-parity object | -|---|---| -| `a.add_tool :x` | appends to `Agent#tools`, same array `tools:` fills | -| `team.add_agent a` | appends to `Agent#agents`, same as `agents:` | -| `a.hands_off_to b, on: 'X'` | `Handoff::OnTextMention.new(target: 'b', text: 'X')` | -| `a.redact %w[...]` | `RegexGuardrail.new(..., position: :output, on_fail: :fix)` | -| `a.stop_when 'X'` / `a.stop_after messages: n` | `Termination::TextMention \| Termination::MaxMessage` | -| `a >> b` | `Agent.new(strategy: :sequential, agents: [a, b])` | diff --git a/docs/agents/concepts/tools.md b/docs/agents/concepts/tools.md deleted file mode 100644 index 5d3e3d0..0000000 --- a/docs/agents/concepts/tools.md +++ /dev/null @@ -1,115 +0,0 @@ -# Tools - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } -end - -agent.add_tool :get_weather -``` - -`tool` receives the Symbol that `def` returns (the same trick as `private def`) and builds a -`ToolDef`: the tool's name is the method name, its description is the humanized name -(`get_weather` -> "Get weather"), its JSON schema comes from the keyword defaults, and its -secrets from `secret('...')` literals in the body. The method stays a normal method, so -`get_weather(city: 'Lisbon')` works in tests. - -## Types - -The default is the type: - -| You write | JSON schema | -|---|---| -| `city: String` | required string | -| `amount: Float` | required number | -| `count: Integer` | required integer | -| `flags: Hash` | required object | -| `tags: [String]` | required array of strings | -| `units: 'metric'` | optional string, default `"metric"` | -| `limit: 10` | optional integer, default `10` | -| `verbose: false` | optional boolean, default `false` | -| `units: %w[metric imperial]` | optional string, one of these, default `"metric"` | -| `extra: nil` | optional, any type | -| `city:` (no default) | required, any type | - -Positional parameters raise at `tool def`. Types are read from the method's AST -(`RubyVM::AbstractSyntaxTree.of`), which works on MRI whenever the source file is on disk. For -methods typed into irb (Ruby 3.2+: set `RubyVM.keep_script_lines = true`) or on other Rubies, -every keyword becomes an untyped property and only keywords without defaults are required. - -Because `city: String` uses the class as the Ruby default, the runtime never calls a tool with a -required argument missing: the task fails with a clear reason instead. - -## Description, approval, options - -```ruby -describe :get_weather, 'Get the current weather for a city.' -requires_approval :issue_refund # a human approves before it runs -tool_credentials :create_issue, 'GH_TOKEN' # when the secret name is not a literal -Weather[:forecast].timeout_seconds = 60 # any ToolDef attribute -tool :ping, description: 'Health check', retry_count: 0 # options on registration -``` - -## Modules - -```ruby -module Weather - extend Conductor::Agents::Tools - - tool def current(city: String) ... end - tool def forecast(city: String, days: 3) ... end -end - -agent.add_tools Weather # both -agent.add_tool Weather[:current] # one -``` - -Inside a class body (an `RSpec.describe` block, for example) `tool def` also works; the method -is bound to a bare instance of the class. - -## RubyLLM tools - -```ruby -class Weather < RubyLLM::Tool - description 'Gets current weather for a location' - param :latitude, type: :number - param :longitude, type: :number - - def execute(latitude:, longitude:) ... end -end - -agent.add_tool Weather # a RubyLLM::Tool class, as-is -``` - -The adapter reads `name`, `description` and `parameters`, and runs `Weather.new.execute(**args)` -as the worker body. RubyLLM is optional and only used when it is loaded. It is not the engine: -Conductor runs the LLM loop server-side with the provider key held as an integration. - -## Server-side tools - -These need no worker in your process: - -```ruby -ToolDef.http('lookup', 'https://api.example.com/orders/${id}', method: 'GET', - headers: { 'Authorization' => 'Bearer ${API_KEY}' }, credentials: ['API_KEY']) -ToolDef.mcp('http://localhost:3001/mcp', tool_names: %w[search fetch]) -ToolDef.human('ask_user', description: 'Ask the user a question and wait for the answer.') -ToolDef.agent(researcher) # another agent as a tool -ToolDef.image('draw', description: '...', llm_provider: 'openai', model: 'dall-e-3') -``` - -`${NAME}` placeholders in headers must be listed in `credentials:`. MCP discovery happens on -the server (`LIST_MCP_TOOLS` before the loop); nothing runs client-side. - -## What the server sends your tool - -The LLM's arguments arrive as top-level task input keys, plus `method`, `_agent_state` and -`_agent_tool_name` which the SDK strips. Strings are coerced to the schema type (`"5"` -> `5`, -`"true"` -> `true`, JSON text -> arrays/objects). A Hash result is returned as-is; anything else -is wrapped as `{ "result" => value }`. A `_state_updates` key in the result is merged into the -agent's durable state by the server. Exceptions fail the task (retried per the tool's -`retry_count`, default 2); missing arguments, missing secrets and unserializable results fail it -terminally. diff --git a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md b/docs/design/AGENTS_IMPLEMENTATION_PLAN.md deleted file mode 100644 index 71d2996..0000000 --- a/docs/design/AGENTS_IMPLEMENTATION_PLAN.md +++ /dev/null @@ -1,661 +0,0 @@ -# Ruby Agents Parity: Implementation Plan - -Status: implementation extended 2026-09-17. Owner: Ruby SDK. - -See [the current parity audit](AGENTS_PARITY_AUDIT.md) for the 19 runnable Python ports, -contract coverage, and real-server playback. The detailed work breakdown below records -the original implementation plan; the status table and scope here supersede its old blockers. - -## Status - -| Slice | State | Notes | -|---|---|---| -| Phase 0 (toolchain, token cache, runtimeMetadata, transport) | done | plus `update-v2` in the runner and `lease_extend_enabled` (2.18) | -| Phase 1 (definition layer, serializer, contract tests) | done | 20 golden contracts plus all 19 example configurations identical to Python, schema-valid | -| Phase 2 (runtime, SSE, dispatch, secrets, system workers) | done | Net::HTTP SSE; polling fallback over `/agent/{id}/status` | -| Phase 3.1 (WireMock replay of `tool_happy_path`) | done | zero unmatched requests; CI job `agents-replay` | -| Phase 3.2 (approval / secrets / team scenarios) | implemented | covered by shared server LLM recordings and real HTTP/MCP services | -| Phase 3.3 (mockLLM functional suite) | implemented | `agents-playback.yml` runs all 19 example files and the feature-branch shared verification action | -| Phase 4 (examples, docs, changelog) | done | Confluence page refresh left to the owner (see 4.1) | - -Decisions taken during implementation that extend section 2: 2.18 (update-v2), 2.19 (swarm -hoisting for `hands_off_to`), 2.20 (run domain applies to every worker). - -Source of truth for *what* we build is the Confluence page -[Ruby Agents Parity Plan](https://orkes.atlassian.net/wiki/spaces/ENG/pages/53739522/Ruby+Agents+Parity+Plan) -and its local mirror `docs/design/AGENTS_PARITY_ONEPAGER.md`, plus the sub-specs -`AGENT_TOOLS_DSL.md`, `AGENT_STREAMING.md`, `AGENT_TEAMS.md`, `AGENT_SECRETS.md`, -`AGENT_TESTING.md`, `SDK_AGENT_TESTING_STRATEGY_V2.md`. This document is *how* and *in what -order*, reconciled against two things the design docs were written before we had verified: - -- the Python SDK agents package (`python-sdk/src/conductor/ai/agents`), the parity source; and -- the server implementation (`conductor/agentspan`, branch `feature/llm_mock_impl`), the wire contract. - -Section 2 lists every place the verified contract changes the design. Everything after that is -the work, split into PR-sized slices with acceptance criteria. - ---- - -## 1. Goal and scope - -Ship `Conductor::Agents` in `conductor_ruby`: define an agent in Ruby, serialize it to the same -`agentConfig` Python sends, start it on the server, run the tool workers locally, stream the -result. The three examples in the one-pager (tools, streaming + approval, team + secret) must run -unchanged against a Conductor server that has `conductor.integrations.ai.enabled=true`. - -### In scope (parity surface) - -| Area | Ruby API (from design docs) | Python source | -|---|---|---| -| Definition | `Agent`, `ToolDef`, `ToolType`, `Tools` DSL (`tool def`, `describe`, `requires_approval`, `[]`), `Guardrail`/`RegexGuardrail`/`LlmGuardrail`, `Termination::*`, `Handoff::*`, `CallbackHandler`, `ConversationMemory`, `PromptTemplate` | `agent.py`, `tool.py`, `guardrail.py`, `termination.py`, `handoff.py`, `callback.py`, `memory.py` | -| Serialization | `ConfigSerializer.serialize(agent)` | `config_serializer.py` | -| Transport | `AgentResourceApi`, `AgentClient`, `OrkesClients#get_agent_client`, `SseClient` | `orkes_agent_client.py` | -| Runtime | `AgentRuntime` (`call_sync`, `call_async`, `deploy`, `serve`, `shutdown`), `AgentConfig.from_env`, `Execution`, `ApprovalRequest`, `ToolRegistry`, `Dispatch`, `Secrets` | `runtime/runtime.py`, `_dispatch.py`, `tool_registry.py`, `config.py`, `result.py` | -| Sugar | `add_tool`/`add_agent`/`hands_off_to`/`redact`/`stop_when`/`stop_after`/`on_approval`, `secret()`/`secrets_env()` auto-declaration | Ruby-only | -| Tests | contract tests (schema + 19 golden configs), runtime tests (replay), CI job | `tests/unit/ai`, `examples/agents/_configs` | -| Docs | README section, `docs/agents/`, three runnable examples, CHANGELOG | `docs/agents/` | - -The implemented surface also includes `plan_execute` with planner/fallback/context, callable -routers, `a >> b`, `prefill_tools`, and structured `output_type`. - -### Out of scope for this plan (Python has them; deliberately deferred) - -Framework agents (`framework`/`rawConfig`: OpenAI Agents, LangGraph, ADK, Claude Agent SDK), -skills, `claude-code` pseudo-provider, -local code execution and CLI tools, OCG retrieval agent, schedules, semantic memory, -`openai_compat.Runner`, OpenTelemetry tracing, liveness monitor / worker restarter, -`scatter_gather`, `@agent`-on-methods (`Agent.from_instance`), -`gate`, `allowed_transitions`, `masked_fields`, `introduction`, -`include_contents`, `thinking_config`, `reasoning_effort`, `context_window_budget`. - -The serializer will be written so any of these can be added as one field + one test later; none -of them changes the architecture. The Confluence page already marks the frameworks row N/A. - ---- - -## 2. Verified contract vs. design docs: what changes - -Each item below is a decision. Where the design docs and the verified contract disagree, the -contract wins and the design docs get a one-line update in Phase 4. - -### 2.1 `agentConfig` wire format (from Python `config_serializer.py`, server `AgentConfig.java`) - -- Keys are camelCase. Every `nil` is dropped. Leaf agents omit `strategy`; it is emitted only - when the agent has sub-agents. `external` and `maxTurns` and `timeoutSeconds` are always - emitted. `approvalRequired`, `stateful`, `enablePlanning` are emitted only when `true`. -- **Credentials asymmetry**: agent-level credentials are top-level `"credentials": [...]`; - tool-level credentials are nested `"config": { "credentials": [...] }`. The server reads - `tool.config.credentials` and falls back to the agent list. -- `strategy` values on the wire are lowercase snake_case: `handoff sequential parallel router - round_robin random swarm manual plan_execute`. Default `handoff`. -- Termination: `{"type": "text_mention"|"stop_message"|"max_message"|"token_usage"|"and"|"or", ...}`. - Handoffs: `{"target", "type": "on_tool_result"|"on_text_mention"|"on_condition", ...}`. - Guardrails: `{"name","position","onFail","maxRetries","guardrailType": "regex"|"llm"|"custom"|"external", ...}`. - Callbacks: `[{"position": "before_model", "taskName": "_before_model"}]`. - Memory: `{"messages": [...], "maxMessages": n}`. Prompt template instructions: - `{"type": "prompt_template", "name", "variables", "version"}`. -- Agent name regex `^[a-zA-Z_][a-zA-Z0-9_-]*$`, `maxTurns >= 1`, duplicate sub-agent names rejected. -- There is **no server-side JSON schema** for `agentConfig`. Python's - `docs/agents/reference/agent-schema.json` (Draft 2020-12, `additionalProperties: false` at the - root) is the contract artifact. We vendor it and the 19 golden configs from - `python-sdk/examples/agents/_configs/`. - -### 2.2 Server requires `model` on every agent, including team parents - -`ModelParser.parse` runs unconditionally on every compile path and raises on a missing or -slash-less model. The one-pager's `team = Agent.new(name: 'bug_desk')` would be rejected. -Python only auto-inherits for `parallel`. - -**Decision D1**: `ConfigSerializer` fills a missing parent `model` from the first sub-agent that -has one, for every multi-agent strategy. A leaf agent with no model raises -`Conductor::Agents::ConfigurationError` at serialize time (unless `external: true`). The -examples stay as written. - -### 2.3 Approval is a `waiting` event plus a `respond` call; `ApprovalRequest` is Ruby-only - -Python has no `ApprovalRequest` and no `on_approval`. On the server, `approvalRequired: true` -on any tool inserts a `HUMAN` task (`_approval_human`, status `IN_PROGRESS`). The SDK -sees SSE `waiting` with `pendingTool = {taskRefName, toolCalls: [{name, args}], response_schema, -...}`. One HUMAN task gates the whole batch of tool calls, so `tool_name`/`parameters` in -`pendingTool` are `null` and `toolCalls` is populated. The reply is -`POST /api/agent/{executionId}/respond` with `{"approved": true}` or -`{"approved": false, "reason": "..."}`. A rejection ends the execution **COMPLETED** with -`output.finishReason == "rejected"` and `output.rejectionReason` set. - -**Decision D2**: `ApprovalRequest` wraps one `waiting` event. It exposes `execution_id`, -`task_ref_name`, `tool_calls` (array of `ToolCall(name:, arguments:)`), and, as sugar, `tool_name` -and argument accessors (`request.amount`) taken from the first tool call. `approve` and -`reject(reason)` post to `respond`. If no `on_approval` block is registered the request is -parked on `execution.pending` and nothing is sent. `finish_reason` becomes `:rejected` from -`output.finishReason`. - -### 2.4 Tool task input carries server-injected keys at the top level - -`inputData` for a worker tool is the LLM's arguments flattened at the top level **plus** -`method` (tool name, always), `_agent_state`, `_agent_tool_name`, and `_allowed_commands` for -`cli` tools. Verified in `conductor-mocks/mocks/agent/tool_happy_path/mappings/04_*.json`. - -**Decision D7**: `Dispatch#coerce_args` strips those four keys before binding keyword arguments, -rejects missing required arguments with `FAILED_WITH_TERMINAL_ERROR` (not a Ruby -`ArgumentError` mid-call; see 2.9 on `city: String` defaults), coerces `String -> Integer/Float/ -Boolean` and `String <-> JSON` for array/object params, returns a Hash as-is and wraps any other -result as `{"result" => value}`. A `_state_updates` key in the result passes through (server -merges it into `_agent_state`). Non-JSON-serializable results fail the task with a clear reason. - -### 2.5 `requiredWorkers` and system workers - -`POST /agent/start` returns `{executionId, agentName, requiredWorkers: [String]}`. The list is -flat task names: every `toolType: worker` tool **by its own name, no prefix**, plus -compiler-generated SIMPLE tasks the SDK must serve: `_termination` whenever -`termination` is set, custom guardrail `taskName`s, callback `_` tasks, -`on_condition` handoff `taskName`s, `stopWhen.taskName`. `_handoff_check`, -`_check_transfer` and `_transfer_to_` are compiled INLINE on this server and -are **not** in the list (Python still registers them for older servers; we do not). - -**Decision**: `ToolRegistry#register_system_workers(required_workers)` serves exactly the -compiler-generated names it recognises by suffix, with Ruby ports of Python's -`TerminationEntry` (`{should_continue, reason}`), `GuardrailEntry` (`{passed, message, on_fail, -fixed_output, guardrail_name, should_continue}`), `CallbackEntry`, and the `on_condition` -handoff worker. An unrecognised name logs a warning listing it, because a task with no worker -sits SCHEDULED forever. Since `stop_when`/`stop_after` serialize to `termination`, the -termination worker is needed for the team example and is Phase 2, not later. - -### 2.6 SSE framing and reconnect (server `AgentStreamRegistry`) - -`GET /api/agent/stream/{executionId}`, `Accept: text/event-stream`. Server writes `:connected` -first, then `id:\nevent:\ndata:\n\n`, ids start at 1 per execution, a -`:heartbeat` comment every 15 s, emitter never times out, stream is closed after `done` or -`error`. Replay buffer of 200 events for 5 minutes after completion. `Last-Event-ID` is an HTTP -request header bound to a Java `Long`: send a bare integer; when absent the server replays from -0, so connecting after start loses nothing. Child sub-workflow events are aliased onto the -parent stream. Event types: `thinking tool_call tool_result handoff waiting guardrail_pass -guardrail_fail error done` plus `context_condensed subagent_start subagent_stop`. `done.output` -is the workflow output `{result, finishReason, context, rejectionReason}`. - -**Decision D3**: `SseClient` reconnects with `Last-Event-ID` after any drop with a fixed 1 s -backoff, raises `SseUnavailableError` if the first connect fails, and the runtime falls back to -polling `GET /agent/{id}/status` (fields `isComplete`, `isWaiting`, `output`, `pendingTool`, -`reasonForIncompletion`) every 0.5 s. This is simpler than Python's task-graph synthesis and -loses only `partial_text` in fallback mode. `partial_text` is the concatenation of `thinking` and -`message` `content` fields; there is no separate delta event. - -### 2.7 SSE cannot go through the existing `RestClient` - -`RestClient` has a 120 s total timeout, retry middleware, and buffers bodies. `SseClient` opens -its own long-lived connection: plain `Net::HTTP` with `read_body` streaming (no adapter -uncertainty, no new gem), read timeout disabled, headers from `ApiClient#get_authentication_headers` -so token refresh stays in one place. Faraday `on_data` through `faraday-net_http_persistent 2.3.1` -is the alternative; spike both in Phase 2 slice 1 and keep whichever streams the recorded -scenario end to end. - -### 2.8 Secrets ride on `runtimeMetadata`; both Ruby models lack the field - -`TaskDef.runtimeMetadata` is `List` (names). `Task.runtimeMetadata` is -`Map` (values), filled at poll time by `RuntimeMetadataResolver` (secret store, -then server env with `CONDUCTOR_SECRET_` / `CONDUCTOR_ENV_` prefixes), silently omitted on miss, -never persisted. Needs conductor-oss with PR #1255 (3.32.0-rc.8+). Neither field exists in -`lib/conductor/http/models/task.rb` or `task_def.rb` today. - -**Decision D5**: add both fields. `Secrets.secret(name)` reads -`TaskContext.current.task.runtime_metadata[name]`, then `ENV[name]`, then raises -`CredentialNotFoundError`. `TaskContext` is stored in `Thread.current[]`, which is fiber-local in -Ruby, so this already satisfies the "fiber-local, never writes ENV" rule with no new binding -mechanism. `secrets_env(*names)` returns a Hash for `system`/`spawn`/`Open3`. - -### 2.9 `tool def` type inference needs the AST, and so does secret scanning - -`Method#parameters` returns kinds and names only, never default values, so `city: String` vs -`units: 'metric'` is indistinguishable by reflection. Both type inference and `secret('...')` -scanning therefore use `RubyVM::AbstractSyntaxTree.of(method)` (MRI 2.6+, works when the source -file is on disk; in irb/`eval` needs `RubyVM.keep_script_lines = true`, Ruby 3.2+). Prism ships -only with Ruby 3.3, and the gem targets Ruby >= 3.0, so no parser dependency. - -**Decision D9**: `Tools#get_json_schema_def` maps `String/Integer/Float/Hash` class defaults to -required typed properties, literal defaults to optional properties with `default`, `[String]` to -`array` with `items`, `%w[a b]` to `enum` (optional, default first), `true/false` to boolean. -When the AST is unavailable the fallback is Python's behaviour: every keyword param gets `{}` -and `:keyreq` params are required; secret declaration then needs the explicit -`add_tool ..., credentials:`. Positional parameters raise at `tool def`. A one-line spike -proving `AbstractSyntaxTree.of` works for a top-level `def` and for a method in a module is the -first task of Phase 1. - -Runtime consequence: because `city: String` has the class as its Ruby default, the worker must -never call the method with a required argument missing (it would receive the `String` class). -That is why D7 validates required arguments before the call. - -### 2.10 Tool task definitions - -Python registers each tool with `retryCount 2, retryDelaySeconds 2, retryLogic LINEAR_BACKOFF, -timeoutSeconds 0, responseTimeoutSeconds 10, timeoutPolicy RETRY, runtimeMetadata [names]`, -`overwrite_task_def: true`, `worker_id "agent-sdk"`. The server also registers a TaskDef for -every required worker (with `responseTimeoutSeconds 3600`); ours overwrites it, exactly as the -recorded scenario shows (`02_put_api_metadata_taskdefs.json`). - -**Decision D6**: same defaults, via `Worker.define(name, register_task_def: true, -overwrite_task_def: true, task_def_template: TaskDef.new(...))`. `TaskDefinitionRegistrar# -build_task_definition` already honours `task_def_template`; it must stop overriding -`timeout_seconds`/`response_timeout_seconds` when the template sets them (it currently uses -`||=`, which is fine for non-nil values but `timeout_seconds: 0` is truthy in Ruby, so no change -needed; add a test). - -### 2.11 MCP discovery is server-side; drop `McpDiscovery` - -The compiler emits `LIST_MCP_TOOLS` and an LLM filter chain before the agent loop for every -`toolType: mcp` tool, and dispatches `CALL_MCP_TOOL` at runtime. Python's `mcp_discovery.py` -served its now-legacy local-compile path. **Decision D4**: no `McpDiscovery` class; `ToolDef.mcp` -serializes `{server_url, headers, tool_names, max_tools}` and nothing else happens client-side. -The one-pager's runtime diagram loses one box. - -### 2.12 Stateful runs, domains, sessions - -`AgentStartRequest.runId` makes the server map every required worker to that task domain. Python -sends `runId = uuid4().hex` only when the agent or any tool is `stateful`, then polls with -`domain == run_id`. `sessionId` is passed straight through. **Decision**: same. `Agent#stateful` -and `ToolDef#stateful` exist from Phase 1; domain wiring is a Phase 2 task with one test. - -### 2.13 Finish reasons and statuses - -Server `finishReason` is uppercase (`STOP TOOL_CALLS MAX_TOKENS CONTENT_FILTER LENGTH`) except -the literal `"rejected"`. Execution status is the Conductor workflow status (`RUNNING COMPLETED -FAILED TIMED_OUT TERMINATED PAUSED`); `executionId == workflowId`. **Decision D8**: -`Execution#finish_reason` is `:stop | :tool_calls | :length | :content_filter | :rejected | -:error | :cancelled | :timeout`, derived the way Python's `_derive_finish_reason` does. - -### 2.14 `redact` needs verification - -`redact` is specified as `RegexGuardrail(position: :output, on_fail: :fix)`. Whether -`GuardrailCompiler` rewrites matched text for a regex guardrail with `onFail: fix` was not -confirmed. Phase 1 serializes it as specified; Phase 3's replay scenario for the team example -verifies behaviour, and if the server does not rewrite, `redact` switches to `on_fail: :retry` -and the doc is updated. - -### 2.15 Testing infrastructure reality - -`SDK_AGENT_TESTING_STRATEGY_V2.md` (in-server `mockLLM` provider, record mode) is **not -implemented on the server yet**: only fixture primitives exist on `feature/llm_mock_impl` -(`ai/.../testing/LlmFixture*.java`, `llm-fixture.schema.json`); no `MockLLM` provider, no -`conductor.ai.mock-llm.record` property, no test-server task. What exists today is -`conductor-mocks` with one normalized WireMock scenario, `agent/tool_happy_path`, recorded -against a real server for this SDK. - -**Decision**: contract tests need no server and land first. Runtime tests use WireMock replay of -`tool_happy_path` now (it is the only recording, and it covers start, TaskDef PUT, poll with -`runtimeMetadata: []`, update-v2, SSE). Approval, secrets and team scenarios are recorded into -`conductor-mocks` as they become runnable. The mockLLM functional suite is a Phase 3 follow-up -gated on the server; the spec helper is written so `mocks:` (WireMock) and `model: -'mockLLM/'` (real server) coexist. - -### 2.16 Prerequisite already called out in the design: instance-level token cache - -`Configuration` keeps `auth_token`/`token_update_time` as class-level state. Two runtimes with -different credentials in one process (tests, multi-tenant workers) would share one token. Only -`ApiClient` reads it. Move to instance level; keep the class accessors as deprecated shims for -one release. - -### 2.17 Local toolchain - -Ruby is not installed on this machine (`ruby: command not found`; no rbenv/rvm/mise). Docker and -podman are. Either install Ruby 3.3 or run the suite in `ruby:3.3-alpine` (the repo's -`Dockerfile` base). This is the first checklist item in Phase 0. - -### 2.18 Task result updates go to update-v2 (added during implementation) - -The recorded scenario shows the tool result posted to `POST /api/tasks/update-v2` with -`extendLease: false`; the Ruby runner posted to `POST /api/tasks`. The Python runner uses -update-v2 by default and falls back to `/tasks` once on 404/405. **Decision**: same in -`TaskRunner#send_task_update`; `Worker` gains `lease_extend_enabled` (tool workers set it, like -Python) so lease extension can follow later. - -### 2.19 `hands_off_to` on a member makes the team a swarm (added during implementation) - -The design puts handoffs on the member (`triage.hands_off_to filer`), but the server reads -`handoffs` on the coordinator and only acts on them under `strategy: swarm`; a member's -handoffs under the default `handoff` strategy would be ignored. **Decision D10**: when members -declare handoffs and the team has no explicit strategy, `ConfigSerializer` emits -`strategy: swarm` and hoists the members' handoffs onto the team. An explicit strategy is -never overridden. Python-style teams (handoffs on the parent, `strategy: :swarm`) serialize -unchanged; goldens 13 and 17 prove it. - -### 2.20 The run domain applies to every worker (added during implementation) - -When a start request carries `runId`, the server maps every name in `requiredWorkers` to that -task domain, not just stateful tools. **Decision**: `ToolRegistry` gives all workers of a -stateful run the run domain; the plan's per-tool rule was wrong. - ---- - -## 3. Target layout - -``` -lib/conductor/agents.rb # require 'conductor/agents'; requires everything below -lib/conductor/agents/ - version.rb? # no: reuse Conductor::VERSION - errors.rb # ConfigurationError, CredentialNotFoundError, AgentApiError, - # AgentNotFoundError, SseUnavailableError, ToolSerializationError - agent.rb # Agent (definition + sugar + call_sync/call_async delegating to runtime) - tool_def.rb # ToolDef, ToolType, PrefillToolCall (minimal), factories http/mcp/human/agent - tools.rb # Tools module: tool, describe, requires_approval, [], schema + secret scan - tools/schema_builder.rb # AST -> JSON schema (D9) - tools/secret_scanner.rb # AST -> literal secret()/secrets_env() names - tools/ruby_llm_adapter.rb # RubyLLM::Tool class -> ToolDef (only if defined?(RubyLLM)) - guardrail.rb # Guardrail, RegexGuardrail, LlmGuardrail, GuardrailResult - termination.rb # Termination::Condition, TextMention, StopMessage, MaxMessage, TokenUsage, And, Or - handoff.rb # Handoff::Condition, OnToolResult, OnTextMention, OnCondition - callback_handler.rb # CallbackHandler + POSITION_TO_METHOD - memory.rb # ConversationMemory - prompt_template.rb # PromptTemplate - config_serializer.rb # ConfigSerializer.serialize(agent) -> Hash (camelCase) - runtime/agent_config.rb # AgentConfig.from_env - runtime/agent_runtime.rb # AgentRuntime: call_sync/call_async/deploy/serve/shutdown - runtime/execution.rb # Execution, ToolCall, TokenUsage - runtime/approval_request.rb # ApprovalRequest - runtime/sse_client.rb # SseClient: each_event(execution_id, last_event_id:), reconnect - runtime/status_poller.rb # polling fallback over /agent/{id}/status - runtime/tool_registry.rb # ToolRegistry: register_tool_workers, register_system_workers - runtime/dispatch.rb # Dispatch.run_tool_task, coerce_args - runtime/system_workers.rb # termination / guardrail / callback / handoff worker bodies - runtime/secrets.rb # Secrets.secret, secrets_env -lib/conductor/http/api/agent_resource_api.rb # AgentResourceApi -lib/conductor/client/agent_client.rb # AgentClient (Hash in / Hash out, like Python) -lib/conductor/orkes/orkes_clients.rb # + get_agent_client -lib/conductor/http/models/task.rb # + runtime_metadata (Hash) -lib/conductor/http/models/task_def.rb # + runtime_metadata (Array) -spec/conductor/agents/** # unit + contract specs (no server) -spec/agents/** # runtime specs (WireMock replay / real server), tagged -spec/fixtures/agents/agent-schema.json # vendored from python-sdk docs/agents/reference -spec/fixtures/agents/configs/*.json # 19 golden configs vendored from python-sdk examples/agents/_configs -examples/agents/{weather,support_approval,bug_desk}.rb -examples/agents/dump_agent_configs.rb # regenerates goldens from Ruby for cross-SDK diff -docs/agents/README.md + concepts/*.md -``` - -Namespacing: definition classes live directly under `Conductor::Agents` so `include -Conductor::Agents` gives `Agent`, `RegexGuardrail`, `Termination`, `Handoff`, and the `tool`, -`describe`, `requires_approval`, `secret`, `secrets_env` methods. `Tools` is included into -`Conductor::Agents` so top-level `include Conductor::Agents` works; `extend -Conductor::Agents::Tools` on a module works because `tool` resolves `method(name)` first and -falls back to `instance_method(name)` + `module_function name`. - ---- - -## 4. Work breakdown - -Each slice is one reviewable PR. Every PR: `bundle exec rubocop`, `bundle exec rspec -spec/conductor/`, coverage not lower than before, `bundle exec ruby -Ilib -e "require -'conductor/agents'"` loads. Sizes: S < 1 day, M 1-3 days, L 3-5 days. - -### Phase 0: prerequisites (no agent code yet) - -**PR 0.1 Toolchain + hygiene (S)** -- Install Ruby 3.3 locally or document `docker run --rm -v $PWD:/app -w /app ruby:3.3 bundle exec rspec`. -- Remove unused `vcr` dev dependency (called out in `AGENT_TESTING.md`); keep `webmock`. -- Add `json_schemer` as a development dependency for contract tests. -- Acceptance: `bundle exec rspec spec/conductor/` green locally. - -**PR 0.2 Instance-level auth token cache (S)** (2.16) -- `Configuration`: `@auth_token`, `@token_update_time` per instance; `update_token`, - `auth_token`, `token_update_time` read instance state. Class-level accessors remain, emit a - deprecation warning once, and read/write a process-wide fallback only when the instance has none. -- `ApiClient` unchanged in behaviour; add a spec that two configurations hold two tokens. -- Acceptance: all existing specs pass; new spec proves isolation. - -**PR 0.3 `runtimeMetadata` on models + registrar (S)** (2.8, 2.10) -- `Task`: `runtime_metadata: 'Hash'` / `:runtimeMetadata`, default `{}`. -- `TaskDef`: `runtime_metadata: 'Array'` / `:runtimeMetadata`, default `[]`. -- `TaskDefinitionRegistrar`: test that `task_def_template` with `timeout_seconds: 0`, - `response_timeout_seconds: 10`, `runtime_metadata: ['GH_TOKEN']` survives `build_task_definition`. -- Acceptance: model round-trip specs; recorded `04_*` poll body deserializes `runtimeMetadata`. - -**PR 0.4 Agent transport (M)** (server §2 of the API report) -- `Http::Api::AgentResourceApi` over `ApiClient#call_api`, `return_type: 'Hash'`: - - `start(body)` POST `/agent/start`; `deploy(body)` POST `/agent/deploy`; `compile(body)` POST `/agent/compile` - - `status(id)` GET `/agent/{executionId}/status`; `execution(id)` GET `/agent/execution/{executionId}` - - `executions(params)` GET `/agent/executions` (`start,size,sort,freeText,status,agentName,sessionId`) - - `respond(id, body)` POST `/agent/{executionId}/respond`; `stop(id)`; `signal(id, message)` - - `pause(id)` PUT `/agent/{executionId}/pause`; `resume(id)` PUT; `cancel(id, reason:)` DELETE `/agent/{executionId}/cancel` - - `list` GET `/agent/list`; `get(name, version:)` GET `/agent/{name}`; `delete(name, version:)` -- `Client::AgentClient` wrapping it with Python's method names (`start_agent`, `deploy_agent`, - `compile_agent`, `get_status`, `get_execution`, `list_executions`, `respond`, `approve`, - `reject`, `send_message`, `stop`, `signal`, `pause`, `resume`, `cancel`). Map 404 to - `AgentNotFoundError`, other `ApiError` to `AgentApiError` carrying the `{error, status}` body. -- `OrkesClients#get_agent_client`. -- Specs mirror `spec/conductor/client/prompt_client_spec.rb` (instance doubles) plus WebMock - specs asserting exact paths and bodies for start/respond/status. -- Acceptance: every endpoint in the table has a spec asserting method + path + body keys. - -### Phase 1: definition layer and serializer (no server) - -**PR 1.1 Spike: AST availability (S, half day, can be a scratch script)** (2.9) -- Prove `RubyVM::AbstractSyntaxTree.of(method)` returns `KW_ARG` defaults and `secret('X')` - call nodes for: a top-level `tool def` in a file, a method in a module with `extend Tools`, a - method defined in irb with `keep_script_lines`. Record the matrix in `tools/schema_builder.rb` - comments. If the top-level-in-file case fails on any supported Ruby (3.0-3.3), the fallback in - D9 becomes the primary path and the DSL doc changes before any more code is written. - -**PR 1.2 ToolDef, ToolType, factories (M)** -- `ToolType` constants: `worker http api mcp human agent_tool generate_image generate_audio - generate_video generate_pdf rag_index rag_search pull_workflow_messages cli`. -- `ToolDef` fields and defaults exactly as Python: `name, description '', input_schema {}, - output_schema {}, func nil, approval_required false, timeout_seconds nil, tool_type 'worker', - config {}, guardrails [], credentials [], stateful false, max_calls nil, retry_count 2, - retry_delay_seconds 2, retry_policy 'linear_backoff'`. -- Factories from the one-pager: `ToolDef.http(name, url, method: 'GET', description:, headers:, - input_schema:, credentials:)`, `ToolDef.mcp(server_url, name: 'mcp_tools', headers:, - tool_names:, max_tools: 64, credentials:)`, `ToolDef.human(name, description:, input_schema:)`, - `ToolDef.agent(agent, name:, description:, retry_count:, retry_delay_seconds:, optional:)`. - `${NAME}` header placeholders must appear in `credentials` or raise. Media/RAG/api factories - are one method each and can ride along if cheap; not required by the examples. -- Specs: config keys per type match Python's `_serialize_tool` (`url method headers accept - contentType`; `server_url headers tool_names max_tools`; `agent` replaced by `agentConfig`). - -**PR 1.3 Tools DSL: `tool def`, `describe`, `requires_approval`, `[]`, schema, secret scan (L)** -- `tool(name)`: resolves the method, builds `ToolDef(name:, description: humanize(name), - input_schema: SchemaBuilder.for(method), credentials: SecretScanner.scan(method), func: ->(**kw) {...})`, - registers it in a per-`self` registry (`Tools#[]`), returns the ToolDef. Positional params raise. -- `describe(name, text)`, `requires_approval(name)` mutate the registered ToolDef. -- `SchemaBuilder` implements the D9 table; `required` only when non-empty; no `$schema` key - (Python emits a bare `{"type":"object","properties":{...},"required":[...]}`). -- `SecretScanner` collects string-literal first arguments of `secret(...)` and every - string-literal argument of `secrets_env(...)`; ignores dynamic names. -- `RubyLlmAdapter`: when `defined?(RubyLLM::Tool)` and `add_tool` receives such a class, build a - `ToolDef` from `.name`, `.description`, `.parameters`-derived schema, `func: ->(**kw) { - klass.new.execute(**kw) }`. -- Specs: one per row of the type table; secret scan positive/negative/dynamic; module-extend - form; RubyLLM adapter behind a stub class. - -**PR 1.4 Guardrails, termination, handoffs, callbacks, memory, prompt template (M)** -- Straight ports with Ruby naming. Validation rules from Python: guardrail `position` in - `input|output`, `on_fail` in `retry|raise|fix|human`, `human` illegal on `input`, - `max_retries >= 1`; `MaxMessage >= 1`; `TokenUsage` needs at least one limit; `&`/`|` - flatten same-type children; `RegexGuardrail(patterns, mode: :block|:allow, message:)`; - `LlmGuardrail(model, policy, max_tokens:)`; `Guardrail.new(name:)` with no block is external. -- `ConversationMemory` with `add_user_message` etc. and `_trim` semantics (keep system messages). -- Specs per class including combinator flattening. - -**PR 1.5 Agent (M)** (2.1, 2.2, 2.12) -- Constructor keywords: `name:, model: nil, instructions: '', tools: [], agents: [], strategy: - :handoff, router: nil, output_type: nil, guardrails: [], memory: nil, termination: nil, - handoffs: [], callbacks: [], credentials: [], max_turns: 25, max_tokens: nil, - timeout_seconds: 0, temperature: nil, stateful: false, metadata: nil, description: nil, - external: false`. -- Validation: name regex, strategy enum (symbols, serialized lowercase), `max_turns >= 1`, - duplicate sub-agent names, `router` required for `:router`. -- Sugar: `add_tool(tool, credentials: nil)` accepts `Symbol` (looks up `Tools` registry of the - caller and of `Conductor::Agents`), `ToolDef`, a `Tools`-extended module (adds all), a - RubyLLM class; `add_tools(*)`, `add_agent(s)`, `hands_off_to(agent, on:)` -> - `Handoff::OnTextMention`, `redact(words)` -> `RegexGuardrail(position: :output, on_fail: :fix, - mode: :block, name: "#{name}_redact")`, `stop_when(text)` -> `Termination::TextMention`, - `stop_after(messages:)` -> `Termination::MaxMessage` (combined with `|` when both set), - `on_approval(&block)`, `strategy=`. -- `call_sync(prompt, session_id: nil)` and `call_async(prompt, session_id: nil, &on_done)` - delegate to `Conductor::Agents.runtime` (Phase 2); until then they raise `NotImplementedError` - with a pointer. -- Specs: validation matrix; sugar produces the parity objects listed in `AGENT_TEAMS.md`. - -**PR 1.6 ConfigSerializer + contract tests (L)** (2.1, 2.2) -- Emission order and conditions per Python's `serialize`, `_serialize_tool`, - `_serialize_guardrail`, `_serialize_termination`, `_serialize_handoff`, `_serialize_memory`, - plus D1 (parent model inheritance) and the `ConfigurationError` for model-less leaves. -- Callbacks serialize to `{position, taskName: "#{name}_#{position}"}` for every - `CallbackHandler` method a handler overrides. -- Vendor `spec/fixtures/agents/agent-schema.json` and the 19 `configs/*.json`. Port - `dump_agent_configs.py` to `examples/agents/dump_agent_configs.rb` so the Ruby definitions of - the same 19 agents live in the repo; the spec compares `serialize(agent)` to each golden file - (`sort_keys`-insensitive Hash equality, model string parametrised by env like Python's dump - script). Validate every serialized config against the schema; assert an unknown root key fails. -- Acceptance: 19/19 golden matches, schema valid, and `examples/agents/dump_agent_configs.rb` - regenerates byte-identical JSON (sorted keys, 2-space indent) for cross-SDK diffs. - -### Phase 2: runtime and transport - -**PR 2.1 Spike + SseClient (M)** (2.6, 2.7) -- Implement over `Net::HTTP` streaming first; if `Faraday` `on_data` via the persistent adapter - streams the recorded `06_get_api_agent_stream_exec_1.json` body (WireMock dribbles it over ~3 s) - with less code, switch. Parser handles `:` comments (heartbeat), `id:`, `event:`, `data:` - (multi-line joined by `\n`), blank-line boundaries, JSON parse failure -> `{"content" => raw}`. -- `each_event(execution_id, last_event_id: nil)` is an Enumerator that stops after `done`/`error`, - reconnects with `Last-Event-ID: ` after a drop, raises `SseUnavailableError` on - first-connect failure or 15 s of heartbeat-only silence before the first real event. -- Specs with WebMock streaming bodies (WebMock can serve a String body; chunk pacing is not - needed for parsing tests) covering comments, multi-line data, reconnect id, terminal events. - -**PR 2.2 Secrets + Dispatch (M)** (2.4, 2.8) -- `Secrets.secret(name)` / `secrets_env(*names)` per D5; `CredentialNotFoundError` message names - the tool and the missing key. -- `Dispatch.run_tool_task(task, tool_def)` per D7; returns a `TaskResult`, sets `worker_id - 'agent-sdk'`, `FAILED` with `reason_for_incompletion` on exceptions, `FAILED_WITH_TERMINAL_ERROR` - on missing required args, missing credentials, or `ToolSerializationError`. Circuit breaker - (10 consecutive failures per tool) is optional; include only if trivial. -- Specs: strips injected keys (uses the recorded `04_*` input verbatim), coercions, missing - required arg, Hash vs scalar result, `_state_updates` passthrough, secret from - `runtime_metadata` vs ENV vs missing, `secrets_env` returns only requested names. - -**PR 2.3 ToolRegistry + system workers (M)** (2.5, 2.10) -- `register_tool_workers(tools, agent_name, domain)` builds one `Worker.define` per `worker`/`cli` - tool with the D6 TaskDef template and `Dispatch` body; server-side tool types are skipped. -- `register_system_workers(required_workers, agent)` matches `_termination`, - `_` callbacks, custom guardrail names, `on_condition` handoff task names; - bodies in `system_workers.rb` are ports of Python's `TerminationEntry`, `GuardrailEntry`, - `CallbackEntry`, `HandoffCondition`. Unknown names warn. -- Workers run on a dedicated `TaskHandler` owned by the runtime with `poll_interval` / - `thread_count` from `AgentConfig`, `register_task_definitions: true`. -- Specs: worker set for the three examples; TaskDef body equals the recorded `02_*` PUT body - (with `runtimeMetadata: []`); termination worker returns `should_continue: false` on - `TextMention` match; unknown required worker logs. - -**PR 2.4 AgentRuntime, Execution, ApprovalRequest (L)** (2.3, 2.6, 2.12, 2.13) -- `AgentConfig.from_env`: `CONDUCTOR_AGENT_WORKER_POLL_INTERVAL` (100 ms), - `CONDUCTOR_AGENT_WORKER_THREADS` (1), `CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER` (false), - `CONDUCTOR_AGENT_STREAMING_ENABLED` (true); Python boolean parsing (`true/1/yes/on`). -- `AgentRuntime#call_async(agent, prompt, session_id:, &on_done)`: serialize; `run_id = - SecureRandom.hex` when stateful; `POST /agent/start` with `{agentConfig, prompt, sessionId, - runId?}`; register tool + system workers from `requiredWorkers` under `domain = run_id`; start - the handler; create `Execution`; spawn the stream thread (`SseClient`, falling back to - `StatusPoller` on `SseUnavailableError` or when streaming is disabled); return immediately. - `call_sync` is `call_async(...).result`. -- Stream thread: `thinking`/`message` append `partial_text`; `tool_call`/`tool_result` pair into - `execution.tool_calls`; `waiting` with `pendingTool.toolCalls` builds an `ApprovalRequest` and - invokes `agent.on_approval` (else parks it on `execution.pending`); `done` sets `result = - output['result']`, `finish_reason` per D8, fetches `tokenUsage` from `GET /agent/execution/{id}` - (recursing `tasks[].subWorkflowId`), marks done, fires `on_done`; `error` sets `:error`. All - callbacks are wrapped so a user exception never kills the stream thread. -- `Execution`: `execution_id, done?, waiting?, result (blocks with optional timeout), - partial_text, finish_reason, tool_calls, token_usage, pending, pause, resume, cancel(reason:), - stop, signal(text)`; `Execution.find(id)` builds from `get_status` + `get_execution`. -- `deploy(*agents)` -> `POST /agent/deploy` each, returns agent names; `serve(*agents)` deploys, - registers workers, blocks until `INT`/`TERM`, then `shutdown`; `shutdown` stops the handler and - stream threads. -- `Conductor::Agents.runtime` memoised default (`Configuration.new` from env), - `Conductor::Agents.configure(configuration:, agent_config:)`. -- Specs (WebMock, no WireMock): start body equals `serialize(agent)` + prompt; workers - registered for `requiredWorkers`; approval approve/reject post the exact bodies; rejection -> - `:rejected`; fallback poller path; `on_done` fires once; user exception in `on_approval` is - logged not raised. - -### Phase 3: end-to-end tests and CI - -**PR 3.1 WireMock replay of `agent/tool_happy_path` (M)** (2.15) -- `spec/agents/agents_helper.rb`: `mocks: 'agent/tool_happy_path'` metadata boots - `wiremock/wiremock:3x` with the scenario dir mounted (Docker via `docker`/`podman` CLI; skip - with a clear message when neither exists or `CONDUCTOR_MOCKS_DIR` is unset), points - `CONDUCTOR_SERVER_URL` at it, asserts `GET /__admin/requests/unmatched` is empty in `after`. -- The weather spec from `AGENT_TESTING.md` passes end to end: start, TaskDef PUT, poll, tool - runs in-process, update-v2, SSE `done`, `call_sync` returns the recorded text. -- CI: new `agents-replay` job that checks out `conductor-oss/conductor-mocks` and runs - `bundle exec rspec spec/agents`. Runs on every PR including forks (no secrets). - -**PR 3.2 Record the remaining scenarios (M, blocked on a server with AI enabled and keys)** -- `approval_approve`, `approval_reject`, `secrets_runtime_metadata`, `team_handoff` using the - three examples, recorded with WireMock's snapshot recorder and `scripts/normalize.py` per the - `conductor-mocks` README; PRs to `conductor-mocks`; specs here tagged with each scenario. - `team_handoff` also settles 2.14 (`redact` semantics). - -**PR 3.3 mockLLM functional suite (L, blocked on server)** -- When `MockLLM` lands in conductor-oss: `spec/agents/functional/*_spec.rb` for the six-scenario - catalog in `SDK_AGENT_TESTING_STRATEGY_V2.md`, asserting persisted LLM task input and tool - tasks via `WorkflowClient#get_workflow`, plus a CI job that builds the pinned server. Not a - blocker for releasing the feature; replay covers the wire contract until then. - -### Phase 4: docs, examples, release - -**PR 4.1 Examples + docs (M)** -- `examples/agents/weather.rb`, `support_approval.rb`, `bug_desk.rb` exactly as in the one-pager, - plus `dump_agent_configs.rb` from Phase 1. -- `docs/agents/README.md` and `concepts/{tools,streaming-hitl,teams,secrets}.md` adapted from the - four design sub-specs; README gets an "Agents" section after "LLM/AI Tasks"; `AGENTS.md` and - `DESIGN.md` get the new layer in the architecture diagrams and directory listing. -- Update the design docs for the decisions in section 2 (D1 model inheritance, D2 approval batch - semantics, D4 no `McpDiscovery`, D9 AST fallback, 2.15 testing reality) and refresh the - Confluence page from the one-pager (read the live page immediately before writing; it is - edited concurrently). -- `CHANGELOG.md` Unreleased: Added `Conductor::Agents`, `AgentClient`, `runtimeMetadata` - fields; Changed instance-level token cache; Removed `vcr`. - -**PR 4.2 Release (S)** -- Version bump (minor), gem build, tag. `conductor/agents` stays an explicit require; `require - 'conductor'` does not load it. - ---- - -## 5. Sequencing and parallelism - -``` -0.1 -> 0.2 -> 0.3 -> 0.4 ----------------------------------. - \ \ - 1.1 -> 1.2 -> 1.3 -> 1.4 -> 1.5 -> 1.6 -> 2.3 -> 2.4 -> 3.1 -> 4.1 -> 4.2 - 2.1 --' \ - 2.2 --' '-> 3.2 (server + keys) - '-> 3.3 (server mockLLM) -``` - -Phase 0 and Phase 1 are independent after 0.3 and can run in parallel. 2.1 and 2.2 depend only -on Phase 0 and 1.2. The critical path is 1.1 -> 1.3 -> 1.6 -> 2.4 -> 3.1. - -Rough total: Phase 0 ~3 days, Phase 1 ~8 days, Phase 2 ~9 days, Phase 3.1 ~2 days, Phase 4 ~3 -days. About five weeks for one engineer to a releasable feature with replay-tested wire contract; -3.2 and 3.3 follow as the server pieces land. - ---- - -## 6. Risks and open questions - -| # | Risk | Mitigation | -|---|---|---| -| R1 | `RubyVM::AbstractSyntaxTree.of` unavailable or lossy on some supported Ruby (3.0-3.3) or non-MRI | PR 1.1 spike first; D9 fallback (`{}` schema, explicit `credentials:`) is always available; document limits | -| R2 | Server rejects team parent without `model` | D1 inherits from first child at serialize time; contract test | -| R3 | `redact` (`regex` + `fix`) may not rewrite on the server | Verify in 3.2; fall back to `retry` | -| R4 | SSE through WireMock is buffered; without `chunkedDribbleDelay` `done` can arrive before the tool executes | `normalize.py` already paces the body; assert unmatched requests empty | -| R5 | `mockLLM` functional testing not available on the server yet | Replay now, functional later (3.3); spec helper supports both | -| R6 | `runtimeMetadata` needs conductor-oss >= 3.32.0-rc.8 (PR #1255); older servers omit it | `secret()` falls back to `ENV`; document minimum version | -| R7 | OSS vs Orkes model string semantics differ (provider type key vs integration name) | Docs say "left side is the integration name"; on OSS that is the provider key with `*_API_KEY` env; note both | -| R8 | Long-lived SSE thread plus worker threads plus user callbacks: exceptions in user blocks | Wrap every user callback; `execution.error` captures; never let the stream thread die silently | -| R9 | Class-level token cache change alters behaviour for users relying on cross-instance sharing | Deprecated shim for one release; changelog entry | -| R10 | `Last-Event-ID` must be an integer; a stray string 400s the reconnect | Parser stores `id` as Integer; reconnect sends `to_s` of it only | -| R11 | Server auto-registers TaskDefs with `responseTimeoutSeconds 3600`; ours overwrites with 10 | Same as Python; document that lease extension keeps long tools alive (follow-up: `lease_extend_enabled`) | - -Open questions for the server team (none block Phases 0-2): - -1. Confirm `regex` guardrail with `onFail: fix` rewrites content (2.14). -2. Timeline for `MockLLM` provider and the test-server entry point (2.15). -3. Whether `_termination` is emitted for `termination` on sub-agents too, so the - registry can register it per agent in a team. diff --git a/docs/design/AGENTS_PARITY_AUDIT.md b/docs/design/AGENTS_PARITY_AUDIT.md deleted file mode 100644 index 3a3a774..0000000 --- a/docs/design/AGENTS_PARITY_AUDIT.md +++ /dev/null @@ -1,71 +0,0 @@ -# Ruby agents parity audit - -Audited 2026-09-17 against version 2 of the live -[Ruby Agents Parity Plan](https://orkes.atlassian.net/wiki/spaces/ENG/pages/53739522/Ruby+Agents+Parity+Plan), -Python SDK `c99e2cf9871c21f7a64d823126ee1b77989b00ad`, and Conductor -`acb7d27750e5dacc6a6334ed0bcabe3ade030533`. - -All 19 requested example ports are implemented and their integration wrappers passed on a -fresh Conductor OSS playback server. The supplied shared action's script reported -**93/93 recordings played back; 0 unmatched requests**. Recordings were not changed. -The GitHub workflow is configured but has not yet been run on GitHub. - -## Plan coherence - -| Plan surface | Implementation and evidence | -|---|---| -| Agent, tools, tool types, schemas, approval metadata | `agent.rb`, `tool_def.rb`, `tools.rb`; serializer and DSL contracts | -| Guardrails, termination, handoffs, callbacks, memory, prompt templates | Definition classes and `runtime/system_workers.rb`; unit contracts, guardrail and team playback | -| Ruby sugar and RubyLLM adapter | `Agent` mutators, `>>`, `Tools`, schema/secret scanner and adapter unit tests | -| Serialization identical to Python | 20 golden contracts plus Python-generated configurations for all 19 runnable ports, validated against the server schema | -| Runtime, deployment, execution and approval | `AgentRuntime`, `Execution`, `ApprovalRequest`; lifecycle unit tests and real-server execution/approval tests | -| Streaming and polling fallback | `SseClient`, `StatusPoller`, `AgentClient#stream_sse`; reconnect/fallback unit tests, playback SSE events and completion | -| Worker dispatch and credentials | `ToolRegistry`, `Dispatch`, `Secrets`; inherited credentials and registration tests, independent external workers and credential-bearing HTTP playback | -| Transport and client factory | `AgentResourceApi`, `AgentClient`, `OrkesClients`; API tests and real-server calls | -| Three original Ruby recipes | `weather.rb`, `support_approval.rb`, `bug_desk.rb` remain available; their features are exercised by the numbered ports | -| Framework adapters marked N/A | Remain outside scope, as specified by the plan | -| Additional requested plan-and-compile example | `Plans#plan_execute`, named planner/fallback slots, inherited model, context, recovery turn limit; golden and real-server compiled-workflow assertions | - -The live plan's diagrams are conceptual and contain names/ownership that differ from the -verified Python/server contract. These are documented mappings, not claims of literal -class-diagram identity: - -- MCP discovery belongs to the server compiler. Ruby sends `ToolDef.mcp`; there is no SDK - `McpDiscovery` class or extra discovery workflow. Example `04` verifies actual discovery. -- `ToolRegistry#tool_workers` / `#system_workers` build workers and the runtime registers - them; these implement the diagram's `register_tool_workers` / `register_system_workers` roles. -- `Dispatch.run_tool_task` runs tool bodies; `SseClient#each_event` reads events while - `AgentRuntime` owns fallback polling through `StatusPoller`. -- A server `waiting` event can gate a batch of tool calls. `ApprovalRequest` exposes that - batch and schema-driven `respond`, alongside `approve` / `reject` and argument access. -- Ruby uses `done?` / `waiting?` predicate methods. Model inheritance and implicit swarm - handoff hoisting follow Python serialization. -- Tool credential declarations use the server's `TaskDef.runtimeMetadata` shape. Provider - credentials stay on the server; the SDK does not provision LLM integration keys. - -These contract decisions are detailed in [the implementation plan](AGENTS_IMPLEMENTATION_PLAN.md), -section 2. The live Confluence page has not been edited; its diagram should adopt these -mappings before describing the implementation as literally identical to the design. - -## Tests and examples - -[The example guide](../../examples/agents/README.md) maps all 19 requested filenames and -provides standalone and playback commands. Integration tests import the actual files and -provide stdin/output/runtime dependencies; definitions and prompts stay only in examples. -Example `22` is expected to fail its strict guardrail, and the wrapper verifies the persisted -rejection rather than accepting arbitrary failures. - -The existing WireMock replay suite remains available. New playback runs the actual server, -MCP service, HTTP dependency and worker processes, then the common action checks the whole -recording set. Unit tests independently cover SSE failures, task retries, credentials, -router errors and worker registration that the successful recordings do not exercise. - -## Local verification - -- Unit suite: 584 examples, zero failures (baseline: 545). -- Full suite without optional service flags: 741 examples, zero failures, 157 expected - environment-gated pending examples. The 19 playback cases were separately enabled and passed. -- RuboCop: 211 files, zero offenses. Library load, Ruby syntax, and gem build passed. -- Ruby standard-library line coverage for loaded `lib/` files during the unit suite increased - from 5,077/6,642 (76.44%) at HEAD to 5,154/6,700 (76.93%). Both were measured with the same - Ruby 3.3 image and `Coverage.start(lines: true)` before loading RSpec. diff --git a/docs/design/AGENTS_PARITY_ONEPAGER.md b/docs/design/AGENTS_PARITY_ONEPAGER.md deleted file mode 100644 index de2dedd..0000000 --- a/docs/design/AGENTS_PARITY_ONEPAGER.md +++ /dev/null @@ -1,587 +0,0 @@ -# Ruby Agents Parity Plan - -Port `python-sdk/src/conductor/ai/agents` → `conductor_ruby` as `Conductor::Agents`. Same `agentConfig` on the wire; server compiles, SDK serializes + runs workers. - -## Classes - -### Definition (serialized to agentConfig) - -```mermaid -classDiagram - direction LR - class Agent { - +new(name:, model:, instructions:, tools: [], agents: [], **opts) - +String name - +String model "provider/model" - +String|PromptTemplate instructions - +List~ToolDef~ tools - +List~Agent~ agents - +Symbol strategy - +Agent|Proc router - +Hash output_type - +List~Guardrail~ guardrails - +ConversationMemory memory - +TerminationCondition termination - +List~HandoffCondition~ handoffs - +List~CallbackHandler~ callbacks - +List~String~ credentials - +Integer max_turns - +Boolean stateful - +add_tool(tool, credentials: nil) add_tools(*tools) - +add_agent(agent) add_agents(*agents) same as agents: - +hands_off_to(agent, on:) - +redact(words) stop_when(text) stop_after(messages:) - +call_sync(prompt, session_id:) String - +call_async(prompt, session_id:, &on_done) Execution - +on_approval() &block - } - class ConfigSerializer { - +serialize(agent) Hash - } - class ToolDef { - +String name - +String description - +Hash input_schema - +Hash output_schema - +Proc func - +ToolType tool_type - +Hash config per type - +Boolean approval_required - +List~String~ credentials - +Integer retry_count - +call(**args) PrefillToolCall - +http(name, url, method) ToolDef$ - +mcp(server_url) ToolDef$ - +human(name) ToolDef$ - +agent(agent) ToolDef$ - } - class Tools { - <> - +tool(method_name) ToolDef - +requires_approval(name) - +describe(name, text) - +[](name) ToolDef - -get_json_schema_def(method) Hash - -scan_secrets(method) List~String~ literal secret() names - } - class ToolType { - <> - worker - http - api - mcp - human - agent_tool - generate_image - generate_audio - generate_video - generate_pdf - rag_index - rag_search - pull_workflow_messages - } - class Guardrail { - +String name - +Symbol position input|output - +Symbol on_fail retry|raise|fix|human - +Integer max_retries - +Proc func - } - class RegexGuardrail - class LlmGuardrail - class TerminationCondition { - <> - +should_terminate(ctx) TerminationResult - +&(other) And - +|(other) Or - } - class TextMention - class StopMessage - class MaxMessage - class TokenUsage - class HandoffCondition { - +String target - +should_handoff(ctx) Boolean - } - class OnToolResult - class OnTextMention - class OnCondition - class CallbackHandler { - +on_agent_start() on_agent_end() - +on_model_start() on_model_end() - +on_tool_start() on_tool_end() - } - class ConversationMemory { - +List~Hash~ messages - +Integer max_messages - } - class PromptTemplate { - +String name - +Hash variables - +Integer version - } - - Agent "1" o-- "*" ToolDef : tools, shared by name - Agent "1" o-- "*" Agent : agents - Agent "1" o-- "*" CallbackHandler : callbacks - Agent "1" *-- "*" Guardrail : guardrails - Agent "1" *-- "0..1" TerminationCondition : termination - Agent "1" *-- "*" HandoffCondition : handoffs - Agent "1" *-- "0..1" ConversationMemory : memory - Agent --> "0..1" PromptTemplate : instructions - Agent --> "0..1" Agent : router - ConfigSerializer ..> Agent : reads - Tools ..> ToolDef : creates - ToolDef --> "1" ToolType : tool_type, one of these enum values - ToolDef "1" *-- "*" Guardrail : tool guardrails - Guardrail <|-- RegexGuardrail - Guardrail <|-- LlmGuardrail - TerminationCondition <|-- TextMention - TerminationCondition <|-- StopMessage - TerminationCondition <|-- MaxMessage - TerminationCondition <|-- TokenUsage - HandoffCondition <|-- OnToolResult - HandoffCondition <|-- OnTextMention - HandoffCondition <|-- OnCondition -``` - -### Runtime + transport - -```mermaid -classDiagram - direction LR - class AgentRuntime { - +call_sync(agent, prompt) String - +call_async(agent, prompt, &on_done) Execution - +deploy(*agents) List~String~ - +serve(*agents) - +shutdown() - } - class AgentConfig { - +Integer worker_poll_interval_ms - +Integer worker_thread_count - +Boolean auto_register_integrations - +Boolean streaming_enabled - +from_env() AgentConfig$ - } - class Execution { - +String execution_id - +done() Boolean - +waiting() Boolean - +result() String blocks until done - +String partial_text so far, non-blocking - +Symbol finish_reason - +List~ToolCall~ tool_calls - +TokenUsage token_usage - +pause() resume() cancel() - +find(execution_id) Execution$ - } - class ApprovalRequest { - +String execution_id - +String tool_name - +Hash arguments - +approve() - +reject(reason) - } - class ToolRegistry { - +register_tool_workers(tools, agent_name, domain) - +register_system_workers(required_workers) - } - class Dispatch { - +run_tool_task(task, tool_def) Hash - -coerce_args(input_data) - -bind_secrets(task, tool_def) - } - class Secrets { - <> - +secret(name) String - +secrets_env(*names) Hash - } - class SseClient { - +each_event(execution_id) Enumerator - -fallback_polling() - } - class McpDiscovery { - +expand(mcp_tool) List~ToolDef~ - } - class AgentClient { - +start_agent(payload) Hash - +deploy_agent(payload) Hash - +compile_agent(payload) Hash - +get_status(id) Hash - +get_execution(id) Hash - +list_executions(params) Hash - +respond(id, body) - +stop(id) signal(id, msg) - +stream_sse(id, last_event_id) Enumerator - } - class AgentResourceApi { - +start(body) POST agent-start - +deploy(body) POST agent-deploy - +compile(body) POST agent-compile - +status(id) GET agent-id-status - +execution(id) GET agent-execution-id - +executions(params) GET agent-executions - +respond(id, body) POST agent-id-respond - +stop(id) POST agent-id-stop - +signal(id, msg) POST agent-id-signal - +stream(id) GET agent-stream-id SSE - } - class Agent { - <> - } - class ConfigSerializer { - <> - } - class TaskHandler { - <> - } - class WorkflowClient { - <> - } - class ApiClient { - <> - } - class OrkesClients { - <> - +get_agent_client() AgentClient - } - - AgentRuntime *-- AgentConfig - AgentRuntime *-- AgentClient - AgentRuntime *-- ToolRegistry - AgentRuntime *-- SseClient - AgentRuntime ..> ConfigSerializer : agentConfig - AgentRuntime ..> Agent : reads on_approval - AgentRuntime ..> McpDiscovery : uses - AgentRuntime ..> Execution : creates - Execution --> AgentClient : pause, cancel - ApprovalRequest --> AgentClient : respond - SseClient ..> Execution : appends partial_text, fires on_done - SseClient ..> ApprovalRequest : creates on waiting - ToolRegistry --> TaskHandler : Worker.define - ToolRegistry ..> Dispatch : worker body - Dispatch ..> Secrets : binds per task - McpDiscovery ..> WorkflowClient : ephemeral LIST_MCP_TOOLS workflow - AgentClient *-- AgentResourceApi - AgentResourceApi --> ApiClient - OrkesClients ..> AgentClient : creates -``` - -> Implementation notes (see `AGENTS_IMPLEMENTATION_PLAN.md`, section 2): `McpDiscovery` was not -> built because the server discovers MCP tools itself at compile time; `ApprovalRequest` wraps the -> server's `waiting` event (one HUMAN task gates a whole turn of tool calls, so it carries -> `tool_calls`); a team parent without a model inherits the first member's model; members' -> `hands_off_to` make a strategy-less team a swarm; the polling fallback uses -> `GET /agent/{id}/status`. - -## Examples - -### 1. Tools - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } -end - -agent = Agent.new( - name: 'weather', - model: 'openai/gpt-4o', - instructions: 'Answer weather questions.' -) -agent.add_tool :get_weather - -puts agent.call_sync('Weather in Lisbon?') -``` - -`tool def` marks a method as a tool. Types come from the keyword defaults: `city: String` = required string, `units: 'metric'` = optional with default. Description = humanized method name (`describe :get_weather, '...'` to override). `RubyLLM::Tool` classes work as-is: `agent.add_tool Weather`. Spec: `docs/design/AGENT_TOOLS_DSL.md`. - -```mermaid -sequenceDiagram - actor You - participant SDK as Ruby SDK - participant Server as Conductor Server - participant LLM as OpenAI - - You->>+SDK: agent.call_sync("Weather in Lisbon?") (blocks until done) - SDK->>+Server: POST /agent/start (agentConfig, prompt) - Server-->>-SDK: executionId, requiredWorkers [get_weather] - SDK->>SDK: start worker for get_weather - SDK->>Server: GET /agent/stream (SSE) - Server->>+LLM: instructions + prompt + tool schema - LLM-->>-Server: call get_weather(city: "Lisbon") - Server->>SDK: task get_weather - SDK->>SDK: get_weather(city: "Lisbon") runs - SDK-->>Server: result - Server->>+LLM: tool result - LLM-->>-Server: "Sunny, 21°C in Lisbon" - Server-->>SDK: SSE text - Server-->>SDK: SSE done - SDK-->>-You: "Sunny, 21°C in Lisbon" -``` - -### 2. Streaming + approval - -```ruby -tool def issue_refund(order_id: String, amount: Float) - Billing.refund(order_id, amount) -end -requires_approval :issue_refund - -agent = Agent.new( - name: 'support', - model: 'anthropic/claude-sonnet-4-5', - instructions: 'Help with orders.' -) -agent.add_tool :issue_refund - -agent.on_approval do |request| - request.amount < 100 ? request.approve : request.reject('Needs a manager') -end - -# blocking -answer = agent.call_sync('Refund order A-1029, it arrived broken') - -# non-blocking, callback when finished -agent.call_async('Refund order A-1029, it arrived broken') do |answer| - Mailer.send(customer, answer) -end - -# non-blocking, poll it yourself -execution = agent.call_async('Refund order A-1029, it arrived broken') -execution.done? # false until finished -execution.partial_text # what has streamed so far -execution.result # blocks for the answer -execution.finish_reason # :stop | :rejected -``` - -`call_sync` blocks and returns the answer. `call_async` returns an `Execution` immediately; SSE runs on a background thread, and the block (if given) runs with the answer when done. Spec: `docs/design/AGENT_STREAMING.md`. - -```mermaid -sequenceDiagram - actor You - participant SDK as Ruby SDK - participant Server as Conductor Server - participant LLM as Anthropic - - You->>+SDK: agent.call_async("Refund order A-1029") with on_done block - SDK->>+Server: POST /agent/start - Server-->>-SDK: executionId, requiredWorkers [issue_refund] - SDK->>SDK: start worker for issue_refund - SDK->>Server: GET /agent/stream (SSE, background thread) - SDK-->>-You: Execution (returns immediately) - - Note over You,SDK: Everything below runs on background thread - - Server->>+LLM: prompt + tool schema - LLM-->>-Server: call issue_refund(order_id: "A-1029", amount: 49.0) - Server->>Server: requires approval → pause - - Server-->>+SDK: SSE waiting (tool, arguments) - SDK->>+You: on_approval(request) - You-->>-SDK: request.approve - SDK->>+Server: POST /agent/id/respond approved - Server-->>-SDK: ok - deactivate SDK - - Server->>+SDK: task issue_refund - SDK->>SDK: issue_refund runs - SDK-->>-Server: result - - Server->>+LLM: tool result - LLM-->>-Server: "Refunded $49." - - Server-->>+SDK: SSE text - SDK->>SDK: execution.partial_text appended - deactivate SDK - - Server-->>+SDK: SSE done - SDK->>SDK: execution.done? = true - SDK->>+You: on_done("Refunded $49.") - You-->>-SDK: block returns - deactivate SDK - - You->>+SDK: execution.finish_reason - SDK-->>-You: :stop -``` - -Had `on_approval` called `request.reject`, the server skips the tool and finishes with `finish_reason == :rejected`. Same flow with `call_sync`: the first activation bar simply stays open until done and the answer is the return value. - -### 3. Team + secret - -```ruby -tool def create_issue(title: String, body: '') - Github.create_issue(title, body, token: secret('GH_TOKEN')) -end - -triage = Agent.new( - name: 'triage', - model: 'openai/gpt-4o-mini', - instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.' -) - -filer = Agent.new( - name: 'filer', - model: 'anthropic/claude-sonnet-4-5', - instructions: 'File the bug as a GitHub issue.' -) -filer.add_tool :create_issue - -triage.hands_off_to filer, on: 'ACTIONABLE' - -team = Agent.new(name: 'bug_desk') -team.add_agent triage -team.add_agent filer - -puts team.call_sync(File.read('report.md')) -``` - -Make the agent, then give it things. `secret('GH_TOKEN')` in the tool body both reads the secret and declares it — the SDK scans for it at `tool def` and the server attaches the value to each task. Optional: `filer.redact %w[password api_key]`, `filer.stop_when 'ISSUE_FILED'`, `filer.stop_after messages: 12`, `team.strategy = :sequential`. Every line is sugar over the Python-parity objects (`Handoff::OnTextMention`, `RegexGuardrail`, `Termination::*`). Spec: `docs/design/AGENT_TEAMS.md`. - -```mermaid -sequenceDiagram - actor You - participant SDK as Ruby SDK - participant Server as Conductor Server - participant Triage as triage (gpt-4o-mini) - participant Filer as filer (claude) - - You->>+SDK: team.call_sync(report) (blocks until done) - SDK->>+Server: POST /agent/start (team agentConfig: triage, filer) - Server-->>-SDK: executionId, requiredWorkers [create_issue] - SDK->>SDK: start worker for create_issue, TaskDef.runtime_metadata = [GH_TOKEN] - SDK->>Server: GET /agent/stream (SSE) - Server->>+Triage: report - Triage-->>-Server: "... ACTIONABLE" - Server->>Server: hands_off_to filer matched - Server->>+Filer: conversation so far + create_issue schema - Filer-->>-Server: call create_issue(title, body) - Server->>Server: resolve GH_TOKEN from secret store - Server->>SDK: task create_issue, runtime_metadata GH_TOKEN=ghp_... - SDK->>SDK: secret("GH_TOKEN") → Github.create_issue - SDK-->>Server: issue url - Server->>+Filer: tool result - Filer-->>-Server: "Filed: github.com/.../issues/42" - Server-->>SDK: SSE text - Server-->>SDK: SSE done - SDK-->>-You: "Filed: github.com/.../issues/42" -``` - -## Secrets - -```mermaid -classDiagram - direction LR - class Integration { - <> - holds LLM provider key - model "openai/gpt-4o" → integration "openai" - } - class Agent { - +String model "openai/gpt-4o" - +List~String~ credentials inherited by sub-agents and tools - } - class Tools { - -scan_secrets(method) literal secret() names - } - class ToolDef { - +List~String~ credentials - } - class TaskDef { - +Hash runtime_metadata names only - } - class Task { - +Hash runtime_metadata name → plaintext, wire-only - } - class Dispatch { - -bind_secrets(task, tool_def) - } - class Secrets { - <> - +secret(name) String - +secrets_env(*names) Hash for system, spawn, Open3 - falls back to ENV - } - class CredentialNotFoundError - - Agent ..> Integration : model prefix names it - Tools ..> ToolDef : scan_secrets at tool def - Agent ..> ToolDef : credentials copied down at serialize - ToolDef ..> TaskDef : credentials copied to runtime_metadata at register - TaskDef ..> Task : server fills values at poll - Dispatch ..> Task : reads runtime_metadata - Dispatch ..> Secrets : binds for this call - Secrets ..> CredentialNotFoundError : missing everywhere -``` - -| Secret | Held by | You write | -|---|---|---| -| Conductor auth | your env | `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET` (existing `Configuration`, token cache moves to instance level) | -| LLM provider key | server Integration | `model: 'openai/gpt-4o'` — `openai` is the integration name. SDK never sees the key. | -| Tool credential | server secret store | `secret('GH_TOKEN')` in the tool body — read and declaration in one | - -### Flow: tool credential - -```mermaid -sequenceDiagram - participant SDK as Ruby SDK - participant Server as Conductor Server - participant Store as Secret Store - - Note over SDK,Server: at startup - SDK->>+Server: register TaskDef create_issue, runtimeMetadata [GH_TOKEN] - Server-->>-SDK: ok - - Note over Server: LLM calls create_issue - Server->>+Store: get GH_TOKEN - Store-->>-Server: ghp_... - Server->>+SDK: task create_issue, runtimeMetadata GH_TOKEN=ghp_... - SDK->>SDK: bind runtimeMetadata for this call (fiber-local) - SDK->>SDK: create_issue runs, secret("GH_TOKEN") → ghp_... - SDK-->>-Server: result (runtimeMetadata dropped) -``` - -### Use a secret in a tool - -```ruby -tool def create_issue(title: String, body: '') - Github.create_issue(title, body, token: secret('GH_TOKEN')) -end -``` - -That's it. `secret('GH_TOKEN')` is the read *and* the declaration: at `tool def` the SDK parses the method body, collects every literal `secret('...')`, and puts the names in the tool's contract (`TaskDef.runtimeMetadata`, `tool.config.credentials`). Python makes you write the list; Ruby reads it off the code. Same wire contract. - -### When the name isn't a literal - -```ruby -filer.add_tool :create_issue, credentials: ['GH_TOKEN'] # this tool -filer = Agent.new(..., credentials: ['GH_TOKEN']) # everything under this agent -``` - -### Tool that shells out - -```ruby -tool def gh_create_issue(title: String) - system(secrets_env('GH_TOKEN'), 'gh', 'issue', 'create', '--title', title) -end -``` - -`secrets_env('GH_TOKEN')` is `{ 'GH_TOKEN' => 'ghp_...' }` for this call, and declares the same way. Ruby's `system` / `spawn` / `Open3` take an env hash as the first argument, so only the child process sees it. We never write to `ENV`. - -### No secret store on the server? - -`secret('GH_TOKEN')` falls back to `ENV['GH_TOKEN']`. - -## Frameworks - -No official Ruby SDK exists for any of these. Not supported. - -| Framework | Ruby | -|---|---| -| OpenAI Agents SDK | N/A | -| Anthropic Claude Agent SDK | N/A | -| LangGraph | N/A | -| Google ADK | N/A | diff --git a/docs/design/AGENT_SECRETS.md b/docs/design/AGENT_SECRETS.md deleted file mode 100644 index 30b50cc..0000000 --- a/docs/design/AGENT_SECRETS.md +++ /dev/null @@ -1,134 +0,0 @@ -# Secrets - -Three kinds. Three different owners. None of them live in your code. - -| Secret | Example | Who holds it | You write | -|---|---|---|---| -| Conductor auth | key + secret for the server | your env | `CONDUCTOR_AUTH_KEY`, `CONDUCTOR_AUTH_SECRET` | -| LLM provider key | OpenAI / Anthropic API key | Conductor server, as an Integration | `model: 'openai/gpt-4o'` | -| Tool credential | GitHub token your tool needs | Conductor server, in the secret store | `credentials :create_issue, 'GH_TOKEN'` + `secret('GH_TOKEN')` | - -## 1. Conductor auth - -``` -export CONDUCTOR_SERVER_URL=https://play.orkes.io/api -export CONDUCTOR_AUTH_KEY=... -export CONDUCTOR_AUTH_SECRET=... -``` - -Existing `Configuration`. Nothing new. - -## 2. LLM provider key - -```ruby -agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', instructions: '...') -``` - -`openai` is not a vendor name, it's the **name of an Integration on the server**. Someone -added the OpenAI key there once (UI, or `IntegrationClient#save_integration`). The SDK never -sees the key. Swap `openai/gpt-4o` for `anthropic/claude-sonnet-4-5` and nothing else changes. - -Same as Python. Same as the existing `llm_chat` DSL task (`llmProvider`). - -## 3. Tool credential - -```ruby -tool def create_issue(title: String, body: '') - Github.create_issue(title, body, token: secret('GH_TOKEN')) -end -``` - -That's the whole thing. `secret('GH_TOKEN')` in the body is both the read *and* the -declaration. At `tool def`, the SDK parses the method body, finds every `secret('...')` with a -literal name, and puts those names in the tool's contract (`TaskDef.runtimeMetadata`, -`tool.config.credentials`). The server then attaches the values to each task for that tool. - -Store the value on the server once: `conductor secrets put GH_TOKEN ghp_...`, or the UI, or -`SecretClient#put_secret`. - -When the name isn't a literal, say it when attaching the tool: - -```ruby -filer.add_tool :create_issue, credentials: ['GH_TOKEN'] -``` - -### What happens - -```mermaid -sequenceDiagram - participant SDK as Ruby SDK - participant Server as Conductor Server - participant Store as Secret Store - - Note over SDK: at startup - SDK->>Server: register TaskDef create_issue, runtimeMetadata: [GH_TOKEN] - - Note over Server: LLM calls create_issue - Server->>Store: get GH_TOKEN - Store-->>Server: ghp_... - Server->>+SDK: task create_issue, runtimeMetadata: GH_TOKEN=ghp_... - SDK->>SDK: bind runtimeMetadata for this call (fiber-local) - SDK->>SDK: create_issue runs, secret('GH_TOKEN') → ghp_... - SDK-->>-Server: result, runtimeMetadata dropped -``` - -- Names go up at registration (`TaskDef.runtimeMetadata`). Values come down per task - (`Task.runtimeMetadata`, wire-only, never persisted with the task). -- `secret('X')` reads the map bound for the current call, else `ENV['X']`, else - `CredentialNotFoundError`. - -Same wire contract as Python's `credentials=[...]` + `get_secret()`. Python makes you write -the list; Ruby reads it off the code. - -## The one Ruby difference - -Python also injects secrets into `os.environ` for the duration of the call, so shell-out -tools (`gh`, `aws`) can find them. It can do that because Python workers are separate -processes. - -Ruby workers are threads in one process. Writing `ENV` from a tool would leak secrets into -every other tool running at the same time. So we don't. For subprocesses: - -```ruby -tool def gh_create_issue(title: String) - system(secrets_env('GH_TOKEN'), 'gh', 'issue', 'create', '--title', title) -end -``` - -`secrets_env('GH_TOKEN')` returns `{ 'GH_TOKEN' => 'ghp_...' }` for the current call, and -the literal name is picked up as a declaration the same way `secret()` is. Ruby's `system`, -`spawn`, and `Open3` all accept an env hash as the first argument — the child sees it, nobody -else does. - -## Explicit declaration - -Two cases where the SDK can't read the name off the code: - -```ruby -filer.add_tool :create_issue, credentials: ['GH_TOKEN'] # name is dynamic, or method typed in irb - -filer = Agent.new(..., credentials: ['GH_TOKEN']) # grant to every tool + sub-agent under this agent -``` - -Both add to whatever was detected. Same as Python's `Agent(credentials=[...])`. - -## No secret store on the server? - -`secret('X')` falls back to `ENV['X']`. Always on. - -## Cheat sheet - -```ruby -secret('KEY') # read — and this alone declares it -secrets_env('KEY_A', 'KEY_B') # Hash for system / spawn / Open3 — also declares -agent.add_tool :x, credentials: ['KEY'] # explicit, when the name isn't a literal -Agent.new(..., credentials: ['KEY']) # grant to everything under the agent -# no server value → ENV['KEY'], automatically -``` - -## Decided - -1. `secret('X')` is a bare helper. No `context:` arg. -2. Tool credentials are detected from `secret('LITERAL')` in the tool body at `tool def` (AST scan). Explicit `add_tool ..., credentials:` for dynamic names; agent-level `credentials:` grants to everything under it. -3. `secret('X')` falls back to `ENV['X']` when the server sends nothing. Always on. -4. `Configuration` auth-token cache moves from class level to instance level. Prerequisite. diff --git a/docs/design/AGENT_STREAMING.md b/docs/design/AGENT_STREAMING.md deleted file mode 100644 index 0cb8ff1..0000000 --- a/docs/design/AGENT_STREAMING.md +++ /dev/null @@ -1,153 +0,0 @@ -# Calling an agent: sync, async, approval - -One idea per step. Each step is a complete program. - -## 1. Sync — call and wait - -```ruby -agent = Agent.new( - name: 'support', - model: 'anthropic/claude-sonnet-4-5', - instructions: 'Help with orders.' -) - -answer = agent.call_sync('What is your return policy?') -puts answer -``` - -`call_sync` blocks until the agent is finished and returns the answer as a String. - -## 2. Async — call and get told when it's done - -```ruby -agent.call_async('What is your return policy?') do |answer| - Mailer.send(customer, answer) -end -``` - -Returns immediately. Your block runs with the answer when the agent finishes. - -## 3. Async — call and check on it yourself - -```ruby -execution = agent.call_async('What is your return policy?') - -execution.done? # false until finished -execution.partial_text # what has streamed in so far -execution.result # blocks until done, returns the answer -execution.cancel # stop it -``` - -Same `call_async`, no block. You get an `Execution` and poll it. - -## 4. Give it a tool - -```ruby -tool def issue_refund(order_id: String, amount: Float) - Billing.refund(order_id, amount) -end - -agent = Agent.new( - name: 'support', - model: 'anthropic/claude-sonnet-4-5', - instructions: 'Help with orders.' -) -agent.add_tool :issue_refund - -puts agent.call_sync('Refund order A-1029, it arrived broken') -``` - -The model calls `issue_refund`, the refund happens, the answer comes back. - -## 5. Make the tool wait for a human - -```ruby -requires_approval :issue_refund -``` - -One line, after the tool. Now the agent stops before running `issue_refund` and waits. - -## 6. Decide - -```ruby -agent.on_approval do |request| - if request.amount < 100 - request.approve - else - request.reject('Needs a manager') - end -end -``` - -`request` has the tool's arguments as methods (`request.amount`, `request.order_id`), plus -`approve` and `reject(reason)`. Works the same with `call_sync` (block runs while you wait) -and `call_async` (block runs on the background thread). - -No `on_approval` registered? The agent waits until something else approves it — a web UI, -another process, `execution.pending.approve`. - -## 7. What happened? - -```ruby -execution = agent.call_async('Refund order A-1029, it arrived broken') -execution.result - -execution.finish_reason # :stop, or :rejected if you rejected the tool -execution.tool_calls # [#] -execution.execution_id # look it up later: Execution.find(id) -``` - -## All together - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def issue_refund(order_id: String, amount: Float) - Billing.refund(order_id, amount) -end -requires_approval :issue_refund - -agent = Agent.new( - name: 'support', - model: 'anthropic/claude-sonnet-4-5', - instructions: 'Help with orders.' -) -agent.add_tool :issue_refund - -agent.on_approval do |request| - request.amount < 100 ? request.approve : request.reject('Needs a manager') -end - -agent.call_async('Refund order A-1029, it arrived broken') do |answer| - Mailer.send(customer, answer) -end -``` - -## Optional - -```ruby -agent.call_sync(question, session_id: 'cust-77') # same conversation across calls - -agent.before_tool_call { |call| log call.name } # hooks -agent.after_tool_result { |result| log result } -``` - -## Fire and forget - -`call_async` with no block, keep the `execution_id`, walk away. One catch: if the agent has -`tool def` tools, they run in *your* process, so it has to stay up. Fire-and-forget only works -when tools are all server-side (http / mcp / human) or you've `deploy`ed the agent and have -workers running elsewhere. Same rule as Python. - -## Under the hood - -`call_async(prompt, &on_done)`: - -1. POST `/agent/start`, register the tool workers the server asks for -2. return an `Execution`; open SSE `/agent/stream/{id}` on a background thread -3. text events append to `execution.partial_text` -4. waiting events build an `ApprovalRequest` and call your `on_approval` block -5. done event sets `done?`, `finish_reason`, `tool_calls`; calls `on_done(answer)` - -`call_sync` = `call_async(prompt).result`. diff --git a/docs/design/AGENT_TEAMS.md b/docs/design/AGENT_TEAMS.md deleted file mode 100644 index 67872ff..0000000 --- a/docs/design/AGENT_TEAMS.md +++ /dev/null @@ -1,75 +0,0 @@ -# Teams - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def create_issue(title: String, body: '') - Github.create_issue(title, body, token: secret('GH_TOKEN')) -end - -triage = Agent.new( - name: 'triage', - model: 'openai/gpt-4o-mini', - instructions: 'Read the bug report. Say ACTIONABLE if it should be filed.' -) - -filer = Agent.new( - name: 'filer', - model: 'anthropic/claude-sonnet-4-5', - instructions: 'File the bug as a GitHub issue.' -) -filer.add_tool :create_issue - -triage.hands_off_to filer, on: 'ACTIONABLE' - -team = Agent.new(name: 'bug_desk') -team.add_agent triage -team.add_agent filer - -puts team.call_sync(File.read('report.md')) -``` - -Make the agent. Then give it things. - -| Line | Does | -|---|---| -| `filer.add_tool :create_issue` | Give filer a tool. | -| `triage.hands_off_to filer, on: 'ACTIONABLE'` | When triage's answer contains ACTIONABLE, filer takes over. | -| `team.add_agent triage` | Give the team a member. First one added starts. `team.add_agents a, b` for several. | -| `secret('GH_TOKEN')` | Reads the secret — and declares it: the SDK sees the literal at `tool def` and tells the server this tool needs GH_TOKEN. Store it once with `conductor secrets put GH_TOKEN ghp_...`. Never in your code or env. | - -`tools:` and `agents:` still work as keyword args in `Agent.new` if you already have the list. -Same object either way. - -## If you need them - -```ruby -filer.redact %w[password api_key] # scrub these from output before anyone sees it -filer.stop_when 'ISSUE_FILED' # stop on this text -filer.stop_after messages: 12 # or after this many messages - -team.strategy = :sequential # one after another (default is handoff) -team.strategy = :parallel # all at once, merged -``` - -## Underneath - -Same Python objects, same `agentConfig`. Sugar only. - -| Sugar | Python-parity object | -|---|---| -| `a.add_tool :x` | appends to `Agent#tools`, same array `tools:` fills | -| `team.add_agent a` | appends to `Agent#agents`, same as `agents:` | -| `a.hands_off_to b, on: 'X'` | `handoffs: [Handoff::OnTextMention.new(target: 'b', text: 'X')]` | -| `a.redact %w[...]` | `guardrails: [RegexGuardrail.new(..., position: :output, on_fail: :fix)]` | -| `a.stop_when 'X'` / `a.stop_after messages: n` | `termination: TextMention \| MaxMessage` | -| `secret('X')` | at `tool def`: AST scan adds `X` to `ToolDef#credentials` → `TaskDef.runtimeMetadata` / `tool.config.credentials`. At run: reads `Task.runtimeMetadata['X']` (fiber-local). Same wire contract as Python `credentials=[...]` + `get_secret` | -| `Agent.new(name: 'bug_desk')` + `add_agent` | `strategy: :handoff` default | - -## Implementation note - -The server acts on `handoffs` only on the coordinator and only under `strategy: swarm`. So -`triage.hands_off_to filer` on a member makes a team with no explicit strategy serialize as a -`swarm` with the members' handoffs hoisted onto it (decision 2.19 in -`AGENTS_IMPLEMENTATION_PLAN.md`). An explicit `team.strategy = ...` is never overridden. diff --git a/docs/design/AGENT_TESTING.md b/docs/design/AGENT_TESTING.md deleted file mode 100644 index c6e78bb..0000000 --- a/docs/design/AGENT_TESTING.md +++ /dev/null @@ -1,109 +0,0 @@ -# Testing strategy - -How we test the agents SDK. The mocking layer is shared by every Conductor SDK — agentic and -plain-workflow tests alike; Ruby is the first adopter. WireMock records a real server once; -WireMock replays it in CI. No SDK ships a custom mock server. - -## Contract tests — no server - -Serializing a Ruby agent must produce exactly what the server (and Python) expect: - -```ruby -expect(JSONSchemer.schema(AGENT_SCHEMA).valid?(config)).to be true -expect(ConfigSerializer.serialize(agent)).to eq JSON.parse(fixture('05_handoffs.json')) -``` - -Fixtures vendored from python-sdk (`agent-schema.json`, the 19 golden `_configs/*.json`), -plus unit specs for schema generation, secret scanning, and serialization. This is also what -makes shared mocks work: replay matches by verb + path + order, so one recording serves every -SDK precisely because the SDKs send equivalent requests. - -## Runtime tests — record once, replay everywhere - -To create or refresh a recording: - -1. Run a Conductor server locally (docker, `localhost:8080`) with real provider keys - configured as integrations. -2. Run your SDK's test suite as usual, with the two record env vars set. Ruby example: - - ``` - CONDUCTOR_RECORD=1 \ - CONDUCTOR_RECORD_URL=http://localhost:8080 \ - bundle exec rspec spec/agents - ``` - - `CONDUCTOR_RECORD=1` switches the test helper into record mode: it boots WireMock between - the SDK and your server. `CONDUCTOR_RECORD_URL` says where your server is (shown value is - the default). -3. The test executes for real — the server calls the actual LLM, your tool methods run. -4. On green, WireMock has saved every response your server sent. A cleanup script swaps - run-specific values (execution ids, timestamps) for placeholders so the recording replays - for anyone. -5. The cleaned files are a scenario folder — open a PR to - [conductor-mocks](https://github.com/conductor-oss/conductor-mocks) with it. - -``` -mocks/ - agent/tool_happy_path/ - agent/approval_approve/ - agent/approval_reject/ - agent/secrets_runtime_metadata/ - agent/team_handoff/ - agent/mcp_discovery/ - workflow/simple_task/ - workflow/dynamic_fork/ -``` - -Re-recording is manual, and happens when we bump the supported server version — a re-record -is a reviewable diff. - -``` -record spec <-> WireMock proxy <-> your conductor-oss <-> REAL model - \-> normalized mappings -> conductor-mocks - -replay spec <-> wiremock/wiremock container + scenario mappings - no conductor, no LLM, no keys -``` - -In CI, the replayer is WireMock itself. Scenario states keep responses in recorded order; -`POST /tasks` only matches the stub with the recorded body, so a wrong tool result matches -nothing; any unmatched request fails the run. WireMock buffers SSE — fine, because an agent -stream terminates at `done`, so the captured body is complete. Recorded model text is that -day's output, so tests assert structure (`finish_reason`, tool args), not prose. - -### The weather test (Ruby) - -```ruby -RSpec.describe 'weather agent', mocks: 'agent/tool_happy_path' do - tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } - end - - it 'answers with the tool' do - agent = Agent.new(name: 'weather', model: 'openai/gpt-4o', - instructions: 'Answer weather questions.') - agent.add_tool :get_weather - expect(agent.call_sync('Weather in Lisbon?')).to include('Lisbon') - end -end -``` - -The `mocks:` tag boots WireMock with that scenario and points the SDK at it. - -## CI wiring - -Each SDK's workflow checks out conductor-mocks and starts the official image — every push, -fork PRs included: - -```yaml -- uses: actions/checkout@v4 - with: { repository: conductor-oss/conductor-mocks, path: conductor-mocks } -- run: docker run -d -p 8080:8080 - -v $PWD/conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock - wiremock/wiremock:3x -- run: CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/agents -``` - -conductor-mocks PRs are reviewed by the conductor-oss team; nothing replay serves was -invented by hand, and no SDK internals are ever stubbed. (Housekeeping: drop the unused -vcr/webmock dev deps from this repo.) diff --git a/docs/design/AGENT_TOOLS_DSL.md b/docs/design/AGENT_TOOLS_DSL.md deleted file mode 100644 index abeff6b..0000000 --- a/docs/design/AGENT_TOOLS_DSL.md +++ /dev/null @@ -1,106 +0,0 @@ -# Tools - -## Whole thing, one file - -```ruby -require 'conductor/agents' -include Conductor::Agents - -tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } -end - -agent = Agent.new( - name: 'weather', - model: 'openai/gpt-4o', - instructions: 'Answer weather questions.' -) -agent.add_tool :get_weather - -puts agent.call_sync('Weather in Lisbon?') -``` - -Run it: `ruby weather.rb`. That's the whole thing. - -## What each line does - -| Line | Does | -|---|---| -| `tool def get_weather(...)` | Makes the method a tool. `tool` sees the name `:get_weather` because `def` returns it (same as `private def`). | -| `city: String` | Required string. | -| `units: 'metric'` | Optional string, default `metric`. | -| `agent.add_tool :get_weather` | Give the agent the tool by name. | -| `agent.call_sync('prompt')` | Run it, wait, return the answer. Server URL / auth come from `CONDUCTOR_*` env vars. | - -The description the LLM sees is the method name: `get_weather` → "Get weather". - -## Types - -Whatever you put as the default is the type: - -| You write | Meaning | -|---|---| -| `city: String` | required string | -| `amount: Float` | required number | -| `count: Integer` | required integer | -| `tags: [String]` | required list of strings | -| `units: 'metric'` | optional, default `"metric"` | -| `limit: 10` | optional, default `10` | -| `verbose: false` | optional, default `false` | -| `units: %w[metric imperial]` | optional, one of these | - -## More than a few tools? Put them in a module - -```ruby -module Weather - extend Conductor::Agents::Tools - - tool def current(city: String) ... end - tool def forecast(city: String, days: 3) ... end -end - -agent.add_tools Weather # both -agent.add_tool Weather[:current] # one -``` - -## Options - -Approval and secrets are one line each, after the tool: - -```ruby -tool def issue_refund(order_id: String, amount: Float) - Stripe.refund(order_id, amount, key: secret('STRIPE_KEY')) -end -requires_approval :issue_refund -``` - -- `requires_approval :name` — a human approves before it runs. See `AGENT_STREAMING.md`. -- `secret('KEY')` — reads a secret. Writing it in the body is also the declaration: the SDK - finds it at `tool def` and tells the server this tool needs `STRIPE_KEY`. See `AGENT_SECRETS.md`. -- `describe :name, '...'` — override the description if the method name isn't enough. - -## Rules - -- Keyword args only (`city:`). Positional args raise at load. -- It's still a normal method: `get_weather(city: 'Lisbon')` works in tests. - -## Already have RubyLLM tools? - -They work as-is: - -```ruby -class Weather < RubyLLM::Tool - desc "Gets current weather for a location" - def execute(latitude:, longitude:) ... end -end - -agent.add_tool Weather # a RubyLLM::Tool class, as-is -``` - -RubyLLM already builds the JSON schema from `execute`'s keyword args and exposes `name` and -`description`; we read those into a `ToolDef` and run `Weather.new.execute(**args)` as the -worker body. Optional dependency — only loaded if `RubyLLM` is defined. - -RubyLLM is not the engine underneath. It runs the LLM loop client-side with your provider key; -Conductor runs it server-side with the key held as an integration. Only the tool class and the -`call_sync` / `call_async` API shape are shared. diff --git a/docs/design/EVENT_INTERCEPTOR_SYSTEM.md b/docs/design/EVENT_INTERCEPTOR_SYSTEM.md deleted file mode 100644 index 3ea5885..0000000 --- a/docs/design/EVENT_INTERCEPTOR_SYSTEM.md +++ /dev/null @@ -1,907 +0,0 @@ -# Event-Driven Interceptor System - Design Document - -## Table of Contents - -1. [Overview](#overview) -2. [Architecture](#architecture) -3. [Core Components](#core-components) -4. [Event Hierarchy](#event-hierarchy) -5. [Event Dispatcher](#event-dispatcher) -6. [Listener Protocol](#listener-protocol) -7. [Listener Registration](#listener-registration) -8. [Metrics Collection](#metrics-collection) -9. [Prometheus Integration](#prometheus-integration) -10. [Usage Examples](#usage-examples) -11. [Advanced Use Cases](#advanced-use-cases) -12. [Performance Considerations](#performance-considerations) -13. [File Structure](#file-structure) - ---- - -## Overview - -### Purpose - -The Event-Driven Interceptor System provides a decoupled, extensible mechanism for observing and reacting to task execution lifecycle events in the Conductor Ruby SDK. This enables: - -- **Metrics Collection** - Track poll times, execution durations, error rates -- **Custom Interceptors** - Add logging, tracing, auditing without modifying core code -- **SLA Monitoring** - Alert on tasks exceeding thresholds -- **Cost Tracking** - Monitor compute costs per task type -- **Error Tracking** - Send failures to external services (Sentry, Bugsnag, etc.) - -### Design Goals - -| Goal | Description | -|------|-------------| -| **Decoupled** | Event publishing is separate from event handling | -| **Thread-Safe** | Safe for concurrent task execution | -| **Extensible** | Add listeners without modifying SDK code | -| **Non-Blocking** | Listener failures never block worker execution | -| **Type-Safe** | Clear event contracts with documented attributes | -| **Pluggable** | Multiple metrics backends (null, Prometheus, custom) | - -### Non-Goals - -- **Distributed Tracing** - OpenTelemetry integration is a separate concern -- **Built-in Dashboards** - Users provide their own visualization -- **Async Dispatch** - Events are dispatched synchronously for simplicity - ---- - -## Architecture - -### High-Level Overview - -``` -┌─────────────────────────────────────────────────────────────────────────┐ -│ Task Execution Layer │ -│ ┌──────────────────┐ ┌──────────────────┐ │ -│ │ TaskRunner │ │ TaskHandler │ │ -│ │ (polling loop) │ │ (orchestrator) │ │ -│ └────────┬─────────┘ └────────┬─────────┘ │ -│ │ publish() │ register() │ -└───────────┼──────────────────────────────┼──────────────────────────────┘ - │ │ - ▼ ▼ -┌─────────────────────────────────────────────────────────────────────────┐ -│ Event Dispatch Layer │ -│ ┌────────────────────────────────────────────────────────────────────┐ │ -│ │ SyncEventDispatcher │ │ -│ │ • Thread-safe listener registration (Mutex) │ │ -│ │ • Synchronous event dispatch │ │ -│ │ • Error isolation (listener failures logged, not propagated) │ │ -│ │ • Type-based routing (event.class → listeners) │ │ -│ └──────────────────────────┬─────────────────────────────────────────┘ │ -│ │ dispatch │ -└──────────────────────────────┼──────────────────────────────────────────┘ - │ -┌──────────────────────────────▼──────────────────────────────────────────┐ -│ Listener/Consumer Layer │ -│ ┌────────────────┐ ┌────────────────┐ ┌─────────────────────────┐ │ -│ │MetricsCollector│ │ CustomListener │ │ SLA Monitor │ │ -│ │ (Prometheus) │ │ (Logging) │ │ (Alerting) │ │ -│ └────────────────┘ └────────────────┘ └─────────────────────────┘ │ -│ ┌────────────────┐ ┌────────────────┐ ┌─────────────────────────┐ │ -│ │ Audit Logger │ │ Cost Tracker │ │ Error Reporter │ │ -│ │ (Compliance) │ │ (FinOps) │ │ (Sentry/Bugsnag) │ │ -│ └────────────────┘ └────────────────┘ └─────────────────────────┘ │ -└─────────────────────────────────────────────────────────────────────────┘ -``` - -### Component Interaction Flow - -``` -TaskRunner SyncEventDispatcher Listeners - │ │ │ - │ publish(PollStarted) │ │ - │─────────────────────────────>│ │ - │ │ call(event) ──────────────>│ MetricsCollector - │ │ call(event) ──────────────>│ CustomListener - │ │<───────────────────────────│ - │<─────────────────────────────│ │ - │ │ │ - │ (execute task) │ │ - │ │ │ - │ publish(TaskExecutionCompleted) │ - │─────────────────────────────>│ │ - │ │ call(event) ──────────────>│ MetricsCollector - │ │ call(event) ──────────────>│ CustomListener - │ │<───────────────────────────│ - │<─────────────────────────────│ │ -``` - ---- - -## Core Components - -### Component Summary - -| Component | Location | Purpose | -|-----------|----------|---------| -| `ConductorEvent` | `events/conductor_event.rb` | Base event class with timestamp | -| `TaskRunnerEvent` | `events/conductor_event.rb` | Base for task runner events | -| `PollStarted`, etc. | `events/task_runner_events.rb` | Specific event types | -| `SyncEventDispatcher` | `events/sync_event_dispatcher.rb` | Thread-safe event router | -| `TaskRunnerEventsListener` | `events/listeners.rb` | Listener protocol (duck typing) | -| `ListenerRegistry` | `events/listener_registry.rb` | Bulk listener registration | -| `MetricsCollector` | `telemetry/metrics_collector.rb` | Canonical metric collector | -| `NullBackend` | `telemetry/metrics_collector.rb` | No-op backend | -| `PrometheusBackend` | `telemetry/prometheus_backend.rb` | Prometheus backend with canonical label schemas | -| `MetricsServer` | `telemetry/prometheus_backend.rb` | WEBrick HTTP server for `/metrics` | - ---- - -## Event Hierarchy - -### Class Hierarchy - -``` -ConductorEvent # Base - provides timestamp -└── TaskRunnerEvent # Base for task runner - adds task_type - ├── PollStarted # Polling started - ├── PollCompleted # Polling completed successfully - ├── PollFailure # Polling failed - ├── TaskExecutionStarted # Task execution started - ├── TaskExecutionCompleted # Task execution completed - ├── TaskExecutionFailure # Task execution failed - └── TaskUpdateFailure # Task result update failed (CRITICAL) -``` - -### Event Attributes - -#### ConductorEvent (Base) - -```ruby -class ConductorEvent - attr_reader :timestamp # Time - UTC timestamp when event was created - - def to_h - { timestamp: @timestamp.iso8601(3) } - end -end -``` - -#### TaskRunnerEvent (Base) - -```ruby -class TaskRunnerEvent < ConductorEvent - attr_reader :task_type # String - Task definition name -end -``` - -#### PollStarted - -Published when polling starts for a task type. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `worker_id` | String | Unique worker identifier | -| `poll_count` | Integer | Number of polls performed so far | - -#### PollCompleted - -Published when polling completes successfully. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `duration_ms` | Float | Duration of poll in milliseconds | -| `tasks_received` | Integer | Number of tasks received | - -#### PollFailure - -Published when polling fails. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `duration_ms` | Float | Duration of poll in milliseconds | -| `cause` | Exception | The exception that caused the failure | - -#### TaskExecutionStarted - -Published when task execution starts. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `task_id` | String | Unique task identifier | -| `worker_id` | String | Unique worker identifier | -| `workflow_instance_id` | String | Workflow instance identifier | - -#### TaskExecutionCompleted - -Published when task execution completes successfully. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `task_id` | String | Unique task identifier | -| `worker_id` | String | Unique worker identifier | -| `workflow_instance_id` | String | Workflow instance identifier | -| `duration_ms` | Float | Duration of execution in milliseconds | -| `output_size_bytes` | Integer | Size of output data in bytes (optional) | - -#### TaskExecutionFailure - -Published when task execution fails. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `task_id` | String | Unique task identifier | -| `worker_id` | String | Unique worker identifier | -| `workflow_instance_id` | String | Workflow instance identifier | -| `duration_ms` | Float | Duration of execution in milliseconds | -| `cause` | Exception | The exception that caused the failure | -| `is_retryable` | Boolean | Whether the error is retryable | - -#### TaskUpdateFailure (CRITICAL) - -Published when task result update fails after all retries. This is a **critical** event - the task result is lost. - -| Attribute | Type | Description | -|-----------|------|-------------| -| `task_type` | String | Task definition name | -| `task_id` | String | Unique task identifier | -| `worker_id` | String | Unique worker identifier | -| `workflow_instance_id` | String | Workflow instance identifier | -| `cause` | Exception | The exception that caused the failure | -| `retry_count` | Integer | Number of retry attempts made | -| `task_result` | TaskResult | The task result that failed to update (for recovery) | - ---- - -## Event Dispatcher - -### SyncEventDispatcher - -The `SyncEventDispatcher` is a thread-safe, synchronous event dispatcher that routes events to registered listeners. - -```ruby -module Conductor::Worker::Events - class SyncEventDispatcher - def initialize - @listeners = Hash.new { |h, k| h[k] = [] } - @mutex = Mutex.new - end - - # Register a listener for an event type - # @param event_type [Class] Event class to listen for - # @param listener [Proc, #call] Callable to invoke when event is published - # @return [self] - def register(event_type, listener) - @mutex.synchronize do - @listeners[event_type] << listener unless @listeners[event_type].include?(listener) - end - self - end - - # Unregister a listener for an event type - # @param event_type [Class] Event class - # @param listener [Proc, #call] Listener to remove - # @return [self] - def unregister(event_type, listener) - @mutex.synchronize do - @listeners[event_type].delete(listener) - end - self - end - - # Publish an event to all registered listeners - # @param event [ConductorEvent] Event to publish - # @return [self] - def publish(event) - listeners = @mutex.synchronize { @listeners[event.class].dup } - - listeners.each do |listener| - listener.call(event) - rescue StandardError => e - # Listener failure is isolated - never breaks the worker - warn "[Conductor] Event listener error for #{event.class}: #{e.message}" - end - - self - end - - # Check if there are listeners registered for an event type - def has_listeners?(event_type) - @mutex.synchronize { @listeners[event_type].any? } - end - - # Get the number of listeners for an event type - def listener_count(event_type) - @mutex.synchronize { @listeners[event_type].size } - end - - # Clear all listeners - def clear - @mutex.synchronize { @listeners.clear } - self - end - end -end -``` - -### Key Design Decisions - -| Decision | Rationale | -|----------|-----------| -| **Synchronous dispatch** | Simpler than async, avoids ordering issues | -| **Mutex for thread safety** | Protects listener list during registration and iteration | -| **Copy listeners before dispatch** | Allows modification during dispatch without deadlock | -| **Error isolation** | Listener exceptions are logged but don't propagate | -| **Type-based routing** | Events routed by class, not inheritance hierarchy | - -### Thread Safety Guarantees - -1. **Registration is thread-safe** - Multiple threads can register listeners concurrently -2. **Publishing is thread-safe** - Multiple threads can publish events concurrently -3. **Listeners are called sequentially** - Within a single publish call -4. **Listener exceptions are isolated** - One listener failure doesn't affect others - ---- - -## Listener Protocol - -### TaskRunnerEventsListener - -The listener protocol uses duck typing - implement only the methods you need: - -```ruby -module Conductor::Worker::Events - # Listener protocol for task runner events - # Include this module to document the expected interface - # All methods are optional - implement only the ones you need - module TaskRunnerEventsListener - # Called when polling starts - # @param event [PollStarted] - def on_poll_started(event); end - - # Called when polling completes successfully - # @param event [PollCompleted] - def on_poll_completed(event); end - - # Called when polling fails - # @param event [PollFailure] - def on_poll_failure(event); end - - # Called when task execution starts - # @param event [TaskExecutionStarted] - def on_task_execution_started(event); end - - # Called when task execution completes successfully - # @param event [TaskExecutionCompleted] - def on_task_execution_completed(event); end - - # Called when task execution fails - # @param event [TaskExecutionFailure] - def on_task_execution_failure(event); end - - # Called when task update fails after all retries (CRITICAL) - # @param event [TaskUpdateFailure] - def on_task_update_failure(event); end - end -end -``` - -### Implementation Example - -```ruby -class MyListener - # Only implement the methods you care about - def on_task_execution_completed(event) - puts "Task #{event.task_id} completed in #{event.duration_ms}ms" - end - - def on_task_execution_failure(event) - puts "Task #{event.task_id} FAILED: #{event.cause.message}" - end -end -``` - ---- - -## Listener Registration - -### ListenerRegistry - -The `ListenerRegistry` provides bulk registration of listener objects: - -```ruby -module Conductor::Worker::Events - class ListenerRegistry - # Mapping of event classes to listener method names - EVENT_METHOD_MAP = { - PollStarted => :on_poll_started, - PollCompleted => :on_poll_completed, - PollFailure => :on_poll_failure, - TaskExecutionStarted => :on_task_execution_started, - TaskExecutionCompleted => :on_task_execution_completed, - TaskExecutionFailure => :on_task_execution_failure, - TaskUpdateFailure => :on_task_update_failure - }.freeze - - # Register a listener object with the dispatcher - # Auto-detects implemented methods via respond_to? - # @param listener [Object] Object implementing TaskRunnerEventsListener methods - # @param dispatcher [SyncEventDispatcher] Event dispatcher - def self.register_task_runner_listener(listener, dispatcher) - EVENT_METHOD_MAP.each do |event_class, method_name| - if listener.respond_to?(method_name) - dispatcher.register(event_class, ->(event) { listener.send(method_name, event) }) - end - end - end - - # Register multiple listeners with the dispatcher - # @param listeners [Array] Array of listener objects - # @param dispatcher [SyncEventDispatcher] Event dispatcher - def self.register_all(listeners, dispatcher) - listeners.each do |listener| - register_task_runner_listener(listener, dispatcher) - end - end - end -end -``` - -### Usage in TaskHandler - -```ruby -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [MyListener.new, AnotherListener.new] -) -``` - ---- - -## Metrics Collection - -### MetricsCollector - -`MetricsCollector.create` returns a collector that emits the canonical -(harmonized) metric surface: - -```ruby -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) -``` - -See [docs/METRICS_AND_INTERCEPTORS.md](../METRICS_AND_INTERCEPTORS.md) for the -full metrics catalog and label reference. - -### Backend Protocol - -Metrics backends must implement these methods: - -```ruby -# Increment a counter -# @param name [String] Metric name -# @param labels [Hash] Metric labels -def increment(name, labels: {}) -end - -# Observe a value (histogram) -# @param name [String] Metric name -# @param value [Numeric] Value to observe -# @param labels [Hash] Metric labels -def observe(name, value, labels: {}) -end - -# Set a gauge value -# @param name [String] Metric name -# @param value [Numeric] Value to set -# @param labels [Hash] Metric labels -def set(name, value, labels: {}) -end -``` - -### NullBackend - -A no-op backend for when metrics are disabled: - -```ruby -class NullBackend - def increment(name, labels: {}); end - def observe(name, value, labels: {}); end - def set(name, value, labels: {}); end -end -``` - ---- - -## Prometheus Integration - -### Prometheus Backends - -The SDK ships `PrometheusBackend` with canonical metric registrations using -`taskType` labels, `status` on time histograms, and canonical bucket -boundaries. It implements `increment`, `observe`, and `set` and integrates -with the `prometheus-client` gem. See -[docs/METRICS_AND_INTERCEPTORS.md](../METRICS_AND_INTERCEPTORS.md) -for the full metric catalog emitted by each backend. - -### MetricsServer - -An optional HTTP server for exposing Prometheus metrics: - -```ruby -module Conductor::Worker::Telemetry - class MetricsServer - DEFAULT_PORT = 9090 - - def initialize(port: DEFAULT_PORT, registry: nil) - @port = port - @registry = registry || Prometheus::Client.registry - end - - def start - require 'webrick' - @server = WEBrick::HTTPServer.new(Port: @port, Logger: WEBrick::Log.new('/dev/null')) - - @server.mount_proc '/metrics' do |_req, res| - res.content_type = 'text/plain; version=0.0.4' - res.body = Prometheus::Client::Formats::Text.marshal(@registry) - end - - @server.mount_proc '/health' do |_req, res| - res.body = '{"status":"healthy"}' - end - - @thread = Thread.new { @server.start } - end - - def stop - @server&.shutdown - @thread&.join(5) - end - end -end -``` - ---- - -## Usage Examples - -### Basic Metrics Collection - -```ruby -require 'conductor' - -# Create configuration -config = Conductor::Configuration.new( - server_api_url: 'https://conductor.example.com/api', - key_id: 'key', - key_secret: 'secret' -) - -# Create metrics collector with Prometheus backend -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) - -# Start metrics server -metrics_server = Conductor::Worker::Telemetry::MetricsServer.new(port: 9090) -metrics_server.start - -# Create task handler with metrics -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [metrics] -) - -# Start workers -handler.start -handler.join -``` - -### Custom Logging Interceptor - -```ruby -class LoggingInterceptor - def initialize(logger = Logger.new($stdout)) - @logger = logger - end - - def on_poll_started(event) - @logger.debug("Polling for #{event.task_type}...") - end - - def on_poll_completed(event) - @logger.debug("Poll for #{event.task_type}: #{event.tasks_received} tasks in #{event.duration_ms}ms") - end - - def on_task_execution_started(event) - @logger.info("Starting task #{event.task_id} (#{event.task_type})") - end - - def on_task_execution_completed(event) - @logger.info("Completed task #{event.task_id} in #{event.duration_ms}ms") - end - - def on_task_execution_failure(event) - @logger.error("Task #{event.task_id} FAILED: #{event.cause.message}") - @logger.error(event.cause.backtrace.first(5).join("\n")) - end - - def on_task_update_failure(event) - @logger.fatal("CRITICAL: Task #{event.task_id} result LOST after #{event.retry_count} retries!") - end -end - -# Use with TaskHandler -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [LoggingInterceptor.new] -) -``` - -### Error Tracking (Sentry Integration) - -```ruby -class SentryInterceptor - def on_task_execution_failure(event) - Sentry.capture_exception(event.cause, extra: { - task_id: event.task_id, - task_type: event.task_type, - workflow_instance_id: event.workflow_instance_id, - duration_ms: event.duration_ms, - is_retryable: event.is_retryable - }) - end - - def on_task_update_failure(event) - Sentry.capture_message( - "CRITICAL: Task result lost", - level: :fatal, - extra: { - task_id: event.task_id, - task_type: event.task_type, - retry_count: event.retry_count - } - ) - end -end -``` - -### Multiple Listeners - -```ruby -# Combine metrics, logging, and error tracking -handler = Conductor::Worker::TaskHandler.new( - configuration: config, - event_listeners: [ - Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus), - LoggingInterceptor.new, - SentryInterceptor.new - ] -) -``` - ---- - -## Advanced Use Cases - -### SLA Monitor - -Monitor task execution times and alert on SLA violations: - -```ruby -class SLAMonitor - def initialize(thresholds:, alerter:) - @thresholds = thresholds # { 'task_type' => max_duration_ms } - @alerter = alerter - end - - def on_task_execution_completed(event) - threshold = @thresholds[event.task_type] - return unless threshold && event.duration_ms > threshold - - @alerter.alert( - type: :sla_violation, - task_type: event.task_type, - task_id: event.task_id, - duration_ms: event.duration_ms, - threshold_ms: threshold - ) - end -end - -# Usage -sla_monitor = SLAMonitor.new( - thresholds: { - 'process_order' => 5000, # 5 seconds - 'send_email' => 2000, # 2 seconds - 'generate_report' => 30000 # 30 seconds - }, - alerter: SlackAlerter.new(webhook_url: ENV['SLACK_WEBHOOK']) -) -``` - -### Cost Tracker - -Track compute costs per task type: - -```ruby -class CostTracker - def initialize(cost_per_ms:, reporting_interval: 60) - @cost_per_ms = cost_per_ms # { 'task_type' => cost_per_ms } - @costs = Hash.new(0.0) - @mutex = Mutex.new - @reporting_interval = reporting_interval - start_reporting_thread - end - - def on_task_execution_completed(event) - cost = (@cost_per_ms[event.task_type] || 0.0001) * event.duration_ms - @mutex.synchronize { @costs[event.task_type] += cost } - end - - private - - def start_reporting_thread - Thread.new do - loop do - sleep @reporting_interval - report_costs - end - end - end - - def report_costs - @mutex.synchronize do - total = @costs.values.sum - puts "Cost Report: Total=$#{format('%.4f', total)}" - @costs.each { |task_type, cost| puts " #{task_type}: $#{format('%.4f', cost)}" } - @costs.clear - end - end -end -``` - -### Audit Logger - -Log all task executions for compliance: - -```ruby -class AuditLogger - def initialize(log_file:) - @logger = Logger.new(log_file) - end - - def on_task_execution_started(event) - log_entry('STARTED', event) - end - - def on_task_execution_completed(event) - log_entry('COMPLETED', event, duration_ms: event.duration_ms) - end - - def on_task_execution_failure(event) - log_entry('FAILED', event, - duration_ms: event.duration_ms, - error: event.cause.class.name, - message: event.cause.message, - retryable: event.is_retryable) - end - - private - - def log_entry(status, event, extra = {}) - @logger.info({ - timestamp: event.timestamp.iso8601(3), - status: status, - task_type: event.task_type, - task_id: event.task_id, - worker_id: event.worker_id, - workflow_instance_id: event.workflow_instance_id, - **extra - }.to_json) - end -end -``` - ---- - -## Performance Considerations - -### Event Publishing Overhead - -Event publishing is synchronous but lightweight: - -1. **Mutex acquisition** - ~100ns on uncontended lock -2. **List copy** - O(n) where n = number of listeners (typically 1-5) -3. **Listener calls** - Dependent on listener implementation - -**Typical overhead**: < 1ms per event with 3 listeners - -### Recommendations - -| Concern | Recommendation | -|---------|----------------| -| **Many listeners** | Keep listener count low (< 10) | -| **Slow listeners** | Offload heavy work to background threads | -| **High-frequency events** | Consider sampling in custom listeners | -| **Logging** | Use async logging (Logger with queue) | -| **Metrics** | Prometheus client is thread-safe and efficient | - -### Thread Pool Sizing - -The event system doesn't use a separate thread pool. Events are processed in the TaskRunner thread. This means: - -- **Listener execution time** directly impacts polling interval -- **Blocking operations** in listeners will block task polling -- **Keep listeners fast** (< 10ms) or offload to background - -### Error Isolation Example - -```ruby -# If listener A fails, listener B still runs -class FailingListener - def on_task_execution_completed(event) - raise "Intentional failure" # This is caught and logged - end -end - -class WorkingListener - def on_task_execution_completed(event) - puts "Still runs!" # This executes even if FailingListener fails - end -end -``` - ---- - -## File Structure - -``` -lib/conductor/worker/ -├── events/ -│ ├── conductor_event.rb # Base event class + TaskRunnerEvent -│ ├── task_runner_events.rb # All task runner event types -│ ├── sync_event_dispatcher.rb # Thread-safe event dispatcher -│ ├── listeners.rb # TaskRunnerEventsListener protocol -│ └── listener_registry.rb # Bulk listener registration helper -├── telemetry/ -│ ├── metrics_collector.rb # MetricsCollector class + NullBackend -│ └── prometheus_backend.rb # PrometheusBackend + MetricsServer -├── task_runner.rb # Publishes events during polling/execution -└── task_handler.rb # Creates dispatcher, registers listeners - -spec/conductor/worker/ -├── events/ -│ ├── conductor_event_spec.rb -│ ├── task_runner_events_spec.rb -│ ├── sync_event_dispatcher_spec.rb -│ └── listener_registry_spec.rb -└── telemetry/ - ├── metrics_collector_spec.rb - └── prometheus_backend_spec.rb -``` - ---- - -## Comparison to Python SDK - -| Aspect | Python SDK | Ruby SDK | -|--------|------------|----------| -| **Dispatch model** | Async (asyncio.create_task) | Sync (same thread) | -| **Thread safety** | asyncio.Lock | Mutex | -| **Listener protocol** | typing.Protocol | Duck typing (respond_to?) | -| **Event classes** | @dataclass(frozen=True) | attr_reader + to_h | -| **Metrics backend** | Prometheus multiprocess | Prometheus single process | -| **Error isolation** | ✅ Caught and logged | ✅ Caught and logged | -| **Event types** | Same 7 event types | Same 7 event types | - -### Why Synchronous in Ruby? - -The Python SDK uses async dispatch because: -1. Python uses asyncio for workers -2. Async dispatch avoids blocking the event loop - -The Ruby SDK uses synchronous dispatch because: -1. Ruby workers use threads (GVL releases on I/O) -2. Simpler implementation with predictable ordering -3. Listeners typically complete in < 1ms -4. Thread-per-worker model already provides isolation diff --git a/docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md b/docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md deleted file mode 100644 index 3dcc014..0000000 --- a/docs/design/SDK_AGENT_TESTING_STRATEGY_V2.md +++ /dev/null @@ -1,554 +0,0 @@ -# SDK Agents Testing Strategy V2 - -> Status: local working draft for rapid iteration. This supersedes the WireMock assumptions in -> the current Confluence draft but does not update Confluence yet. - -## Decision - -SDK agent tests run against a real Conductor server and a deterministic `mockLLM` provider built -directly into the Conductor OSS test runtime. - -There is no WireMock, HTTP recording proxy, replay server, or separate request normalizer. Conductor -OSS owns the small mock provider implementation and its canonical JSON fixtures under -`src/test/resources/llm-mocks/`. Each SDK owns its agent definitions, workers, and assertions. - -The server supports two modes over that one fixture format: - -- **Replay mode** is the default for SDK development and CI. Agents use `mockLLM/`, and - `mockLLM` validates the provider-neutral request before returning the fixture's response. -- **Record mode** is an explicit local authoring mode. A JVM system property names the scenario, - and Conductor records the provider-neutral requests and responses from any real integration other - than `mockLLM` into the canonical fixture. - -Local development and CI use the same path: - -1. Start `conductor-server` from a pinned Conductor OSS revision with the test runtime enabled. -2. Register the no-secret `mockLLM` integration provider. -3. Register an agent whose model points to a named `mockLLM` scenario. -4. Run the agent through the real Conductor APIs and real task queues. -5. Assert the prompt, tool calls, tool results, guardrail behavior, and final execution result. - -Record mode is not part of the normal test run. It exists to create or refresh the single expected -fixture for a scenario, after which that fixture is replayed everywhere through `mockLLM`. - -## Goals - -- Exercise each SDK against the real Conductor server, agent compiler, execution engine, task - queues, and persistence layer. -- Make agent tests deterministic, fast, credential-free, and runnable on fork pull requests. -- Verify the exact SDK/server contract: serialized definitions, prompt wording, tool schemas and - arguments, tool results, guardrail outcomes, and final output. -- Keep the mocking implementation intentionally small and colocated with the server behavior it - replaces. -- Run identical scenarios locally and in CI. -- Make expected model wording, tool calls, and follow-up turns easy to create from a real provider - when a scenario is first authored or intentionally refreshed. - -## Non-goals - -- Testing OpenAI, Anthropic, or another provider's wire protocol. -- Recording or replaying raw provider HTTP traffic. -- Maintaining separate OpenAI, Anthropic, model, or provider-specific mock sets. -- Running record mode in CI. -- Verifying the natural-language quality of a real model response. -- Mocking Conductor's HTTP API, task queue, SSE implementation, or execution engine. -- Using a separate `conductor-mocks` repository for functional-test infrastructure. - -## Test layers - -### 1. Contract tests: SDK only, no server - -Contract tests remain ordinary unit tests. They catch SDK serialization and schema regressions -quickly without booting Conductor. - -Each SDK should assert that its agent configuration: - -- conforms to the published agent schema; -- serializes into the expected Conductor agent/workflow definition; -- generates the expected tool input schema; -- includes guardrails, secrets, handoffs, and other agent settings in the correct shape. - -Ruby example: - -```ruby -expect(JSONSchemer.schema(AGENT_SCHEMA).valid?(config)).to be true -expect(ConfigSerializer.serialize(agent)).to eq JSON.parse(fixture('weather_agent.json')) -``` - -These tests answer: **did the SDK build the correct definition?** - -### 2. Functional tests: SDK plus real Conductor plus `mockLLM` - -Functional tests boot a real server and exercise an agent end to end. Only the external model is -replaced. The `mockLLM` provider returns deterministic chat responses from a named fixture while all -orchestration remains real. - -These tests answer: **does this SDK-defined agent behave correctly when Conductor executes it?** - -### 3. Fixture authoring: real provider through Conductor record mode - -Record mode runs the same server and SDK scenario with a real, non-`mockLLM` integration. Conductor -captures the request and response after provider-specific translation has been removed from the -equation and writes the canonical fixture consumed by `mockLLM`. - -This is a developer workflow, not a test topology or CI dependency. A recorded fixture must be -reviewed and then replayed with `mockLLM` before it is committed. - -## Ownership and layout - -### Conductor OSS - -Conductor OSS owns the provider and the model responses because prompt construction, tool-call -translation, and guardrail orchestration happen in the server. - -Proposed layout (the exact module prefix may change during implementation): - -```text -conductor-oss/ -└── ai/ - └── src/ - └── test/ - ├── java/.../MockLLM.java - ├── java/.../MockLLMRecorder.java - └── resources/ - └── llm-mocks/ - ├── weather_tool_call.json - ├── input_guardrail_block.json - ├── output_guardrail_fix.json - └── team_handoff.json -``` - -`MockLLM` is a normal Conductor AI model implementation available only in the test runtime. It: - -- advertises the provider name `mockLLM`; -- uses the requested model name as the fixture name; -- loads that fixture from `src/test/resources/llm-mocks`; -- finds the single fixture turn whose `expect` block matches the provider-neutral request; -- returns the response through the same Conductor AI model interface as a real provider; and -- performs no network I/O and requires no credentials. - -Replay uses request content rather than a global call counter. That keeps concurrent SDK scenarios -isolated and makes a prompt, tool schema, tool call, or tool-result mismatch fail at the relevant -request rather than shifting every later response. - -A minimal fixture shape could be: - -```json -{ - "schemaVersion": 1, - "scenario": "weather_tool_call", - "turns": [ - { - "expect": { - "messages": [ - { - "role": "system", - "content": "Answer weather questions using the weather tool." - }, - { "role": "user", "content": "Weather in Lisbon?" } - ], - "tools": [ - { - "name": "get_weather", - "description": "Get the weather for a city" - } - ] - }, - "respond": { - "toolCall": { - "id": "call_weather_1", - "name": "get_weather", - "arguments": { "city": "Lisbon", "units": "metric" } - } - } - }, - { - "expect": { - "lastMessage": { - "role": "tool", - "name": "get_weather", - "content": { "temp_c": 21.0, "summary": "Sunny in Lisbon" } - } - }, - "respond": { "text": "Sunny in Lisbon, 21C." } - } - ] -} -``` - -The fixture is a provider-neutral request/response script, not a capture of an HTTP exchange. It -contains only semantic fields required to validate and drive the Conductor behavior under test. -Provider request envelopes, URLs, headers, authentication, completion IDs, timestamps, token usage, -and provider/model names are never written. - -### Record mode in Conductor - -Record mode is enabled when Conductor starts. The proposed JVM properties are: - -```text --Dconductor.ai.mock-llm.record= --Dconductor.ai.mock-llm.output-dir= -``` - -The first property enables recording and supplies the canonical scenario name. The output directory -may default to the Conductor OSS test-resource location for a source checkout, but CI and scripts -should pass it explicitly so the destination is unambiguous. - -When record mode is enabled, Conductor decorates its normal AI model invocation at the -provider-neutral boundary: - -1. The agent must reference a real integration such as OpenAI or Anthropic. Selecting `mockLLM` - fails immediately because recording a mock would create a mock of a mock. -2. Conductor builds the normal prompt, messages, tool schemas, and guardrail request. -3. The recorder forwards the call to the selected real provider. -4. The recorder retains only the stable `expect` request fields and semantic `respond` fields. -5. Tool calls execute normally. Their results naturally appear in the next recorded request. -6. Each completed turn is written atomically to `/.json`. - -Recording OpenAI and recording Anthropic both write exactly the same file shape and path. The chosen -provider is merely how a developer generates the desired wording and tool calls. It is not part of -fixture identity, and there is only one supported fixture set. - -Record one scenario at a time. Re-recording intentionally replaces that scenario's canonical file; -the resulting source diff is the review surface. - -### Each SDK - -Each SDK owns: - -- the test agent definition written through that SDK's public API; -- any SDK-hosted tool implementation or worker; -- the input used to start the agent; -- assertions over the registered definition and completed execution; and -- a small test helper that starts or connects to the Conductor test server and registers the - `mockLLM` integration. - -SDK tests do not start a mock HTTP service and do not understand the mock fixture internals. A test -selects a scenario only by using a model such as `mockLLM/weather_tool_call`. - -## Replay mode: one execution model for local development and CI - -```mermaid -sequenceDiagram - actor R as Local developer or SDK CI - participant C as conductor-server - participant M as mockLLM (in-process) - participant S as SDK test + tool workers - - R->>C: Start server with test runtime - C->>C: Load src/test/resources/llm-mocks - R->>C: Register mockLLM integration (no secret) - R->>S: Run SDK agent suite - - S->>C: Register agent using mockLLM/weather_tool_call - S->>C: Start agent with test input - C->>M: Chat request with system prompt, messages, and tool schemas - M-->>C: Scripted get_weather(city: Lisbon) tool call - C-->>S: Queue get_weather task - S->>S: Execute the real SDK tool implementation - S->>C: Complete task with tool result - C->>M: Chat request containing the tool result - M-->>C: Scripted final answer - C-->>S: Complete agent execution - - S->>C: Read execution and task details - S->>S: Assert prompt wording, tool call, tool result, guardrails, and final output - S-->>R: Pass or fail -``` - -This is the only functional-test topology. SSE and polling may both be exercised by SDK tests, but -they are merely two clients of the same real execution; neither needs a separate mock strategy or -sequence diagram. - -## Record mode: local fixture authoring - -```mermaid -sequenceDiagram - actor D as Developer - participant S as SDK scenario + tool workers - participant C as conductor-server - participant R as In-process recorder - participant L as Real non-mockLLM provider - - D->>C: Start with -Dconductor.ai.mock-llm.record=weather_tool_call - D->>S: Run one SDK scenario using OpenAI or Anthropic integration - S->>C: Register and start the agent - C->>R: Provider-neutral prompt, messages, and tool schemas - R->>L: Invoke the configured real provider - L-->>R: Generated tool call - R->>R: Save canonical expect + respond turn - R-->>C: Return the real provider response - C-->>S: Queue the generated tool call - S->>C: Complete the real tool with its result - C->>R: Next request containing that tool result - R->>L: Invoke the same configured provider - L-->>R: Generated final text - R->>R: Save canonical expect + respond turn atomically - R-->>C: Return final text - C-->>S: Complete agent execution - S-->>D: Scenario result; review the fixture diff - - Note over D,L: Provider-specific HTTP is never recorded.
OpenAI and Anthropic produce the same single fixture format. -``` - -After recording, restart Conductor without the record property, change the test agent back to -`mockLLM/weather_tool_call`, and run the same scenario in replay mode. The fixture is ready only when -that deterministic replay and the SDK assertions pass. - -## Guardrail execution - -Guardrails also run in the real server. A scenario supplies deterministic model output only when a -guardrail or the main agent needs an LLM response. - -```mermaid -sequenceDiagram - participant S as SDK test - participant C as conductor-server - participant M as mockLLM (in-process) - - S->>C: Register agent + guardrail using mockLLM scenario - S->>C: Start agent with controlled input - C->>C: Execute the real guardrail path - opt Guardrail requires an LLM decision - C->>M: Guardrail prompt and controlled content - M-->>C: Scripted allow, block, or fix decision - end - C->>C: Continue, stop, or rewrite according to guardrail result - C-->>S: Completed or blocked execution - S->>C: Read execution and guardrail task details - S->>S: Assert prompt text, decision, action, status, and visible output -``` - -For a blocked input guardrail, the test should also assert that the main agent LLM task and tool task -were never scheduled. For a fixing guardrail, it should assert both the original guardrail decision -and the exact rewritten text passed to the next stage. - -## What functional tests assert - -Assertions come from Conductor's persisted execution and task data plus observations made by the -SDK tool worker. They must not depend only on the final answer. - -For each scenario, assert the applicable items: - -1. **Registration** - - the agent definition registered successfully; - - the agent references the expected `mockLLM` integration and scenario; - - generated tool schemas and guardrail configuration match the test. -2. **Prompt construction** - - system/developer instructions contain the required exact wording; - - user messages and relevant conversation history appear in the expected order; - - tool descriptions and schemas exposed to the model are correct; - - runtime-only values are asserted structurally or by stable substring, not as fixed IDs. -3. **Tool execution** - - the expected tool is called exactly once unless the scenario says otherwise; - - arguments match exactly; - - the SDK worker returns the expected result; - - that result is present in the next LLM task input. -4. **Guardrails** - - the expected input/output/tool guardrail runs; - - its decision and configured action are correct; - - blocked paths do not execute forbidden downstream work; - - fixed content, rejection reasons, and final statuses are preserved. -5. **Completion** - - the execution reaches the expected terminal status; - - final text and finish reason match the deterministic fixture; - - no unexpected tool or LLM tasks were scheduled. - -Ruby-style example: - -```ruby -agent = Agent.new( - name: 'weather', - model: 'mockLLM/weather_tool_call', - instructions: 'Answer weather questions using the weather tool.' -) -agent.add_tool :get_weather - -execution = agent.call_async('Weather in Lisbon?') -result = execution.result -details = conductor.workflow_client.get_workflow(execution.execution_id, true) - -expect(recorded_tool_calls).to contain_exactly({ - name: 'get_weather', - arguments: { 'city' => 'Lisbon', 'units' => 'metric' } -}) -expect(llm_task(details).input_data.fetch('messages').to_json) - .to include('Answer weather questions using the weather tool.') -expect(result).to eq('Sunny in Lisbon, 21C.') -``` - -The final helper names will follow each SDK's existing APIs. The important contract is that tests -inspect the real execution rather than a mock server's request log. - -## Initial shared scenario catalog - -All SDKs should implement the same small behavioral matrix. The named model fixture in Conductor OSS -drives the response, while each SDK expresses and verifies the scenario through its own public API. - -| Scenario | Main behavior under test | Required assertions | -|---|---|---| -| `direct_answer` | Agent completes without a tool | Prompt text, final text, no tool tasks | -| `weather_tool_call` | One SDK tool call followed by an answer | Tool schema, exact args, tool result in next prompt, final text | -| `input_guardrail_block` | Input guardrail prevents execution | Guardrail prompt/result, blocked status, no main LLM/tool task | -| `output_guardrail_fix` | Output guardrail rewrites model text | Original output, fix decision, corrected final output | -| `approval_reject` | Tool call waits for and receives rejection | Pending approval, rejection reason, tool body not run, final status | -| `team_handoff` | One agent hands execution to another | Handoff target, prompts for both agents, final owner and output | - -Provider-specific protocol cases do not belong in this suite. Those remain Conductor provider adapter -tests. - -## Local workflow - -Provide one script or Gradle task in Conductor OSS that starts the test server with `mockLLM` on the -classpath. The SDK test helper should support either starting that command or connecting to an -already-running server. - -### Normal replay - -```bash -# terminal 1: conductor-oss -./gradlew - -# terminal 2: ruby-sdk -CONDUCTOR_SERVER_URL=http://localhost:8080/api \ -bundle exec rspec spec/agents -``` - -No LLM API key or mock-service URL is set. The SDK helper registers `mockLLM` idempotently before the -suite and waits for the Conductor health endpoint before running tests. - -### Create or refresh a fixture - -Start the built server with the record properties and configure one real provider integration in -the normal Conductor way: - -```bash -java \ - -Dconductor.ai.mock-llm.record=weather_tool_call \ - -Dconductor.ai.mock-llm.output-dir=/path/to/conductor/ai/src/test/resources/llm-mocks \ - -jar conductor-server.jar -``` - -Then run only the corresponding SDK scenario with its agent temporarily configured for a real -integration: - -```bash -CONDUCTOR_SERVER_URL=http://localhost:8080/api \ -bundle exec rspec spec/agents/weather_tool_call_spec.rb -``` - -Review the resulting `weather_tool_call.json`, stop the recording server, and rerun the scenario -against `mockLLM/weather_tool_call`. The recording invocation requires credentials only for the real -integration chosen by the developer. - -## CI wiring - -Each SDK agent job performs the same operations: - -1. Check out the SDK and a pinned Conductor OSS revision. -2. Build and start `conductor-server` with the test runtime and `mockLLM` fixtures. -3. Wait until the server is healthy. -4. Register the `mockLLM` integration. -5. Install the SDK toolchain and run its agent tests against the server. -6. Upload server logs and execution details on failure. - -Record mode must not be enabled in CI. CI consumes the committed canonical fixtures and never -configures a real LLM integration or provider credential. - -Illustrative workflow: - -```yaml -jobs: - agents: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@v4 - - uses: actions/checkout@v4 - with: - repository: conductor-oss/conductor - ref: - path: conductor-oss - - run: conductor-oss/gradlew - - run: bundle install - - run: bundle exec rspec spec/agents - env: - CONDUCTOR_SERVER_URL: http://localhost:8080/api -``` - -The final server-start step should run in the background, include a health wait, and guarantee cleanup. -It is shown compactly here because the exact Gradle task is part of the Conductor OSS implementation. - -## Failure behavior - -Failures should be direct and local: - -- an unknown model/scenario fails with the missing fixture path; -- no fixture turn matching the current request fails with a concise summary of message roles and - available selectors; -- malformed fixture JSON fails server startup or the first scenario use; -- enabling record mode while selecting `mockLLM` fails with an explicit unsupported-operation - message; -- record mode refuses to mix multiple scenario names into one output file; -- an SDK assertion shows the actual persisted prompt, tool call, guardrail result, or final output; -- an unexpected external provider selection fails because CI has no credentials and only `mockLLM` - is registered for these tests. - -## Removed from the previous design - -The following components and concepts are deleted: - -- WireMock and its Gradle record/replay modes; -- recorded HTTP mappings and `__files` payloads; -- post-processing raw provider traffic through a request normalizer; -- WireMock scenario state, request logs, and unmatched-request verification; -- `CONDUCTOR_RECORD` and `CONDUCTOR_RECORD_URL`; -- provider/model-specific recording directories; -- a reusable workflow whose primary purpose is to run WireMock; and -- the separate recording lifecycle in `conductor-mocks`. - -The replacement is one in-process replay provider, one in-process provider-neutral recorder, a -single set of deterministic JSON fixtures, and ordinary assertions against a real Conductor -execution. - -## Implementation slices - -1. **Conductor OSS test provider** - - implement `MockLLM` through the existing AI model interface; - - load named fixtures from `src/test/resources/llm-mocks`; - - expose it only in the test server profile/classpath; - - add provider unit tests for text, tool-call, missing-fixture, and concurrent scenario behavior. -2. **Conductor OSS record mode** - - add the `conductor.ai.mock-llm.record` and output-directory JVM properties; - - decorate non-`mockLLM` model calls at the provider-neutral boundary; - - atomically emit the canonical request/response fixture without provider metadata; - - reject `mockLLM` as a recording source; - - test that OpenAI-shaped and Anthropic-shaped adapters produce the same fixture schema. -3. **Conductor OSS test-server entry point** - - add a documented Gradle/script entry point; - - make no-secret `mockLLM` integration registration automatic or idempotent; - - expose a reliable health check. -4. **Ruby SDK first adopter** - - add the server-backed agent spec helper; - - implement the initial scenario catalog; - - assert persisted prompts, tool calls, guardrails, and results; - - add the agent job to CI. -5. **Other SDKs** - - reuse the same fixture names and behavioral assertions; - - implement only the language-specific agent definition, tool worker, and test adapter. - -## Acceptance criteria - -- The full Ruby agent suite passes locally with no provider credentials. -- The same suite passes in CI against a freshly started real Conductor server. -- No WireMock process, dependency, configuration, or mock repository is involved. -- Starting Conductor with `-Dconductor.ai.mock-llm.record=` and a real provider creates or - refreshes that scenario's canonical fixture. -- Recording the same scenario through OpenAI or Anthropic produces the same provider-neutral schema - and the same output path; only one version may be committed. -- Record mode rejects `mockLLM`, and CI never enables record mode. -- At least one test proves a real SDK tool is invoked and its result reaches the next LLM turn. -- At least one test proves a guardrail blocks or fixes content and the forbidden path does not run. -- At least one test asserts stable prompt wording from the persisted LLM task input. -- Tests fail clearly when the SDK changes a prompt, tool schema/argument, tool result, or guardrail - configuration unexpectedly. -- The test server makes zero outbound LLM network calls. diff --git a/docs/design/WORKER_DESIGN.md b/docs/design/WORKER_DESIGN.md deleted file mode 100644 index 1830ba2..0000000 --- a/docs/design/WORKER_DESIGN.md +++ /dev/null @@ -1,1781 +0,0 @@ -# Conductor Ruby SDK - Worker Infrastructure Design - -## Table of Contents - -1. [Overview & Goals](#overview--goals) -2. [Architecture Overview](#architecture-overview) -3. [Ruby vs Python Concurrency](#ruby-vs-python-concurrency) -4. [Core Components](#core-components) -5. [Three Runner Models](#three-runner-models) -6. [Worker Definition Patterns](#worker-definition-patterns) -7. [Task Context System](#task-context-system) -8. [Behavioral Algorithms](#behavioral-algorithms) -9. [Event System](#event-system) -10. [Configuration System](#configuration-system) -11. [Task Definition Auto-Registration](#task-definition-auto-registration) -12. [File Structure](#file-structure) -13. [Implementation Phases](#implementation-phases) - ---- - -## Overview & Goals - -### Purpose - -This document specifies the design for the Conductor Ruby SDK's worker infrastructure - the system that polls for tasks from a Conductor server, executes them using user-defined workers, and reports results back. - -### Goals - -1. **Full parity with Python SDK** - Support all features from the Python worker SDK including batch polling, adaptive backoff, capacity management, events, and metrics -2. **Ruby-idiomatic API** - Use Ruby conventions (blocks, mixins, snake_case) while maintaining the same capabilities -3. **Production-grade** - Handle edge cases, failures, and high-throughput scenarios reliably -4. **Extensible** - Support custom event listeners, metrics backends, and execution models -5. **Multiple concurrency models** - Support threads (default), Ractors (opt-in), and fibers (opt-in) - -### Non-Goals - -- Workflow definition DSL (covered in separate design) -- HTTP client implementation (already exists) -- Model serialization (already exists) - ---- - -## Architecture Overview - -### Component Hierarchy - -``` -┌─────────────────────────────────────────────────────────────────────┐ -│ User Code │ -│ (Worker classes, @worker_task methods, Worker.define blocks) │ -└─────────────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────────┐ -│ TaskHandler │ -│ • Discovers workers (registry + auto-scan) │ -│ • Resolves configuration (3-tier hierarchy) │ -│ • Creates one Thread/Ractor per worker │ -│ • Manages lifecycle (start/stop/join) │ -│ • Aggregates events/metrics │ -└─────────────────────────────────────────────────────────────────────┘ - │ - ┌─────────────┼─────────────┐ - ▼ ▼ ▼ -┌─────────────────────┐ ┌─────────────────────┐ ┌─────────────────────┐ -│ TaskRunner │ │ TaskRunner │ │ RactorTaskRunner │ -│ (Thread-based) │ │ (Thread-based) │ │ (Ractor-based) │ -│ │ │ │ │ │ -│ • ThreadPoolExecutor│ │ • FiberExecutor │ │ • Ractor isolation │ -│ • Batch polling │ │ (async gem) │ │ • Own HTTP client │ -│ • Capacity mgmt │ │ • Batch polling │ │ • Message passing │ -│ • Event publishing │ │ • Capacity mgmt │ │ • Event publishing │ -└─────────────────────┘ └─────────────────────┘ └─────────────────────┘ - │ │ │ - ▼ ▼ ▼ -┌─────────────────────────────────────────────────────────────────────┐ -│ TaskResourceApi │ -│ • poll_task / batch_poll │ -│ • update_task │ -│ • HTTP communication via ApiClient │ -└─────────────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────────┐ -│ Conductor Server │ -└─────────────────────────────────────────────────────────────────────┘ -``` - -### Comparison to Python SDK - -| Aspect | Python SDK | Ruby SDK | -|--------|-----------|----------| -| **Worker isolation** | One OS process per worker (`multiprocessing.Process`) | One Thread per worker (default), Ractor opt-in | -| **Task concurrency** | `ThreadPoolExecutor` (sync) or `asyncio` (async) | `Concurrent::ThreadPoolExecutor` (default), Fiber opt-in | -| **Why different** | Python GIL blocks ALL threads during CPU work | Ruby GVL releases during I/O (HTTP, sleep) | -| **Async model** | `asyncio` event loop, `async/await` | `async` gem with Fibers (opt-in) | -| **Event dispatch** | `SyncEventDispatcher` with `threading.Lock` | `SyncEventDispatcher` with `Mutex` | -| **Config resolution** | 3-tier: worker env → global env → code | Same | -| **HTTP client** | `requests` (sync), `httpx` (async) | `Faraday` with `net_http_persistent` | - ---- - -## Ruby vs Python Concurrency - -### Why Threads Work for Ruby Workers - -Python uses processes because the GIL (Global Interpreter Lock) prevents true thread parallelism for **any** Python code. Ruby's GVL (Global VM Lock) is similar but crucially different: - -**Ruby GVL releases during:** -- Network I/O (HTTP requests, socket operations) -- File I/O -- `sleep` calls -- C extension calls that release the GVL - -**Worker operations are I/O-bound:** -1. Poll HTTP endpoint (GVL released) -2. Execute worker (may be CPU-bound, but typically I/O) -3. Update HTTP endpoint (GVL released) - -This means Ruby threads provide **real concurrency** for typical worker workloads, making process-per-worker unnecessary overhead for most use cases. - -### When to Use Ractors - -Ractors provide true parallelism (no GVL sharing) but with restrictions: -- No shared mutable state between Ractors -- Limited gem compatibility (many gems use global state) -- Requires Ruby 3.1+ - -**Use Ractors when:** -- Worker performs CPU-intensive computation -- Worker doesn't need shared state -- All dependencies are Ractor-safe - -### When to Use Fibers - -Fibers provide lightweight cooperative concurrency within a single thread: -- ~400 bytes per fiber vs ~8KB per thread -- Can handle thousands of concurrent I/O operations -- Requires non-blocking I/O throughout - -**Use Fibers when:** -- Extremely high concurrency (hundreds of concurrent tasks) -- All operations are non-blocking (no blocking gem calls) -- Memory is constrained - ---- - -## Core Components - -### 1. TaskHandler - -The top-level orchestrator that manages all workers. - -```ruby -module Conductor - module Worker - class TaskHandler - # Initialize with optional workers and configuration - # @param workers [Array] Pre-created worker instances - # @param configuration [Configuration] Conductor configuration - # @param scan_for_annotated_workers [Boolean] Auto-discover @worker_task methods - # @param import_modules [Array] Ruby files/modules to require (triggers registration) - # @param event_listeners [Array] Custom event listeners - # @param metrics_settings [MetricsSettings] Metrics configuration - def initialize( - workers: nil, - configuration: nil, - scan_for_annotated_workers: true, - import_modules: nil, - event_listeners: nil, - metrics_settings: nil - ) - end - - # Start all worker threads/ractors - # @return [self] - def start - end - - # Stop all workers gracefully - # @param timeout [Integer] Seconds to wait before force-killing (default: 5) - # @return [self] - def stop(timeout: 5) - end - - # Wait for all workers to complete (blocking) - # @return [self] - def join - end - - # Check if handler is running - # @return [Boolean] - def running? - end - - # Get list of registered workers - # @return [Array] - def workers - end - end - end -end -``` - -**Responsibilities:** -1. Discover workers from registry + auto-scan -2. Resolve configuration for each worker (3-tier hierarchy) -3. Create appropriate runner (TaskRunner or RactorTaskRunner) based on config -4. Create one Thread (or Ractor) per worker -5. Manage lifecycle (start/stop/join) -6. Create shared EventDispatcher and register listeners -7. Optionally start MetricsProvider - -**Context Manager Pattern:** -```ruby -Conductor::Worker::TaskHandler.new(configuration: config) do |handler| - handler.start - handler.join -end -# Automatically calls stop on block exit -``` - -### 2. TaskRunner (Thread-based) - -The polling loop that runs in a dedicated Thread. - -```ruby -module Conductor - module Worker - class TaskRunner - # Initialize runner for a specific worker - # @param worker [Worker] The worker to run - # @param configuration [Configuration] Conductor configuration - # @param event_dispatcher [SyncEventDispatcher] Shared event dispatcher - # @param executor [Symbol] :thread_pool (default) or :fiber - def initialize(worker, configuration:, event_dispatcher:, executor: :thread_pool) - end - - # Main polling loop (runs until stopped) - def run - end - - # Single iteration of the polling loop - def run_once - end - - # Signal the runner to stop - def shutdown - end - - # Check if runner is running - # @return [Boolean] - def running? - end - end - end -end -``` - -**Internal State:** -```ruby -@worker # Worker instance -@configuration # Conductor configuration -@task_client # TaskClient for HTTP operations -@event_dispatcher # SyncEventDispatcher for publishing events -@executor # Concurrent::ThreadPoolExecutor or FiberExecutor -@running_tasks # Set of running futures/fibers -@consecutive_empty_polls # Counter for adaptive backoff -@auth_failures # Counter for auth failure backoff -@shutdown # AtomicBoolean for graceful shutdown -@last_poll_time # Time of last poll (for backoff calculation) -``` - -### 3. RactorTaskRunner - -The Ractor-based runner for CPU-bound workers requiring true parallelism. - -```ruby -module Conductor - module Worker - class RactorTaskRunner - # Initialize runner for a specific worker (runs inside Ractor) - # @param worker [Worker] The worker to run (must be Ractor-safe) - # @param configuration [Configuration] Conductor configuration (serializable parts only) - def initialize(worker, configuration:) - end - - # Main polling loop (creates HTTP client inside Ractor) - def run - end - - # Called by TaskHandler to receive events from Ractor - # @return [Array] Events from this poll cycle - def drain_events - end - end - end -end -``` - -**Key Differences from TaskRunner:** -1. Creates `TaskClient` **inside** `run()` (Ractors can't share objects) -2. Uses Ractor-local storage for TaskContext (not `Thread.current`) -3. Events are collected and sent to main Ractor via `Ractor.yield` for aggregation -4. No ThreadPoolExecutor - sequential execution within the Ractor (parallelism comes from multiple Ractors) - -### 4. Worker - -The user-facing worker definition that wraps an execute function. - -```ruby -module Conductor - module Worker - class Worker - attr_reader :task_definition_name, :execute_function, :config - attr_accessor :domain, :poll_interval, :thread_count, :worker_id, - :register_task_def, :overwrite_task_def, :strict_schema, - :paused, :poll_timeout, :isolation, :executor - - # Initialize a worker - # @param task_definition_name [String] Task type name in Conductor - # @param execute_function [Proc, Method] Function to execute tasks - # @param options [Hash] Worker configuration options - def initialize(task_definition_name, execute_function = nil, **options, &block) - end - - # Execute a task - # @param task [Task] The task to execute - # @return [TaskResult, TaskInProgress, Hash] Execution result - def execute(task) - end - - # Get polling interval in seconds - # @return [Float] - def polling_interval_seconds - end - - # Check if worker is async (for auto-detection, not used in Ruby) - # @return [Boolean] - def async? - end - end - end -end -``` - -**Execute Function Return Type Handling:** - -| Return Type | Behavior | -|-------------|----------| -| `TaskResult` | Use directly (set task_id, workflow_instance_id) | -| `TaskInProgress` | Create `IN_PROGRESS` result with `callback_after_seconds` | -| `Hash` | Wrap in `COMPLETED` TaskResult as output_data | -| `true` | `COMPLETED` with empty output | -| `false` | `FAILED` with empty output | -| `nil` | `COMPLETED` with empty output | -| Any other object | `COMPLETED` with `{ result: object }` output | -| Raises `NonRetryableError` | `FAILED_WITH_TERMINAL_ERROR` | -| Raises any `StandardError` | `FAILED` with error message | - -### 5. WorkerConfig - -Configuration resolver with 3-tier hierarchy. - -```ruby -module Conductor - module Worker - class WorkerConfig - # Resolve configuration for a worker - # @param worker_name [String] Task definition name - # @param defaults [Hash] Code-level defaults from worker definition - # @return [Hash] Resolved configuration - def self.resolve(worker_name, defaults = {}) - end - - # Configuration properties with types and defaults - PROPERTIES = { - poll_interval: { type: :integer, default: 100 }, # milliseconds - thread_count: { type: :integer, default: 1 }, - domain: { type: :string, default: nil }, - worker_id: { type: :string, default: -> { generate_worker_id } }, - poll_timeout: { type: :integer, default: 100 }, # milliseconds - register_task_def: { type: :boolean, default: false }, - overwrite_task_def: { type: :boolean, default: true }, - strict_schema: { type: :boolean, default: false }, - paused: { type: :boolean, default: false }, - isolation: { type: :symbol, default: :thread }, # :thread or :ractor - executor: { type: :symbol, default: :thread_pool } # :thread_pool or :fiber - }.freeze - end - end -end -``` - -**Resolution Priority (highest to lowest):** - -1. **Worker-specific environment variable:** - - `conductor.worker.{task_name}.{property}` (dotted) - - `CONDUCTOR_WORKER_{TASK_NAME}_{PROPERTY}` (uppercase) - -2. **Global worker environment variable:** - - `conductor.worker.all.{property}` (dotted) - - `CONDUCTOR_WORKER_ALL_{PROPERTY}` (uppercase) - -3. **Legacy environment variable:** - - `CONDUCTOR_WORKER_{PROPERTY}` (old format) - -4. **Code-level default:** - - Value passed to `worker_task` or `Worker.new` - -**Boolean Parsing:** Accepts `true/1/yes` and `false/0/no` (case-insensitive). - ---- - -## Three Runner Models - -### Model 1: TaskRunner with ThreadPoolExecutor (Default) - -``` -┌─────────────────────────────────────────────────────────────┐ -│ Worker Thread │ -│ ┌─────────────────────────────────────────────────────┐ │ -│ │ TaskRunner.run() │ │ -│ │ ┌────────────────────────────────────────────────┐ │ │ -│ │ │ Polling Loop │ │ │ -│ │ │ 1. Check capacity │ │ │ -│ │ │ 2. Adaptive backoff │ │ │ -│ │ │ 3. Batch poll │ │ │ -│ │ │ 4. Submit tasks to executor ─────────────────┐ │ │ │ -│ │ │ 5. Loop │ │ │ │ -│ │ └───────────────────────────────────────────────┘ │ │ │ -│ └──────────────────────────────────────────────────│─┘ │ -│ │ │ -│ ┌──────────────────────────────────────────────────▼─┐ │ -│ │ Concurrent::ThreadPoolExecutor │ │ -│ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │ -│ │ │ Thread 1│ │ Thread 2│ │ Thread 3│ │ Thread N│ │ │ -│ │ │ Task A │ │ Task B │ │ Task C │ │ idle │ │ │ -│ │ └─────────┘ └─────────┘ └─────────┘ └─────────┘ │ │ -│ │ (thread_count = N) │ │ -│ └────────────────────────────────────────────────────┘ │ -└─────────────────────────────────────────────────────────────┘ -``` - -**Configuration:** -```ruby -worker_task 'my_task', thread_count: 5, poll_interval: 100 -# or -Worker.new('my_task', thread_count: 5, executor: :thread_pool) -``` - -**Characteristics:** -- One dedicated thread for the polling loop -- ThreadPoolExecutor with `thread_count` threads for task execution -- GVL released during HTTP I/O, so threads provide real concurrency -- Best for: Most workloads (I/O-bound or mixed) - -### Model 2: TaskRunner with FiberExecutor (Opt-in) - -``` -┌─────────────────────────────────────────────────────────────┐ -│ Worker Thread │ -│ ┌─────────────────────────────────────────────────────┐ │ -│ │ TaskRunner.run() │ │ -│ │ ┌────────────────────────────────────────────────┐ │ │ -│ │ │ Async Event Loop (via async gem) │ │ │ -│ │ │ ┌───────┐ ┌───────┐ ┌───────┐ ┌───────┐ │ │ │ -│ │ │ │Fiber 1│ │Fiber 2│ │Fiber 3│ │Fiber N│ │ │ │ -│ │ │ │Task A │ │Task B │ │Task C │ │ poll │ │ │ │ -│ │ │ └───────┘ └───────┘ └───────┘ └───────┘ │ │ │ -│ │ │ (cooperative scheduling, single thread) │ │ │ -│ │ │ (thread_count = concurrency limit via Semaphore) │ │ -│ │ └────────────────────────────────────────────────┘ │ │ -│ └─────────────────────────────────────────────────────┘ │ -└─────────────────────────────────────────────────────────────┘ -``` - -**Configuration:** -```ruby -worker_task 'my_task', thread_count: 100, executor: :fiber -# Requires: gem 'async' in Gemfile -``` - -**Characteristics:** -- Single thread with fiber-based cooperative concurrency -- Requires `async` gem (optional dependency, loaded lazily) -- `thread_count` becomes fiber concurrency limit (semaphore) -- All I/O must be non-blocking (async gem provides non-blocking HTTP) -- Best for: Very high concurrency I/O-bound tasks (hundreds/thousands) - -**Lazy Loading:** -```ruby -# In fiber_executor.rb -def self.load_async_gem - require 'async' - require 'async/http' -rescue LoadError - raise Conductor::ConfigurationError, - "The 'async' gem is required for fiber executor. Add `gem 'async'` to your Gemfile." -end -``` - -### Model 3: RactorTaskRunner (Opt-in) - -``` -┌────────────────────────────────────────────────────────────────────────┐ -│ Main Thread (TaskHandler) │ -│ ┌──────────────────────────────────────────────────────────────────┐ │ -│ │ Event Aggregation Loop │ │ -│ │ • Receives events from Ractors via Ractor.receive │ │ -│ │ • Dispatches to shared EventDispatcher │ │ -│ └──────────────────────────────────────────────────────────────────┘ │ -│ ▲ ▲ ▲ │ -│ │ events │ events │ events │ -└───────────┼────────────────────┼────────────────────┼────────────────────┘ - │ │ │ -┌───────────┴──────┐ ┌─────────┴────────┐ ┌───────┴──────────┐ -│ Ractor 1 │ │ Ractor 2 │ │ Ractor 3 │ -│ ┌──────────────┐ │ │ ┌──────────────┐ │ │ ┌──────────────┐ │ -│ │RactorTaskRunner│ │ │ │RactorTaskRunner│ │ │ │RactorTaskRunner│ │ -│ │ Worker A │ │ │ │ Worker B │ │ │ │ Worker C │ │ -│ │ │ │ │ │ │ │ │ │ │ │ -│ │ Own HTTP │ │ │ │ Own HTTP │ │ │ │ Own HTTP │ │ -│ │ client │ │ │ │ client │ │ │ │ client │ │ -│ │ │ │ │ │ │ │ │ │ │ │ -│ │ Sequential │ │ │ │ Sequential │ │ │ │ Sequential │ │ -│ │ task exec │ │ │ │ task exec │ │ │ │ task exec │ │ -│ └──────────────┘ │ │ └──────────────┘ │ │ └──────────────┘ │ -│ (no GVL sharing) │ │ (no GVL sharing) │ │ (no GVL sharing) │ -└──────────────────┘ └──────────────────┘ └──────────────────┘ - True parallel execution across Ractors -``` - -**Configuration:** -```ruby -worker_task 'cpu_intensive_task', isolation: :ractor, thread_count: 4 -# Creates 4 Ractors, each running the same worker -``` - -**Characteristics:** -- True parallelism (each Ractor has its own GVL) -- HTTP client created inside Ractor (can't be shared) -- Events sent to main thread via Ractor messaging -- `thread_count` = number of Ractors (each processes one task at a time) -- Worker must be Ractor-safe (no shared mutable state) -- Requires Ruby 3.1+ -- Best for: CPU-intensive workers - -**Ractor Constraints:** -```ruby -# These will NOT work in Ractor-based workers: -- Global variables (@@, $) -- Class instance variables -- Mutating shared objects -- Many gems that use global state - -# These WILL work: -- Pure functions -- Immutable data -- Ractor-local state via Ractor.current[:key] -``` - ---- - -## Worker Definition Patterns - -### Pattern 1: Class-based with Module Mixin - -```ruby -class OrderProcessor - include Conductor::Worker::WorkerMixin - - worker_task 'process_order', - poll_interval: 200, - thread_count: 5, - domain: 'orders' - - def execute(task) - order_id = task.input_data['order_id'] - amount = task.input_data['amount'] - - # Process the order... - result = process(order_id, amount) - - # Return hash (auto-wrapped in COMPLETED TaskResult) - { processed: true, order_id: order_id, total: result.total } - end -end - -# Usage -handler = TaskHandler.new(workers: [OrderProcessor.new]) -``` - -### Pattern 2: Block-based with `Worker.define` - -```ruby -Conductor::Worker.define('send_notification', poll_interval: 100, thread_count: 3) do |task| - recipient = task.input_data['recipient'] - message = task.input_data['message'] - - NotificationService.send(to: recipient, body: message) - - { sent: true, recipient: recipient } -end - -# Workers registered automatically, discovered by TaskHandler -handler = TaskHandler.new(scan_for_annotated_workers: true) -``` - -### Pattern 3: Method Annotation with `worker_task` - -```ruby -module MyWorkers - extend Conductor::Worker::Annotatable - - worker_task 'greet_user', poll_interval: 50 - def self.greet(name:, greeting: 'Hello') - "#{greeting}, #{name}!" - end - - worker_task 'calculate_total', thread_count: 10 - def self.calculate(items:) - total = items.sum { |item| item['price'] * item['quantity'] } - { total: total, item_count: items.size } - end -end - -# Usage -handler = TaskHandler.new( - scan_for_annotated_workers: true, - import_modules: ['./lib/my_workers'] -) -``` - -### Pattern 4: Keyword Argument Mapping - -When the execute function has keyword arguments, they are automatically mapped from `task.input_data`: - -```ruby -# Worker definition with keyword args -worker_task 'process_payment' -def process_payment(order_id:, amount:, currency: 'USD') - # order_id, amount extracted from task.input_data - # currency uses default if not in input_data - PaymentGateway.charge(order_id, amount, currency) -end - -# Task input_data: { "order_id" => "123", "amount" => 99.99 } -# Mapped to: process_payment(order_id: "123", amount: 99.99, currency: 'USD') -``` - -### Pattern 5: Full Task Access - -```ruby -worker_task 'audit_task' -def audit(task) - # Full access to task object - puts "Task ID: #{task.task_id}" - puts "Workflow: #{task.workflow_instance_id}" - puts "Retry count: #{task.retry_count}" - puts "Poll count: #{task.poll_count}" - - # Access input - data = task.input_data - - # Return result - { audited: true } -end -``` - ---- - -## Task Context System - -TaskContext provides execution context accessible from anywhere in the worker code. - -### Thread-local Storage (TaskRunner) - -```ruby -module Conductor - module Worker - class TaskContext - # Get current context (thread-local) - # @return [TaskContext, nil] - def self.current - Thread.current[:conductor_task_context] - end - - # Set current context (internal use) - def self.current=(context) - Thread.current[:conductor_task_context] = context - end - - # Clear current context (internal use) - def self.clear - Thread.current[:conductor_task_context] = nil - end - - attr_reader :task, :task_result - - def initialize(task, task_result) - @task = task - @task_result = task_result - end - - # Convenience accessors - def task_id - @task.task_id - end - - def workflow_instance_id - @task.workflow_instance_id - end - - def retry_count - @task.retry_count || 0 - end - - def poll_count - @task.poll_count || 0 - end - - def input - @task.input_data || {} - end - - def task_def_name - @task.task_def_name - end - - # Mutable context methods - def add_log(message) - @task_result.log(message) - end - - def set_callback_after(seconds) - @task_result.callback_after_seconds = seconds - end - - def set_output(output_data) - @task_result.output_data = output_data - end - - def callback_after_seconds - @task_result.callback_after_seconds - end - end - end -end -``` - -### Fiber Storage (FiberExecutor) - -```ruby -# When using fiber executor, context stored in Fiber.current.storage -def self.current - if defined?(Fiber.current.storage) - Fiber.current.storage[:conductor_task_context] - else - Thread.current[:conductor_task_context] - end -end -``` - -### Ractor Storage (RactorTaskRunner) - -```ruby -# Ractors use Ractor.current for isolation -def self.current - Ractor.current[:conductor_task_context] -rescue - Thread.current[:conductor_task_context] -end -``` - -### Usage in Worker Code - -```ruby -worker_task 'my_task' -def execute(task) - ctx = Conductor::Worker::TaskContext.current - - # Log progress - ctx.add_log("Starting processing for #{ctx.task_id}") - - # Check retry count to avoid infinite loops - if ctx.retry_count > 3 - raise NonRetryableError, "Too many retries" - end - - # Long-running task - set callback - if will_take_long? - ctx.set_callback_after(60) # Check back in 60 seconds - return TaskInProgress.new(output: { status: 'processing' }) - end - - # Process... - result = do_work(ctx.input) - - ctx.add_log("Completed processing") - result -end -``` - ---- - -## Behavioral Algorithms - -### Algorithm 1: Main Polling Loop (`run_once`) - -```ruby -def run_once - # 1. Cleanup completed tasks (removes done futures from tracking set) - cleanup_completed_tasks - - # 2. Check capacity - current_capacity = @running_tasks.size - if current_capacity >= @max_workers - sleep(0.001) # 1ms sleep to prevent busy-waiting - return - end - - available_slots = @max_workers - current_capacity - - # 3. Adaptive backoff for empty polls - if @consecutive_empty_polls > 0 - backoff_ms = [1 * (2 ** [@consecutive_empty_polls, 10].min), @poll_interval].min - elapsed_ms = (Time.now - @last_poll_time) * 1000 - if elapsed_ms < backoff_ms - sleep((backoff_ms - elapsed_ms) / 1000.0) - return - end - end - - # 4. Batch poll for tasks - @last_poll_time = Time.now - tasks = batch_poll(available_slots) - - # 5. Submit tasks for execution - if tasks.empty? - @consecutive_empty_polls += 1 - else - @consecutive_empty_polls = 0 - tasks.each do |task| - future = @executor.post { execute_and_update(task) } - @running_tasks << future - end - end -end -``` - -### Algorithm 2: Batch Poll with Auth Backoff - -```ruby -def batch_poll(count) - # Skip if worker is paused - return [] if @worker.paused - - # Auth failure exponential backoff (capped at 60 seconds) - if @auth_failures > 0 - backoff_seconds = [2 ** @auth_failures, 60].min - elapsed = Time.now - @last_auth_failure_time - if elapsed < backoff_seconds - return [] - end - end - - # Publish PollStarted event - @event_dispatcher.publish(Events::PollStarted.new( - task_type: @worker.task_definition_name, - worker_id: @worker_id, - poll_count: @poll_count - )) - - start_time = Time.now - - begin - # HTTP batch poll - tasks = @task_client.batch_poll( - @worker.task_definition_name, - count: count, - timeout: @worker.poll_timeout, - worker_id: @worker_id, - domain: @worker.domain.presence # nil if empty string - ) - - duration_ms = (Time.now - start_time) * 1000 - @poll_count += 1 - - # Publish PollCompleted event - @event_dispatcher.publish(Events::PollCompleted.new( - task_type: @worker.task_definition_name, - duration_ms: duration_ms, - tasks_received: tasks.size - )) - - # Reset auth failures on success - @auth_failures = 0 - - tasks - rescue AuthorizationError => e - @auth_failures += 1 - @last_auth_failure_time = Time.now - duration_ms = (Time.now - start_time) * 1000 - - @event_dispatcher.publish(Events::PollFailure.new( - task_type: @worker.task_definition_name, - duration_ms: duration_ms, - cause: e - )) - - @logger.warn("Auth failure ##{@auth_failures}, backing off #{[2 ** @auth_failures, 60].min}s") - [] - rescue StandardError => e - duration_ms = (Time.now - start_time) * 1000 - - @event_dispatcher.publish(Events::PollFailure.new( - task_type: @worker.task_definition_name, - duration_ms: duration_ms, - cause: e - )) - - @logger.error("Poll failed: #{e.message}") - [] - end -end -``` - -### Algorithm 3: Task Execution - -```ruby -def execute_and_update(task) - task_result = execute_task(task) - - # Skip update for TaskInProgress (task stays in IN_PROGRESS state) - return if task_result.nil? || task_result.is_a?(TaskInProgress) - - update_task_with_retry(task_result) -end - -def execute_task(task) - # Create initial TaskResult for context - initial_result = TaskResult.new - initial_result.task_id = task.task_id - initial_result.workflow_instance_id = task.workflow_instance_id - initial_result.worker_id = @worker_id - - # Set task context (thread-local) - TaskContext.current = TaskContext.new(task, initial_result) - - start_time = Time.now - - # Publish TaskExecutionStarted - @event_dispatcher.publish(Events::TaskExecutionStarted.new( - task_type: @worker.task_definition_name, - task_id: task.task_id, - worker_id: @worker_id, - workflow_instance_id: task.workflow_instance_id - )) - - begin - # Execute worker - output = @worker.execute(task) - - duration_ms = (Time.now - start_time) * 1000 - - # Handle different return types - task_result = case output - when TaskResult - output - when TaskInProgress - result = TaskResult.in_progress - result.callback_after_seconds = output.callback_after_seconds - result.output_data = output.output - result - when Hash - result = TaskResult.complete - result.output_data = output - result - when true - TaskResult.complete - when false - TaskResult.failed('Worker returned false') - when nil - TaskResult.complete - else - result = TaskResult.complete - result.output_data = { 'result' => output } - result - end - - # Set IDs and merge context modifications - task_result.task_id = task.task_id - task_result.workflow_instance_id = task.workflow_instance_id - task_result.worker_id = @worker_id - - # Merge logs and callback_after from context - ctx = TaskContext.current - task_result.logs ||= [] - task_result.logs.concat(ctx.task_result.logs || []) - task_result.callback_after_seconds ||= ctx.callback_after_seconds - - output_size = task_result.output_data.to_json.bytesize rescue 0 - - # Publish TaskExecutionCompleted - @event_dispatcher.publish(Events::TaskExecutionCompleted.new( - task_type: @worker.task_definition_name, - task_id: task.task_id, - worker_id: @worker_id, - workflow_instance_id: task.workflow_instance_id, - duration_ms: duration_ms, - output_size_bytes: output_size - )) - - task_result - - rescue NonRetryableError => e - duration_ms = (Time.now - start_time) * 1000 - task_result = TaskResult.failed_with_terminal_error(e.message) - task_result.task_id = task.task_id - task_result.workflow_instance_id = task.workflow_instance_id - task_result.log("NonRetryableError: #{e.class}: #{e.message}") - - @event_dispatcher.publish(Events::TaskExecutionFailure.new( - task_type: @worker.task_definition_name, - task_id: task.task_id, - worker_id: @worker_id, - workflow_instance_id: task.workflow_instance_id, - duration_ms: duration_ms, - cause: e, - is_retryable: false - )) - - task_result - - rescue StandardError => e - duration_ms = (Time.now - start_time) * 1000 - task_result = TaskResult.failed(e.message) - task_result.task_id = task.task_id - task_result.workflow_instance_id = task.workflow_instance_id - task_result.log("Error: #{e.class}: #{e.message}\n#{e.backtrace.first(5).join("\n")}") - - @event_dispatcher.publish(Events::TaskExecutionFailure.new( - task_type: @worker.task_definition_name, - task_id: task.task_id, - worker_id: @worker_id, - workflow_instance_id: task.workflow_instance_id, - duration_ms: duration_ms, - cause: e, - is_retryable: true - )) - - task_result - - ensure - TaskContext.clear - end -end -``` - -### Algorithm 4: Task Update with Retry - -```ruby -RETRY_BACKOFFS = [0, 10, 20, 30].freeze # seconds - -def update_task_with_retry(task_result) - RETRY_BACKOFFS.each_with_index do |backoff, attempt| - sleep(backoff) if backoff > 0 - - start_time = Time.now - begin - @task_client.update_task(task_result) - duration_ms = (Time.now - start_time) * 1000 - - publish_task_update_completed(task_result, duration_ms) - return # Success - rescue StandardError => e - duration_ms = (Time.now - start_time) * 1000 - @logger.error("Task update failed (attempt #{attempt + 1}/#{RETRY_BACKOFFS.size}): #{e.message}") - - if attempt == RETRY_BACKOFFS.size - 1 - # All retries exhausted - CRITICAL: task result is lost - @logger.fatal("CRITICAL: Task update failed after #{RETRY_BACKOFFS.size} attempts. " \ - "Task #{task_result.task_id} result is LOST.") - publish_task_update_failure(task_result, e, duration_ms) - end - end - end -end -``` - -### Algorithm 5: Capacity Management - -The semaphore/capacity is held during BOTH execute AND update to prevent over-polling: - -```ruby -# In ThreadPoolExecutor model: -def execute_and_update(task) - # This entire method runs in a thread pool thread - # The "slot" is occupied from the moment the task is submitted - # until this method returns (after execute + update) - - task_result = execute_task(task) - return if task_result.nil? || task_result.is_a?(TaskInProgress) - - # Still occupying the slot during update retries - update_task_with_retry(task_result) - - # Slot released when this method returns and future completes -end - -# In FiberExecutor model: -def execute_and_update_fiber(task) - @semaphore.acquire # Acquire slot - - begin - task_result = execute_task(task) - return if task_result.nil? || task_result.is_a?(TaskInProgress) - - update_task_with_retry(task_result) - ensure - @semaphore.release # Release slot only after update - end -end -``` - ---- - -## Event System - -### Event Hierarchy - -```ruby -module Conductor - module Worker - module Events - # Base event with timestamp - class ConductorEvent - attr_reader :timestamp - - def initialize - @timestamp = Time.now.utc - end - end - - # Base for task runner events - class TaskRunnerEvent < ConductorEvent - attr_reader :task_type - - def initialize(task_type:) - super() - @task_type = task_type - end - end - - # Poll started - class PollStarted < TaskRunnerEvent - attr_reader :worker_id, :poll_count - - def initialize(task_type:, worker_id:, poll_count:) - super(task_type: task_type) - @worker_id = worker_id - @poll_count = poll_count - end - end - - # Poll completed successfully - class PollCompleted < TaskRunnerEvent - attr_reader :duration_ms, :tasks_received - - def initialize(task_type:, duration_ms:, tasks_received:) - super(task_type: task_type) - @duration_ms = duration_ms - @tasks_received = tasks_received - end - end - - # Poll failed - class PollFailure < TaskRunnerEvent - attr_reader :duration_ms, :cause - - def initialize(task_type:, duration_ms:, cause:) - super(task_type: task_type) - @duration_ms = duration_ms - @cause = cause - end - end - - # Task execution started - class TaskExecutionStarted < TaskRunnerEvent - attr_reader :task_id, :worker_id, :workflow_instance_id - - def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:) - super(task_type: task_type) - @task_id = task_id - @worker_id = worker_id - @workflow_instance_id = workflow_instance_id - end - end - - # Task execution completed - class TaskExecutionCompleted < TaskRunnerEvent - attr_reader :task_id, :worker_id, :workflow_instance_id, - :duration_ms, :output_size_bytes - - def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:, - duration_ms:, output_size_bytes:) - super(task_type: task_type) - @task_id = task_id - @worker_id = worker_id - @workflow_instance_id = workflow_instance_id - @duration_ms = duration_ms - @output_size_bytes = output_size_bytes - end - end - - # Task execution failed - class TaskExecutionFailure < TaskRunnerEvent - attr_reader :task_id, :worker_id, :workflow_instance_id, - :duration_ms, :cause, :is_retryable - - def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:, - duration_ms:, cause:, is_retryable: true) - super(task_type: task_type) - @task_id = task_id - @worker_id = worker_id - @workflow_instance_id = workflow_instance_id - @duration_ms = duration_ms - @cause = cause - @is_retryable = is_retryable - end - end - - # Task update failed (CRITICAL - result lost) - class TaskUpdateFailure < TaskRunnerEvent - attr_reader :task_id, :worker_id, :workflow_instance_id, - :cause, :retry_count, :task_result - - def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:, - cause:, retry_count:, task_result:) - super(task_type: task_type) - @task_id = task_id - @worker_id = worker_id - @workflow_instance_id = workflow_instance_id - @cause = cause - @retry_count = retry_count - @task_result = task_result # For recovery - end - end - end - end -end -``` - -### SyncEventDispatcher - -```ruby -module Conductor - module Worker - class SyncEventDispatcher - def initialize - @listeners = Hash.new { |h, k| h[k] = [] } - @mutex = Mutex.new - end - - # Register a listener for an event type - # @param event_type [Class] Event class to listen for - # @param listener [Proc, #call] Callable to invoke - def register(event_type, listener) - @mutex.synchronize do - @listeners[event_type] << listener unless @listeners[event_type].include?(listener) - end - end - - # Unregister a listener - def unregister(event_type, listener) - @mutex.synchronize do - @listeners[event_type].delete(listener) - end - end - - # Publish an event to all registered listeners - # @param event [ConductorEvent] Event to publish - def publish(event) - listeners = @mutex.synchronize { @listeners[event.class].dup } - - listeners.each do |listener| - begin - listener.call(event) - rescue StandardError => e - # Listener failure is isolated - never breaks the worker - warn "Event listener error for #{event.class}: #{e.message}" - end - end - end - - # Check if there are listeners for an event type - def has_listeners?(event_type) - @mutex.synchronize { @listeners[event_type].any? } - end - - # Clear all listeners (for testing) - def clear - @mutex.synchronize { @listeners.clear } - end - end - end -end -``` - -### Listener Protocol (Duck Typing) - -```ruby -module Conductor - module Worker - # Listener protocol - implement any/all of these methods - # Methods are optional - only implemented methods are called - module TaskRunnerEventsListener - # @param event [PollStarted] - def on_poll_started(event); end - - # @param event [PollCompleted] - def on_poll_completed(event); end - - # @param event [PollFailure] - def on_poll_failure(event); end - - # @param event [TaskExecutionStarted] - def on_task_execution_started(event); end - - # @param event [TaskExecutionCompleted] - def on_task_execution_completed(event); end - - # @param event [TaskExecutionFailure] - def on_task_execution_failure(event); end - - # @param event [TaskUpdateFailure] - def on_task_update_failure(event); end - end - end -end -``` - -### Listener Registration Helper - -```ruby -module Conductor - module Worker - class ListenerRegistry - # Register a listener object with the dispatcher - # Auto-detects implemented methods via respond_to? - # @param listener [Object] Object implementing TaskRunnerEventsListener methods - # @param dispatcher [SyncEventDispatcher] Event dispatcher - def self.register_task_runner_listener(listener, dispatcher) - EVENT_METHOD_MAP.each do |event_class, method_name| - if listener.respond_to?(method_name) - dispatcher.register(event_class, ->(event) { listener.send(method_name, event) }) - end - end - end - - EVENT_METHOD_MAP = { - Events::PollStarted => :on_poll_started, - Events::PollCompleted => :on_poll_completed, - Events::PollFailure => :on_poll_failure, - Events::TaskExecutionStarted => :on_task_execution_started, - Events::TaskExecutionCompleted => :on_task_execution_completed, - Events::TaskExecutionFailure => :on_task_execution_failure, - Events::TaskUpdateFailure => :on_task_update_failure - }.freeze - end - end -end -``` - -### MetricsCollector - -`MetricsCollector.create` returns a collector that emits the canonical -(harmonized) metric surface: - -```ruby -metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) -``` - -See [docs/METRICS_AND_INTERCEPTORS.md](../METRICS_AND_INTERCEPTORS.md) for the -full legacy and canonical metrics catalogs, label reference, and migration -guide. - ---- - -## Configuration System - -### Environment Variable Formats - -For a worker named `process_order` and property `poll_interval`: - -**Worker-specific (highest priority):** -```bash -# Dotted format (preferred) -conductor.worker.process_order.poll_interval=200 - -# Uppercase format -CONDUCTOR_WORKER_PROCESS_ORDER_POLL_INTERVAL=200 -``` - -**Global (applies to all workers):** -```bash -# Dotted format -conductor.worker.all.poll_interval=100 - -# Uppercase format -CONDUCTOR_WORKER_ALL_POLL_INTERVAL=100 -``` - -**Legacy (lowest priority):** -```bash -CONDUCTOR_WORKER_POLL_INTERVAL=100 -``` - -### Configuration Properties - -| Property | Type | Default | Env Var Suffix | Description | -|----------|------|---------|----------------|-------------| -| `poll_interval` | Integer | 100 | `POLL_INTERVAL` | Polling interval in milliseconds | -| `thread_count` | Integer | 1 | `THREAD_COUNT` | Max concurrent tasks (or Ractor count) | -| `domain` | String | nil | `DOMAIN` | Task domain for isolation | -| `worker_id` | String | auto | `WORKER_ID` | Unique worker identifier | -| `poll_timeout` | Integer | 100 | `POLL_TIMEOUT` | Server-side long poll timeout (ms) | -| `register_task_def` | Boolean | false | `REGISTER_TASK_DEF` | Auto-register task definition | -| `overwrite_task_def` | Boolean | true | `OVERWRITE_TASK_DEF` | Overwrite existing task defs | -| `strict_schema` | Boolean | false | `STRICT_SCHEMA` | Enforce strict JSON schema | -| `paused` | Boolean | false | `PAUSED` | Pause worker (stop polling) | -| `isolation` | Symbol | :thread | `ISOLATION` | `:thread` or `:ractor` | -| `executor` | Symbol | :thread_pool | `EXECUTOR` | `:thread_pool` or `:fiber` | - -### Auto-generated Worker ID - -```ruby -def self.generate_worker_id - hostname = Socket.gethostname rescue 'unknown' - pid = Process.pid - thread_id = Thread.current.object_id.to_s(16) - "#{hostname}-#{pid}-#{thread_id}" -end -``` - ---- - -## Task Definition Auto-Registration - -When `register_task_def: true`, the worker automatically registers its task definition on startup. - -### Registration Flow - -```ruby -def register_task_definition - return unless @worker.register_task_def - - # Build TaskDef from worker config or template - task_def = @worker.task_def_template&.dup || TaskDef.new - task_def.name = @worker.task_definition_name - - # Generate JSON schemas if possible - if @worker.execute_function.respond_to?(:parameters) - input_schema = generate_input_schema(@worker.execute_function) - output_schema = generate_output_schema(@worker.execute_function) - - register_schemas(input_schema, output_schema) if input_schema || output_schema - end - - # Register or update task definition - if @worker.overwrite_task_def - begin - @metadata_client.update_task_def(task_def) - rescue ApiError => e - if e.status == 404 - @metadata_client.register_task_def([task_def]) - else - raise - end - end - else - # Check if exists first - begin - @metadata_client.get_task_def(@worker.task_definition_name) - @logger.info("Task definition '#{@worker.task_definition_name}' already exists, skipping registration") - rescue ApiError => e - if e.status == 404 - @metadata_client.register_task_def([task_def]) - else - raise - end - end - end - - @logger.info("Registered task definition: #{@worker.task_definition_name}") -rescue StandardError => e - # Graceful degradation - worker still starts - @logger.warn("Failed to register task definition: #{e.message}") -end -``` - -### JSON Schema Generation - -```ruby -def generate_input_schema(func) - return nil unless func.respond_to?(:parameters) - - properties = {} - required = [] - - func.parameters.each do |type, name| - next if name == :task # Skip if taking full task object - - properties[name.to_s] = { 'type' => 'string' } # Default type - - case type - when :keyreq # Required keyword arg - required << name.to_s - when :key # Optional keyword arg - # Not required - end - end - - return nil if properties.empty? - - schema = { - '$schema' => 'http://json-schema.org/draft-07/schema#', - 'type' => 'object', - 'properties' => properties - } - schema['required'] = required unless required.empty? - schema['additionalProperties'] = !@worker.strict_schema - - schema -end -``` - ---- - -## File Structure - -``` -lib/conductor/ -├── worker/ -│ ├── worker.rb # Worker class -│ ├── worker_mixin.rb # WorkerMixin module for class-based workers -│ ├── worker_config.rb # Configuration resolver -│ ├── worker_registry.rb # Global worker registry -│ ├── task_handler.rb # Top-level orchestrator -│ ├── task_runner.rb # Thread-based runner -│ ├── ractor_task_runner.rb # Ractor-based runner (Phase 2) -│ ├── fiber_executor.rb # Fiber executor (Phase 3) -│ ├── task_context.rb # Execution context -│ ├── task_in_progress.rb # TaskInProgress return type -│ └── exceptions.rb # Worker-specific exceptions -├── worker/events/ -│ ├── conductor_event.rb # Base event class -│ ├── task_runner_events.rb # All 7 event classes -│ ├── sync_event_dispatcher.rb # Thread-safe event dispatcher -│ ├── listener_registry.rb # Listener registration helper -│ └── listeners.rb # Listener protocol module -├── worker/telemetry/ -│ ├── metrics_collector.rb # MetricsCollector class + NullBackend -│ └── prometheus_backend.rb # PrometheusBackend + MetricsServer -└── exceptions.rb # Add NonRetryableError -``` - ---- - -## Implementation Phases - -### Phase 1: Core Thread-based Runner (MVP) - -**Goal:** Production-ready thread-based workers with all Python SDK behavioral parity. - -**Components:** -1. `Worker` class with execute function handling -2. `WorkerMixin` module for class-based workers -3. `worker_task` DSL method and global registry -4. `WorkerConfig` with 3-tier resolution -5. `TaskRunner` with ThreadPoolExecutor -6. `TaskHandler` orchestrator -7. `TaskContext` (thread-local) -8. `TaskInProgress` return type -9. All 7 event classes -10. `SyncEventDispatcher` -11. `ListenerRegistry` -12. `MetricsCollector` with null backend - -**Algorithms:** -- Batch polling with dynamic sizing -- Adaptive backoff for empty polls -- Auth failure exponential backoff -- Task update with 4 retries -- Capacity management (semaphore during execute + update) - -**Tests:** -- Unit tests for each component -- Integration tests against local Conductor server -- All Python SDK test scenarios ported - -### Phase 2: Ractor-based Runner (Work-in-Progress) - -**Goal:** True parallelism for CPU-bound workers. - -**Status:** Partially implemented. The `RactorTaskRunner` can poll and execute -tasks inside Ractors, but the event bridge to the main thread is not yet -wired. This means metrics, interceptors, and custom event listeners receive -**no events** from Ractor workers. See -[Ractor Runner Limitations](../METRICS_AND_INTERCEPTORS.md#ractor-runner-limitations-work-in-progress----untested) -for details. - -**Components:** -1. `RactorTaskRunner` -- implemented, untested end-to-end -2. Ractor-local TaskContext storage -- implemented -3. Event aggregation via Ractor messaging -- **not yet implemented** -4. `isolation: :ractor` configuration -- implemented - -**Constraints:** -- Requires Ruby 3.1+ -- Worker must be Ractor-safe -- HTTP client created inside Ractor - -### Phase 3: Fiber Executor - -**Goal:** High-concurrency I/O-bound workers. - -**Components:** -1. `FiberExecutor` using `async` gem -2. Fiber-local TaskContext storage -3. `executor: :fiber` configuration -4. Async-compatible HTTP client - -**Constraints:** -- Requires `async` gem (optional dependency) -- All I/O must be non-blocking -- Worker must not use blocking operations - -### Phase 4: Metrics & Observability - -**Goal:** Production observability. - -**Components:** -1. `PrometheusBackend` with full metric set -2. HTTP metrics endpoint -3. Health check endpoint -4. Datadog backend (optional) - -### Phase 5: Advanced Features - -**Goal:** Feature parity with all Python SDK capabilities. - -**Components:** -1. Task definition auto-registration with JSON schemas -2. Worker auto-discovery (directory scanning) -3. Graceful shutdown with drain -4. Dynamic worker pause/resume - ---- - -## Appendix A: Complete Example - -```ruby -require 'conductor' - -# Configure -config = Conductor::Configuration.new( - server_api_url: 'http://localhost:8080/api' -) - -# Define workers - -# Class-based worker -class OrderProcessor - include Conductor::Worker::WorkerMixin - - worker_task 'process_order', - poll_interval: 200, - thread_count: 5 - - def execute(task) - order = task.input_data - ctx = Conductor::Worker::TaskContext.current - - ctx.add_log("Processing order #{order['id']}") - - # Simulate processing - result = process_order(order) - - ctx.add_log("Order processed successfully") - - { status: 'completed', total: result.total } - end - - private - - def process_order(order) - # Business logic here - OpenStruct.new(total: order['amount'] * 1.1) - end -end - -# Block-based worker -Conductor::Worker.define('send_notification', thread_count: 3) do |task| - recipient = task.input_data['recipient'] - message = task.input_data['message'] - - # Send notification - NotificationService.send(to: recipient, body: message) - - { sent: true } -end - -# Method-based worker with keyword args -module PaymentWorkers - extend Conductor::Worker::Annotatable - - worker_task 'charge_card', poll_interval: 100 - def self.charge_card(card_token:, amount:, currency: 'USD') - result = PaymentGateway.charge(card_token, amount, currency) - { transaction_id: result.id, status: result.status } - end -end - -# Custom event listener -class AuditLogger - def on_task_execution_completed(event) - puts "[AUDIT] Task #{event.task_id} completed in #{event.duration_ms}ms" - end - - def on_task_execution_failure(event) - puts "[AUDIT] Task #{event.task_id} FAILED: #{event.cause.message}" - end -end - -# Start workers -Conductor::Worker::TaskHandler.new( - configuration: config, - workers: [OrderProcessor.new], - scan_for_annotated_workers: true, - event_listeners: [AuditLogger.new] -) do |handler| - puts "Starting workers..." - handler.start - - # Handle shutdown gracefully - trap('INT') { handler.stop } - trap('TERM') { handler.stop } - - handler.join - puts "Workers stopped." -end -``` - ---- - -## Appendix B: Migration from Current Implementation - -The current `lib/conductor/worker/` has basic implementations that need to be replaced: - -| Current | New | -|---------|-----| -| `worker.rb` (WorkerModule) | `worker.rb` + `worker_mixin.rb` | -| `task_runner.rb` (simple) | `task_runner.rb` (full algorithms) | -| N/A | `task_handler.rb` | -| N/A | `worker_config.rb` | -| N/A | `task_context.rb` | -| N/A | `events/*` | -| N/A | `telemetry/*` | - -**Migration Strategy:** -1. Keep existing files during development -2. Build new implementation alongside -3. Update `lib/conductor.rb` to use new implementation -4. Remove old files after testing - ---- - -*Last Updated: February 2026* -*Status: Design Complete - Ready for Implementation* diff --git a/docs/design/WORKFLOW_DSL.md b/docs/design/WORKFLOW_DSL.md deleted file mode 100644 index d8fb8e3..0000000 --- a/docs/design/WORKFLOW_DSL.md +++ /dev/null @@ -1,561 +0,0 @@ -# Conductor Ruby SDK - Workflow DSL Design - -## Table of Contents - -1. [Overview](#overview) -2. [Design Goals](#design-goals) -3. [Architecture](#architecture) -4. [Core Components](#core-components) -5. [DSL Syntax Reference](#dsl-syntax-reference) -6. [Reference Types](#reference-types) -7. [Control Flow](#control-flow) -8. [Task Types](#task-types) -9. [Implementation Details](#implementation-details) -10. [Examples](#examples) - ---- - -## Overview - -The Conductor Ruby SDK provides a clean, Ruby-idiomatic DSL for defining workflows. The DSL uses blocks, method chaining, and Ruby's dynamic features to create a natural syntax for workflow definition. - -### Basic Example - -```ruby -workflow = Conductor.workflow :order_processing, version: 1, executor: executor do - # Access workflow inputs with wf[:param] - user = simple :get_user, user_id: wf[:user_id] - - # Reference task outputs with task[:field] - order = simple :create_order, user_email: user[:email] - - # Parallel execution - parallel do - simple :ship_order, order_id: order[:id] - simple :send_confirmation, email: user[:email] - end - - # Set workflow output - output order_id: order[:id], status: 'completed' -end - -# Register and execute -workflow.register(overwrite: true) -result = workflow.execute(input: { user_id: 123 }, wait_for_seconds: 60) -``` - ---- - -## Design Goals - -1. **Ruby-idiomatic** - Use Ruby conventions (blocks, symbols, keyword arguments) -2. **Type-safe references** - `OutputRef` and `InputRef` for compile-time safety -3. **Minimal boilerplate** - Auto-generate reference names, convert types automatically -4. **Composable** - Nested blocks for control flow (`parallel`, `decide`, `loop_over`) -5. **Full feature coverage** - Support all Conductor task types including LLM tasks - ---- - -## Architecture - -``` -┌─────────────────────────────────────────────────────────────┐ -│ User Code │ -│ Conductor.workflow :name do ... end │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ instance_eval(&block) -┌─────────────────────────────────────────────────────────────┐ -│ WorkflowBuilder │ -│ • Holds workflow metadata (name, version, description) │ -│ • Contains task method implementations │ -│ • Collects TaskRefs during DSL evaluation │ -│ • Resolves OutputRef/InputRef to expression strings │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ returns WorkflowDefinition -┌─────────────────────────────────────────────────────────────┐ -│ WorkflowDefinition │ -│ • Wraps WorkflowBuilder │ -│ • Provides .register(), .execute(), .call() methods │ -│ • Delegates to WorkflowExecutor for execution │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ to_workflow_def -┌─────────────────────────────────────────────────────────────┐ -│ Conductor::Http::Models::WorkflowDef │ -│ • Serializable workflow definition │ -│ • Contains WorkflowTask array │ -│ • Ready for API submission │ -└─────────────────────────────────────────────────────────────┘ -``` - ---- - -## Core Components - -### 1. WorkflowBuilder (`lib/conductor/workflow/dsl/workflow_builder.rb`) - -The core DSL engine containing: - -- **Workflow metadata methods**: `description`, `timeout`, `owner_email`, `restartable`, `failure_workflow`, `output` -- **Task methods**: `simple`, `http`, `wait`, `terminate`, `sub_workflow`, etc. -- **Control flow methods**: `parallel`, `decide`, `when_true`, `when_false`, `loop_over`, `loop_while` -- **LLM methods**: `llm_chat`, `llm_embed`, `generate_image`, etc. -- **Value resolution**: `resolve_value`, `resolve_hash` for OutputRef/InputRef conversion - -**Key responsibility**: During `instance_eval`, task methods create `TaskRef` objects and return them, allowing chaining like `user[:email]`. - -### 2. WorkflowDefinition (`lib/conductor/workflow/dsl/workflow_definition.rb`) - -Wrapper class returned by `Conductor.workflow`: - -```ruby -class WorkflowDefinition - def initialize(builder, executor: nil) - @builder = builder - @executor = executor - end - - def register(overwrite: false) - # Register workflow via executor - end - - def execute(input: {}, wait_for_seconds: nil, ...) - # Execute workflow via executor - end - - def call(...) - execute(...) - end - - def to_workflow_def - @builder.to_workflow_def - end -end -``` - -### 3. TaskRef (`lib/conductor/workflow/dsl/task_ref.rb`) - -Stores task metadata during DSL evaluation: - -```ruby -class TaskRef - attr_reader :ref_name, :task_name, :task_type, :input_parameters, :options - - def [](field) - OutputRef.new(ref_name, field.to_s) - end - - def to_workflow_task - # Convert to Conductor::Http::Models::WorkflowTask - end -end -``` - -### 4. OutputRef (`lib/conductor/workflow/dsl/output_ref.rb`) - -Enables `task[:field]` syntax: - -```ruby -class OutputRef - def initialize(task_ref, path) - @task_ref = task_ref - @path = path - end - - def [](field) - OutputRef.new(@task_ref, "#{@path}.#{field}") - end - - def to_s - "${#{@task_ref}.output.#{@path}}" - end -end -``` - -### 5. InputRef (`lib/conductor/workflow/dsl/input_ref.rb`) - -Enables `wf[:param]` syntax: - -```ruby -class InputRef - def [](field) - InputFieldRef.new("workflow.input.#{field}") - end - - def var(name) - InputFieldRef.new("workflow.variables.#{name}") - end -end - -class InputFieldRef - def [](field) - InputFieldRef.new("#{@path}.#{field}") - end - - def to_s - "${#{@path}}" - end -end -``` - -### 6. Control Flow Builders - -**ParallelBuilder** - Collects branches for `parallel do...end`: - -```ruby -class ParallelBuilder - def initialize(parent_builder) - @parent = parent_builder - @branches = [[]] - end - - def method_missing(name, *args, **kwargs, &block) - # Delegate to parent builder, collect tasks into current branch - end - - def finalize - @branches # Array of task arrays - end -end -``` - -**SwitchBuilder** - Handles `decide expr do...end`: - -```ruby -class SwitchBuilder - def on(value, &block) - # Add case branch - end - - def otherwise(&block) - # Add default branch - end -end -``` - ---- - -## DSL Syntax Reference - -### Workflow Definition - -```ruby -Conductor.workflow :name, version: 1, executor: executor do - description 'Workflow description' - timeout 3600 # Timeout in seconds - owner_email 'owner@example.com' - restartable true - failure_workflow 'failure_handler' - - # ... tasks ... - - output key: value # Workflow output parameters -end -``` - -### Input/Output References - -```ruby -# Workflow inputs -wf[:user_id] # "${workflow.input.user_id}" -wf[:data][:items] # "${workflow.input.data.items}" -wf.var(:counter) # "${workflow.variables.counter}" - -# Task outputs -task[:result] # "${task_ref.output.result}" -task[:data][:nested][:field] # "${task_ref.output.data.nested.field}" -``` - ---- - -## Control Flow - -### Parallel Execution - -```ruby -parallel do - simple :task_a - simple :task_b - simple :task_c -end -``` - -Generates FORK_JOIN → tasks → JOIN structure. - -### Conditional Branching - -```ruby -decide user[:tier] do - on 'gold' do - simple :apply_gold_discount - end - on 'silver' do - simple :apply_silver_discount - end - otherwise do - simple :no_discount - end -end -``` - -### Conditional Shortcuts - -```ruby -when_true order[:is_premium] do - simple :apply_discount -end - -when_false order[:validated] do - terminate :failed, 'Validation failed' -end -``` - -### Loops - -```ruby -# Loop N times -loop_times 3 do - simple :process_batch -end - -# Loop with condition -loop_while '$.has_more == true' do - simple :fetch_page -end - -# Loop over items -loop_over users[:list] do - simple :process_user, user: iteration[:item] -end -``` - ---- - -## Task Types - -### Basic Tasks - -| Method | Type | Description | -|--------|------|-------------| -| `simple :name, **inputs` | SIMPLE | Worker task execution | -| `http :name, url:, method:, body:, headers:` | HTTP | HTTP request | -| `javascript :name, script:, **bindings` | INLINE | Inline JavaScript | -| `jq :name, query:, **inputs` | JSON_JQ_TRANSFORM | JQ transformation | -| `set var: value` | SET_VARIABLE | Set workflow variables | -| `human :name, assignee:, display_name:` | HUMAN | Human/manual task | - -### Wait and Events - -| Method | Type | Description | -|--------|------|-------------| -| `wait seconds` | WAIT | Wait for duration | -| `wait until_time: 'ISO8601'` | WAIT | Wait until time | -| `event :name, sink:, **payload` | EVENT | Publish event | -| `wait_for_webhook :name, matches: {}` | WAIT_FOR_WEBHOOK | Wait for callback | - -### Workflow Control - -| Method | Type | Description | -|--------|------|-------------| -| `terminate :status, 'reason'` | TERMINATE | End workflow | -| `sub_workflow :name, workflow:, version:` | SUB_WORKFLOW | Call workflow | -| `start_workflow :name, workflow:, **inputs` | START_WORKFLOW | Fire-and-forget | -| `inline_workflow :name do...end` | SUB_WORKFLOW | Inline sub-workflow | - -### Dynamic Tasks - -| Method | Type | Description | -|--------|------|-------------| -| `dynamic :name, dynamic_task_param:` | DYNAMIC | Runtime task name | -| `dynamic_fork :name, tasks_param:, tasks_input_param:` | FORK_JOIN_DYNAMIC | Dynamic parallel | -| `http_poll :name, url:, termination_condition:` | HTTP_POLL | Poll until condition | - -### LLM/AI Tasks - -| Method | Type | Description | -|--------|------|-------------| -| `llm_chat :name, provider:, model:, messages:` | LLM_CHAT_COMPLETE | Chat completion | -| `llm_complete :name, provider:, model:, prompt:` | LLM_TEXT_COMPLETE | Text completion | -| `llm_embed :name, provider:, model:, text:` | LLM_GENERATE_EMBEDDINGS | Generate embeddings | -| `llm_store_embeddings :name, vector_db:, index:, embeddings:` | LLM_STORE_EMBEDDINGS | Store in vector DB | -| `llm_search_embeddings :name, vector_db:, index:, embeddings:` | LLM_SEARCH_EMBEDDINGS | Search vector DB | -| `generate_image :name, provider:, model:, prompt:` | GENERATE_IMAGE | Image generation | -| `generate_audio :name, provider:, model:, text:, voice:` | GENERATE_AUDIO | Text-to-speech | -| `list_mcp_tools :name, mcp_server:` | LIST_MCP_TOOLS | List MCP tools | -| `call_mcp_tool :name, mcp_server:, method:, arguments:` | CALL_MCP_TOOL | Call MCP tool | -| `get_document :name, url:, media_type:` | GET_DOCUMENT | Retrieve document | - ---- - -## Implementation Details - -### Reference Resolution - -The DSL automatically converts references to Conductor expression strings: - -```ruby -def resolve_value(value) - case value - when OutputRef, InputRef - value.to_s # "${task_ref.output.field}" - when Hash - resolve_hash(value) # Recursively resolve - when Array - value.map { |v| resolve_value(v) } - else - value # Literals pass through - end -end -``` - -### Task Reference Name Generation - -```ruby -def generate_ref_name(task_name) - base = "#{task_name}_ref" - @ref_counter[base] += 1 - @ref_counter[base] == 1 ? base : "#{base}_#{@ref_counter[base]}" -end -``` - -This ensures unique reference names: -- First `simple :get_user` → `get_user_ref` -- Second `simple :get_user` → `get_user_ref_2` - -### Workflow Task Conversion - -`TaskRef#to_workflow_task` converts DSL representation to `WorkflowTask` model: - -```ruby -def to_workflow_task - wf_task = Conductor::Http::Models::WorkflowTask.new( - name: @task_name, - task_reference_name: @ref_name, - type: @task_type, - input_parameters: @input_parameters - ) - - # Handle special task types (FORK_JOIN, SWITCH, DO_WHILE, etc.) - case @task_type - when TaskType::FORK_JOIN - wf_task.fork_tasks = convert_branches(@options[:fork_branches]) - when TaskType::SWITCH - wf_task.expression = @options[:expression] - wf_task.decision_cases = @options[:decision_cases] - wf_task.default_case = @options[:default_case] - # ... etc - end - - wf_task -end -``` - ---- - -## Examples - -### E-commerce Order Processing - -```ruby -workflow = Conductor.workflow :order_processing, version: 1, executor: executor do - description 'Process customer orders' - timeout 3600 - - # Validate order - validation = simple :validate_order, - order_id: wf[:order_id], - customer_id: wf[:customer_id] - - # Check inventory - inventory = simple :check_inventory, - items: validation[:items] - - # Conditional based on stock - decide inventory[:in_stock] do - on 'true' do - # Process payment - payment = simple :process_payment, - amount: validation[:total], - customer_id: wf[:customer_id] - - # Parallel fulfillment - parallel do - simple :ship_order, order_id: wf[:order_id] - simple :send_confirmation, email: validation[:customer_email] - simple :update_analytics, order_data: validation[:result] - end - end - - otherwise do - simple :notify_backorder, items: inventory[:missing_items] - terminate :failed, 'Items out of stock' - end - end - - output order_id: wf[:order_id], status: 'completed' -end -``` - -### AI Document Processing - -```ruby -workflow = Conductor.workflow :document_analysis, version: 1, executor: executor do - # Fetch document - doc = get_document :fetch_doc, - url: wf[:document_url], - media_type: wf[:media_type] - - # Generate embeddings - embeddings = llm_embed :embed_doc, - provider: 'openai', - model: 'text-embedding-3-small', - text: doc[:content] - - # Store in vector database - llm_store_embeddings :store_vectors, - vector_db: 'pinecone', - index: 'documents', - embeddings: embeddings[:embeddings], - id: wf[:doc_id], - metadata: { source: wf[:document_url] } - - # Analyze with LLM - analysis = llm_chat :analyze, - provider: 'openai', - model: 'gpt-4', - messages: [ - { role: :system, message: 'Analyze the following document and extract key insights.' }, - { role: :user, message: doc[:content] } - ], - temperature: 0.3 - - output analysis: analysis[:content], doc_id: wf[:doc_id] -end -``` - ---- - -## Testing - -DSL tests are in `spec/conductor/workflow/dsl/workflow_builder_spec.rb`: - -```bash -bundle exec rspec spec/conductor/workflow/dsl/ --format documentation -``` - -Current coverage: 52 test cases covering: -- All basic task methods -- All LLM task methods -- Control flow (parallel, decide, loops) -- Reference resolution (OutputRef, InputRef) -- WorkflowDef conversion - ---- - -## Related Documents - -- [DESIGN.md](../../DESIGN.md) - High-level architecture -- [WORKER_DESIGN.md](WORKER_DESIGN.md) - Worker infrastructure -- [README.md](../../README.md) - User documentation diff --git a/examples/agents/README.md b/examples/agents/README.md deleted file mode 100644 index 1b308a4..0000000 --- a/examples/agents/README.md +++ /dev/null @@ -1,88 +0,0 @@ -# Agent examples - -These are standalone Ruby ports of the [Python agent examples](https://github.com/conductor-oss/python-sdk/tree/c99e2cf9871c21f7a64d823126ee1b77989b00ad/examples/agents). -Each file contains its own tools, agents, and prompts. Run or copy that file; no test harness is required. - -```bash -bundle install -export CONDUCTOR_SERVER_URL=http://localhost:8080/api -export CONDUCTOR_AGENT_LLM_MODEL=openai/gpt-4o-mini -bundle exec ruby -Ilib examples/agents/01_basic_agent.rb -``` - -Configure the selected LLM integration on the server. The SDK uses the usual -`CONDUCTOR_AUTH_KEY` / `CONDUCTOR_AUTH_SECRET` settings when authentication is required. - -| File | Demonstrates | -|---|---| -| [01_basic_agent.rb](01_basic_agent.rb) | Basic question and answer | -| [02a_simple_tools.rb](02a_simple_tools.rb) | Local Ruby tools | -| [02c_tool_retry_config.rb](02c_tool_retry_config.rb) | Task retry policies | -| [04_http_and_mcp_tools.rb](04_http_and_mcp_tools.rb) | HTTP and discovered MCP tools | -| [05_handoffs.rb](05_handoffs.rb) | Agent handoffs | -| [06_sequential_pipeline.rb](06_sequential_pipeline.rb) | Sequential team | -| [07_parallel_agents.rb](07_parallel_agents.rb) | Parallel team | -| [09_human_in_the_loop.rb](09_human_in_the_loop.rb) | Interactive approval and feedback | -| [09c_hitl_streaming.rb](09c_hitl_streaming.rb) | Streaming events and interactive approval | -| [103_plan_and_compile.rb](103_plan_and_compile.rb) | Planner, compiled workflow, and recovery agent | -| [10_guardrails.rb](10_guardrails.rb) | Ruby guardrail functions | -| [13_hierarchical_agents.rb](13_hierarchical_agents.rb) | Nested agent teams | -| [17_swarm_orchestration.rb](17_swarm_orchestration.rb) | Swarm routing | -| [21_regex_guardrails.rb](21_regex_guardrails.rb) | Regex redaction and validation | -| [22_llm_guardrails.rb](22_llm_guardrails.rb) | LLM validation and expected rejection | -| [64_swarm_with_tools.rb](64_swarm_with_tools.rb) | Swarm members with worker tools | -| [66_handoff_to_parallel.rb](66_handoff_to_parallel.rb) | Handoff into a parallel team | -| [33_external_workers.rb](33_external_workers.rb) | External services and one local formatting tool | -| [16e_credentials_http_tool.rb](16e_credentials_http_tool.rb) | Server-side HTTP credential substitution | - -For `04`, run `mcp-testkit==1.0.4 --transport http --auth ...` on port 3001 and -configure `MCP_TEST_API_KEY` and `HTTP_TEST_API_KEY` in the server secret store. -For `16e`, configure `GITHUB_TOKEN` in the server secret store. `GITHUB_REPOS_URL` -can override the GitHub URL when using a local test dependency. -For `33`, start `bundle exec ruby -Ilib examples/agents/external_workers.rb` in another -terminal. Its three worker services run independently of the example runtime. -The two approval examples read their response fields from stdin. Example `22` deliberately -uses strict guardrails: its recorded outcome is a verified content-safety rejection. - -## Playback integration suite - -[The integration wrapper](../../spec/integration/agents/examples_spec.rb) requires these -files through the filename-only [catalog](catalog.rb) and calls their `run` methods. -It never redefines agents, tools, or prompts. It supplies approval input, checks final states, -and inspects persisted tasks for mock LLM execution, approvals, guards, compiled factorial -tasks, and external workers. - -[CI](../../.github/workflows/agents-playback.yml) builds Conductor from -`feature/llm_mock_impl`, uses that branch's unchanged shared recordings, -and invokes its `check-playback` composite action from the same branch. Local verification -used branch head `acb7d27750e5dacc6a6334ed0bcabe3ade030533`. The test server uses a fresh SQLite database. -HTTP/MCP services are real local dependencies; the HTTP fixture returns the recorded GitHub -response and validates the substituted dummy bearer credential. No provider keys are needed. - -To reproduce with Java 21, Ruby 3.3, Python, and `mcp-testkit==1.0.4` installed: - -```bash -# In a separate Conductor checkout at the feature branch revision: -./gradlew --no-daemon :conductor-server:bootJar -x test - -# Back in ruby-sdk; use a new work directory for every run: -export CONDUCTOR_PLAYBACK_WORK_DIR="$PWD/tmp/agent-playback" -export CONDUCTOR_SERVER_URL=http://localhost:18080/api -export CONDUCTOR_AGENT_LLM_MODEL=mock/mockLLM -export CONDUCTOR_AGENTS_PLAYBACK=true -export GITHUB_REPOS_URL='http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated' -bash .github/scripts/start-agent-playback.sh /path/to/conductor -bundle exec ruby -Ilib examples/agents/external_workers.rb > "$CONDUCTOR_PLAYBACK_WORK_DIR/workers.log" 2>&1 & -echo $! > "$CONDUCTOR_PLAYBACK_WORK_DIR/workers.pid" -bundle exec rspec spec/integration/agents/ --format documentation -sh /path/to/conductor/.github/actions/check-playback/check-playback.sh "$CONDUCTOR_SERVER_URL" - -# Stop only the services started for this run: -for file in "$CONDUCTOR_PLAYBACK_WORK_DIR"/*.pid; do - kill "$(cat "$file")" 2>/dev/null || true -done -``` - -The shared check requires all 93 recordings to be consumed and zero unmatched LLM requests. -The GitHub action has been configured; its verifier was also run locally after all 19 tests passed. -WireMock replay remains a separate transport test; it is not a substitute for this suite. diff --git a/harness/README.md b/harness/README.md deleted file mode 100644 index b50a56f..0000000 --- a/harness/README.md +++ /dev/null @@ -1,50 +0,0 @@ -# Ruby SDK Docker Harness - -A long-running worker harness built from the root `Dockerfile`. - -## Worker Harness - -A self-feeding worker that runs indefinitely. On startup it registers five simulated tasks (`ruby_worker_0` through `ruby_worker_4`) and the `ruby_simulated_tasks_workflow`, then runs two background services: - -- **WorkflowGovernor** -- starts a configurable number of `ruby_simulated_tasks_workflow` instances per second (default 2), indefinitely. -- **SimulatedTaskWorkers** -- five task handlers, each with a codename and a default sleep duration. Each worker supports configurable delay types, failure simulation, and output generation via task input parameters. The workflow chains them in sequence: quickpulse (1s) → whisperlink (2s) → shadowfetch (3s) → ironforge (4s) → deepcrawl (5s). - -```bash -docker build --target harness -t ruby-sdk-harness . - -docker run -d \ - -e CONDUCTOR_SERVER_URL=https://your-cluster.example.com/api \ - -e CONDUCTOR_AUTH_KEY=$CONDUCTOR_AUTH_KEY \ - -e CONDUCTOR_AUTH_SECRET=$CONDUCTOR_AUTH_SECRET \ - -e HARNESS_WORKFLOWS_PER_SEC=4 \ - ruby-sdk-harness -``` - -You can also run the harness locally without Docker: - -```bash -export CONDUCTOR_SERVER_URL=https://your-cluster.example.com/api -export CONDUCTOR_AUTH_KEY=$CONDUCTOR_AUTH_KEY -export CONDUCTOR_AUTH_SECRET=$CONDUCTOR_AUTH_SECRET - -ruby harness/main.rb -``` - -Override defaults with environment variables as needed: - -```bash -HARNESS_WORKFLOWS_PER_SEC=4 HARNESS_BATCH_SIZE=10 ruby harness/main.rb -``` - -All resource names use a `ruby_` prefix so multiple SDK harnesses (Python, Java, Go, C#, etc.) can coexist on the same cluster. - -### Environment Variables - -| Variable | Required | Default | Description | -|---|---|---|---| -| `CONDUCTOR_SERVER_URL` | yes | -- | Conductor API base URL | -| `CONDUCTOR_AUTH_KEY` | no | -- | Orkes auth key | -| `CONDUCTOR_AUTH_SECRET` | no | -- | Orkes auth secret | -| `HARNESS_WORKFLOWS_PER_SEC` | no | 2 | Workflows to start per second | -| `HARNESS_BATCH_SIZE` | no | 20 | Number of tasks each worker polls per batch | -| `HARNESS_POLL_INTERVAL_MS` | no | 100 | Milliseconds between poll cycles | diff --git a/harness/manifests/README.md b/harness/manifests/README.md deleted file mode 100644 index 3b3d943..0000000 --- a/harness/manifests/README.md +++ /dev/null @@ -1,132 +0,0 @@ -# Kubernetes Manifests - -This directory contains Kubernetes manifests for deploying the Ruby SDK harness worker to the certification clusters. - -## Prerequisites - -**Set your namespace environment variable:** -```bash -export NS=your-namespace-here -``` - -All kubectl commands below use `-n $NS` to specify the namespace. The manifests intentionally do not include hardcoded namespaces. - -**Note:** The harness worker images are published as public packages on GHCR and do not require authentication to pull. No image pull secrets are needed. - -## Files - -| File | Description | -|---|---| -| `deployment.yaml` | Deployment (single file, works on all clusters) | -| `configmap-aws.yaml` | Conductor URL + auth key for certification-aws | -| `configmap-azure.yaml` | Conductor URL + auth key for certification-az | -| `configmap-gcp.yaml` | Conductor URL + auth key for certification-gcp | -| `secret-conductor.yaml` | Conductor auth secret (placeholder template) | - -## Quick Start - -### 1. Create the Conductor Auth Secret - -The `CONDUCTOR_AUTH_SECRET` must be created as a Kubernetes secret before deploying. - -```bash -kubectl create secret generic conductor-credentials \ - --from-literal=auth-secret=YOUR_AUTH_SECRET \ - -n $NS -``` - -If the `conductor-credentials` secret already exists in the namespace (e.g. from the e2e-testrunner-worker), it can be reused as-is. - -See `secret-conductor.yaml` for more details. - -### 2. Apply the ConfigMap for Your Cluster - -```bash -# AWS -kubectl apply -f manifests/configmap-aws.yaml -n $NS - -# Azure -kubectl apply -f manifests/configmap-azure.yaml -n $NS - -# GCP -kubectl apply -f manifests/configmap-gcp.yaml -n $NS -``` - -### 3. Deploy - -```bash -kubectl apply -f manifests/deployment.yaml -n $NS -``` - -### 4. Verify - -```bash -# Check pod status -kubectl get pods -n $NS -l app=ruby-sdk-harness-worker - -# Watch logs -kubectl logs -n $NS -l app=ruby-sdk-harness-worker -f -``` - -## Building and Pushing the Image - -From the repository root: - -```bash -# Build the harness target and push to GHCR -docker buildx build \ - --platform linux/amd64,linux/arm64 \ - --target harness \ - -t ghcr.io/conductor-oss/ruby-sdk/harness-worker:latest \ - --push . -``` - -After pushing a new image with the same tag, restart the deployment to pull it: - -```bash -kubectl rollout restart deployment/ruby-sdk-harness-worker -n $NS -kubectl rollout status deployment/ruby-sdk-harness-worker -n $NS -``` - -## Tuning - -The harness worker accepts these optional environment variables (set in `deployment.yaml`): - -| Variable | Default | Description | -|---|---|---| -| `HARNESS_WORKFLOWS_PER_SEC` | 2 | Workflows to start per second | -| `HARNESS_BATCH_SIZE` | 20 | Tasks each worker polls per batch | -| `HARNESS_POLL_INTERVAL_MS` | 100 | Milliseconds between poll cycles | - -Edit `deployment.yaml` to change these, then re-apply: - -```bash -kubectl apply -f manifests/deployment.yaml -n $NS -``` - -## Troubleshooting - -### Pod not starting - -```bash -kubectl describe pod -n $NS -l app=ruby-sdk-harness-worker -kubectl logs -n $NS -l app=ruby-sdk-harness-worker --tail=100 -``` - -### Secret not found - -```bash -kubectl get secret conductor-credentials -n $NS -``` - -## Resource Limits - -Default resource allocation: -- **Memory**: 256Mi (request) / 512Mi (limit) -- **CPU**: 100m (request) / 500m (limit) - -Adjust in `deployment.yaml` based on workload. Higher `HARNESS_WORKFLOWS_PER_SEC` values may need more CPU/memory. - -## Service - -The harness worker does **not** need a Service or Ingress. It connects to Conductor via outbound HTTP polling. All communication is outbound. From 6fd05d8a89693f09d4d109ec909202e23494b32c Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 09:34:13 -0700 Subject: [PATCH 11/20] Restore documentation files inherited from main --- DESIGN.md | 304 ++++ docs/METRICS_AND_INTERCEPTORS.md | 827 +++++++++++ docs/design/EVENT_INTERCEPTOR_SYSTEM.md | 907 ++++++++++++ docs/design/WORKER_DESIGN.md | 1781 +++++++++++++++++++++++ docs/design/WORKFLOW_DSL.md | 561 +++++++ harness/README.md | 50 + harness/manifests/README.md | 132 ++ 7 files changed, 4562 insertions(+) create mode 100644 DESIGN.md create mode 100644 docs/METRICS_AND_INTERCEPTORS.md create mode 100644 docs/design/EVENT_INTERCEPTOR_SYSTEM.md create mode 100644 docs/design/WORKER_DESIGN.md create mode 100644 docs/design/WORKFLOW_DSL.md create mode 100644 harness/README.md create mode 100644 harness/manifests/README.md diff --git a/DESIGN.md b/DESIGN.md new file mode 100644 index 0000000..86c4ca1 --- /dev/null +++ b/DESIGN.md @@ -0,0 +1,304 @@ +# Conductor Ruby SDK - Design Document + +## Overview + +This SDK is a Ruby port of the [Conductor Python SDK](https://github.com/conductor-sdk/conductor-python), designed to provide a Ruby-native interface to Conductor OSS while maintaining architectural parity with the Python implementation. + +## Design Principles + +1. **Follow Python SDK architecture** - Maintain the same layers, patterns, and behaviors +2. **Ruby-idiomatic API** - Use Ruby conventions (snake_case, blocks, symbols) while keeping the same capabilities +3. **Feature parity** - Support all features from the Python SDK +4. **Clean DSL** - Block-based workflow definition with method chaining for output references + +## Architecture Layers + +``` +┌─────────────────────────────────────────────────────────────┐ +│ User Code │ +│ Conductor.workflow :name do ... end │ +│ class MyWorker; include WorkerModule; end │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Workflow DSL │ +│ • Conductor.workflow - Entry point for workflow definition │ +│ • WorkflowBuilder - Core DSL engine with task methods │ +│ • WorkflowDefinition - Wrapper with .register/.execute │ +│ • OutputRef/InputRef - Reference expression builders │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Worker Framework │ +│ • WorkerModule - Mixin for class-based workers │ +│ • Worker.define - Block-based worker definition │ +│ • TaskRunner - Polling and execution │ +│ • TaskHandler - Worker orchestration │ +│ • Event system - Lifecycle hooks │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ High-Level Clients (Facades) │ +│ • WorkflowClient, TaskClient, MetadataClient │ +│ • SchedulerClient, PromptClient, SecretClient │ +│ • IntegrationClient, AuthorizationClient, SchemaClient │ +│ • WorkflowExecutor - Synchronous workflow execution │ +│ • OrkesClients - Factory for all clients │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Resource APIs (17 classes) │ +│ • WorkflowResourceApi, TaskResourceApi │ +│ • MetadataResourceApi, SchedulerResourceApi │ +│ • EventResourceApi, WorkflowBulkResourceApi │ +│ • PromptResourceApi, SecretResourceApi, etc. │ +│ • Direct mapping to Conductor REST API │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ HTTP Transport │ +│ • ApiClient - Auth, serialization, dispatch │ +│ • RestClient - Faraday with HTTP/2, connection pooling │ +│ • BaseModel - SWAGGER_TYPES pattern for serialization │ +│ • 50+ Model classes for request/response types │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Configuration │ +│ • Server URL, auth settings │ +│ • Token management (class-level cache, TTL refresh) │ +│ • Environment variable support │ +└─────────────────────────────────────────────────────────────┘ +``` + +## Workflow DSL Design + +The SDK provides a clean, Ruby-idiomatic DSL for building workflows: + +### Entry Point + +```ruby +workflow = Conductor.workflow :order_processing, version: 1, executor: executor do + # Workflow definition using DSL methods + user = simple :get_user, user_id: wf[:user_id] + simple :send_email, email: user[:email] + output result: user[:name] +end + +workflow.register(overwrite: true) +result = workflow.execute(input: { user_id: 123 }) +``` + +### DSL Components + +| Component | Purpose | +|-----------|---------| +| `WorkflowBuilder` | Core DSL engine with all task methods | +| `WorkflowDefinition` | Wrapper returned by `Conductor.workflow` | +| `TaskRef` | Stores task metadata, converts to WorkflowTask | +| `OutputRef` | Enables `task[:field]` syntax | +| `InputRef` | Enables `wf[:param]` syntax | +| `ParallelBuilder` | Handles `parallel do...end` blocks | +| `SwitchBuilder` | Handles `decide do...end` blocks | + +### Reference Resolution + +The DSL automatically converts references to Conductor expression strings: + +```ruby +task[:field] # → "${task_ref.output.field}" +task[:nested][:path] # → "${task_ref.output.nested.path}" +wf[:param] # → "${workflow.input.param}" +wf.var(:counter) # → "${workflow.variables.counter}" +``` + +### Task Types (25+) + +**Basic Tasks:** +- `simple` - Worker task execution +- `http` - HTTP request +- `javascript` - Inline JavaScript execution +- `jq` - JSON JQ transformation +- `set` - Set workflow variables +- `human` - Human/manual task + +**Control Flow:** +- `parallel do...end` - Fork/Join execution +- `decide expr do...end` - Switch/case branching +- `when_true/when_false` - Conditional shortcuts +- `loop_over`, `loop_while`, `loop_times` - Iteration +- `sub_workflow` - Call another workflow +- `inline_workflow` - Define sub-workflow inline + +**System Tasks:** +- `wait` - Wait for duration or time +- `terminate` - End workflow +- `event` - Publish event +- `wait_for_webhook` - Wait for external callback +- `http_poll` - Poll HTTP endpoint +- `dynamic`, `dynamic_fork` - Runtime task resolution +- `kafka_publish` - Publish to Kafka +- `start_workflow` - Fire-and-forget workflow start + +**LLM/AI Tasks:** +- `llm_chat` - Chat completion +- `llm_complete` - Text completion +- `llm_embed` - Generate embeddings +- `llm_index`, `llm_search` - Vector operations +- `llm_store_embeddings`, `llm_search_embeddings` - Vector DB operations +- `generate_image`, `generate_audio` - Media generation +- `list_mcp_tools`, `call_mcp_tool` - MCP integration +- `get_document` - Document retrieval + +## Worker Framework Design + +See [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) for detailed design. + +### Worker Definition Patterns + +```ruby +# Pattern 1: Class-based (recommended for complex workers) +class OrderProcessor + include Conductor::Worker::WorkerModule + worker_task 'process_order', poll_interval: 200, thread_count: 5 + + def execute(task) + order_id = get_input(task, 'order_id') + # Process... + { processed: true, order_id: order_id } + end +end + +# Pattern 2: Block-based (quick workers) +Conductor::Worker.define('send_email', poll_interval: 100) do |task| + EmailService.deliver(task.input_data['to'], task.input_data['subject']) + { sent: true } +end +``` + +### Concurrency Model + +**Python SDK:** Uses `multiprocessing.Process` (one OS process per worker) to avoid GIL. + +**Ruby SDK:** Uses `Thread` + `concurrent-ruby`'s `ThreadPoolExecutor`. Ruby's GVL releases during I/O operations, making threads sufficient for typical worker workloads. + +## Auth Flow (Matches Python SDK) + +| Scenario | Behavior | +|----------|----------| +| No auth configured | Skip auth. No `X-Authorization` header. | +| Initial token fetch | POST `/api/token` → cache token (class-level). | +| `/token` returns 404 | **Disable auth.** Set `auth_configured? = false`. Log "OSS mode". | +| Token TTL expired | Proactively refresh before next request. | +| Server 401/403 with `EXPIRED_TOKEN` | Force-refresh, retry request once. | +| Server 401/403 with `INVALID_TOKEN` | Force-refresh, retry request once. | +| Token refresh fails | Exponential backoff: 2^n seconds, max 5 attempts. | + +## File Structure + +``` +lib/conductor/ +├── version.rb # VERSION constant +├── configuration.rb # Configuration class +├── exceptions.rb # Exception hierarchy +├── client/ # High-level client facades (9) +│ ├── workflow_client.rb +│ ├── task_client.rb +│ ├── metadata_client.rb +│ └── ... +├── http/ +│ ├── api/ # Resource API classes (17) +│ │ ├── workflow_resource_api.rb +│ │ ├── task_resource_api.rb +│ │ └── ... +│ ├── models/ # HTTP models (50+) +│ │ ├── workflow_def.rb +│ │ ├── workflow_task.rb +│ │ └── ... +│ ├── api_client.rb # Auth + serialization +│ └── rest_client.rb # Faraday HTTP client +├── orkes/ # Orkes Cloud specific +│ ├── orkes_clients.rb # Main factory +│ └── models/ +├── worker/ # Worker framework +│ ├── task_runner.rb # Polling loop +│ ├── task_handler.rb # Worker management +│ ├── worker.rb # Worker module +│ └── events/ # Event system +└── workflow/ + ├── dsl/ # Workflow DSL + │ ├── workflow_builder.rb # Core DSL engine + │ ├── workflow_definition.rb # Wrapper class + │ ├── task_ref.rb # Task reference + │ ├── output_ref.rb # Output reference + │ ├── input_ref.rb # Input reference + │ ├── parallel_builder.rb # parallel do...end + │ └── switch_builder.rb # decide do...end + ├── llm/ # LLM helper classes + │ ├── chat_message.rb + │ ├── tool_call.rb + │ ├── tool_spec.rb + │ └── embedding_model.rb + ├── task_type.rb # Task type constants + ├── timeout_policy.rb + └── workflow_executor.rb +lib/conductor/agents.rb # require 'conductor/agents' +lib/conductor/agents/ # Agents (see docs/agents/README.md) +├── agent.rb, tool_def.rb, tools.rb # definition layer + `tool def` DSL +├── guardrail.rb, termination.rb, handoff.rb, callback_handler.rb, memory.rb, prompt_template.rb +├── config_serializer.rb # -> agentConfig, identical to the Python SDK +└── runtime/ # AgentRuntime, SseClient, Execution, ApprovalRequest, ToolRegistry, Dispatch, Secrets +``` + +## Agents + +`Conductor::Agents` ports the Python SDK's `conductor.ai.agents` package. An `Agent` tree is +serialized by `ConfigSerializer` to the same `agentConfig` JSON Python sends; the server compiles +it into a workflow and runs the LLM loop. `AgentRuntime` starts the execution +(`POST /api/agent/start`), registers a Conductor worker for every task the server lists in +`requiredWorkers` (the user's `tool def` tools plus `_termination`, custom guardrails and +callbacks), and follows the run over SSE (`GET /api/agent/stream/{id}`, polling fallback). +Secrets travel as `TaskDef.runtimeMetadata` names and come back as `Task.runtimeMetadata` +values, read inside tools with `secret('NAME')`. The transport for `/api/agent/*` is +`AgentResourceApi` / `AgentClient` like every other resource. Decisions and the verified wire +contract are in `docs/design/AGENTS_IMPLEMENTATION_PLAN.md`. + +## Dependencies + +```ruby +# Core HTTP +gem 'faraday', '~> 2.0' +gem 'faraday-net_http_persistent', '~> 2.0' # Connection pooling +gem 'faraday-retry', '~> 2.0' # Automatic retries + +# Concurrency +gem 'concurrent-ruby', '~> 1.2' # ThreadPoolExecutor + +# JSON +gem 'json', '>= 2.0' +``` + +## Testing Strategy + +1. **Unit tests** - Mock HTTP, test logic in isolation (~400 tests) +2. **Integration tests** - Against Conductor server (~137 tests) +3. **Example-driven** - All examples in `examples/` directory + +```bash +# Run all tests +bundle exec rspec + +# Run DSL tests +bundle exec rspec spec/conductor/workflow/dsl/ + +# Integration tests (requires server) +CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ +``` + +## Related Documents + +- [AGENTS.md](AGENTS.md) - Guide for AI coding agents +- [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) - Worker infrastructure design +- [docs/design/WORKFLOW_DSL.md](docs/design/WORKFLOW_DSL.md) - Workflow DSL design +- [README.md](README.md) - User documentation +- [CONTRIBUTING.md](CONTRIBUTING.md) - Development guidelines diff --git a/docs/METRICS_AND_INTERCEPTORS.md b/docs/METRICS_AND_INTERCEPTORS.md new file mode 100644 index 0000000..fd2f289 --- /dev/null +++ b/docs/METRICS_AND_INTERCEPTORS.md @@ -0,0 +1,827 @@ +# Metrics and Interceptors Guide + +The Conductor Ruby SDK can expose Prometheus metrics for worker polling, task +execution, task result updates, payload sizes, workflow starts, and HTTP API +client latency. It also provides an event-driven interceptor system for custom +logging, error tracking, and observability. + +This document covers the Ruby SDK metrics emitted by `MetricsCollector`. It +does not cover Conductor server metrics or metrics emitted by other SDKs. + +## Table of Contents + +- [Quick Start](#quick-start) +- [Metrics Catalog](#metrics-catalog) +- [Metrics Not Applicable to Ruby](#metrics-not-applicable-to-ruby) +- [Labels](#labels) +- [Prometheus Integration](#prometheus-integration) +- [Custom Metrics Backends](#custom-metrics-backends) +- [Troubleshooting](#troubleshooting) +- [Interceptor System](#interceptor-system) +- [Event Types](#event-types) +- [Creating Custom Interceptors](#creating-custom-interceptors) +- [Advanced Use Cases](#advanced-use-cases) +- [Best Practices](#best-practices) +- [Reference](#reference) + +--- + +### Payload Size Metrics + +Recording `workflow_input_size_bytes` requires JSON-serializing the workflow +input to measure its byte size. This is enabled by default. Override explicitly +via the factory: + +```ruby +# Skip the JSON serialization for large payloads +metrics = MetricsCollector.create(backend: :prometheus, measure_payload_size: false) +``` + +### Collector Lifecycle + +`MetricsCollector` responds to `stop`. Call `stop` to unsubscribe from +process-wide dispatchers (e.g. the `GlobalDispatcher` used for HTTP metrics). +`TaskHandler#stop` calls `stop` on all registered event listeners +automatically. If you manage a collector outside of `TaskHandler`, call `stop` +when the collector is no longer needed. + +--- + +## Quick Start + +```ruby +require 'conductor' + +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) + +# Start metrics HTTP server +metrics_server = Conductor::Worker::Telemetry::MetricsServer.new(port: 9090) +metrics_server.start + +# Create handler with metrics +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [metrics] +) + +handler.start +handler.join + +# Cleanup +metrics_server.stop +``` + +--- + +## Metrics Catalog + +Timing values are seconds. Size values are bytes. Label names use camelCase. +Metric definitions are pre-registered when the Prometheus backend is +created, but specific label-combination time series only appear in +`/metrics` after the corresponding event first records them. + +### Counters + +| Metric | Labels | Description | +|---|---|---| +| `task_poll_total` | `taskType` | Incremented each time the worker issues a poll request. | +| `task_execution_started_total` | `taskType` | Incremented when a polled task is dispatched to the worker function. | +| `task_poll_error_total` | `taskType`, `exception` | Incremented when a poll request fails client-side. | +| `task_execute_error_total` | `taskType`, `exception` | Incremented when the worker function throws. | +| `task_update_error_total` | `taskType`, `exception` | Incremented when updating the task result fails. | +| `task_paused_total` | `taskType` | Incremented when a worker is paused and skips acting on a poll. | +| `thread_uncaught_exceptions_total` | `exception` | Incremented when a worker thread raises an uncaught exception. | +| `workflow_start_error_total` | `workflowType`, `exception` | Incremented when starting a workflow fails client-side. | + +### Time Histograms + +All time histograms use buckets (in seconds): + +```text +0.001, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10 +``` + +| Metric | Labels | Description | +|---|---|---| +| `task_poll_time_seconds` | `taskType`, `status` | Poll request latency. `status` is `SUCCESS` or `FAILURE`. | +| `task_execute_time_seconds` | `taskType`, `status` | Worker function execution duration. `status` is `SUCCESS` or `FAILURE`. | +| `task_update_time_seconds` | `taskType`, `status` | Task-result update latency. `status` is `SUCCESS` or `FAILURE`. | +| `http_api_client_request_seconds` | `method`, `uri`, `status` | HTTP API client request latency. `status` is the HTTP status code as a string, or `"0"` on network failure. | + +Each histogram exposes Prometheus series such as: + +```prometheus +task_execute_time_seconds_bucket{taskType="my_task",status="SUCCESS",le="0.1"} 42.0 +task_execute_time_seconds_count{taskType="my_task",status="SUCCESS"} 50.0 +task_execute_time_seconds_sum{taskType="my_task",status="SUCCESS"} 2.3 +``` + +### Size Histograms + +All size histograms use buckets (in bytes): + +```text +100, 1000, 10000, 100000, 1000000, 10000000 +``` + +| Metric | Labels | Description | +|---|---|---| +| `task_result_size_bytes` | `taskType` | Serialized task result output size. | +| `workflow_input_size_bytes` | `workflowType`, `version` | Serialized workflow input size. `version` is an empty string when the workflow version is absent. | + +### Gauges + +| Metric | Labels | Description | +|---|---|---| +| `active_workers` | `taskType` | Current number of worker threads actively executing tasks. | + +--- + +## Metrics Not Applicable to Ruby + +The cross-SDK canonical catalog defines additional metrics that are not +applicable to the Ruby SDK's runtime model: + +| Canonical metric | Why N/A for Ruby | +|---|---| +| `task_ack_error_total` | Batch-poll response is the ack; there is no separate ack call. | +| `task_ack_failed_total` | Same reason. | +| `task_execution_queue_full_total` | Ruby uses `fallback_policy: :caller_runs` which back-pressures the polling thread instead of rejecting tasks. | +| `worker_restart_total` | Python-only. Its multi-process supervisor restarts child processes. Ruby uses threads, fibers, or ractors. | +| `external_payload_used_total` | Ruby SDK has no external-payload-storage integration. | + +Users cross-referencing the harmonization spec or documentation from other +Conductor SDKs may notice these metrics in other catalogs. Their absence in +the Ruby SDK is intentional. + +### `thread_uncaught_exceptions_total` + +The `thread_uncaught_exceptions_total` counter and its collector handler exist +in the Ruby SDK for API completeness, but the metric is not currently wired to +any runtime event. The harmonization spec defines it as "incremented when a +worker thread terminates with an uncaught exception." In SDKs that implement it +(Java, Go, Rust), it fires only at the thread/goroutine death boundary. Python +and JavaScript also define the metric surface but do not wire it. Future Ruby +SDK versions may connect it to `Thread.report_on_exception` or a similar +mechanism. + +### Ractor Runner Limitations (Work-in-Progress -- Untested) + +> **Warning:** The `RactorTaskRunner` is an experimental work-in-progress. +> Metrics collection from Ractor-based workers is **partially implemented +> and not yet functional**. In practice, **no metrics are delivered** from Ractor workers +> in the current implementation due to the incomplete event bridge described +> below. Do not rely on Ractor worker metrics in production. + +The event communication channel between Ractor workers and the main-thread +`SyncEventDispatcher` is not yet implemented. The `TaskHandler` method +`create_event_receiver_ractor` returns `nil` (with the comment _"Ractor +event communication needs more work"_), and the spawned event-receiver thread +exits immediately without forwarding events. Because `RactorTaskRunner` +receives `nil` as its `event_queue`, the `publish_event` guard +(`return unless @event_queue`) causes **every event to be silently +discarded** -- poll events, execution events, update events, and failure +events alike. As a result, none of the metrics documented in this guide +(counters, histograms, or gauges) will be populated for Ractor-based +workers. + +Additionally: + +- The `active_workers` gauge is **never emitted** by `RactorTaskRunner`. + Unlike `TaskRunner` and `FiberTaskRunner`, which call + `publish_active_workers` when tasks start and complete, the Ractor runner + has no equivalent tracking or publishing logic. +- The `cleanup` method is a no-op stub. +- Ractor shutdown in `TaskHandler#stop` is rudimentary (`ractor.take` with + rescued errors) because Ractors lack a clean shutdown mechanism. + +These limitations will be resolved in a future release when proper +Ractor-to-main-thread event forwarding is implemented. Until then, use +`:thread` isolation (the default) or `:fiber` execution if you need metrics +and observability. + +--- + +## Labels + +| Label | Used by | Values | +|---|---|---| +| `taskType` | Worker metrics | Task definition name. | +| `workflowType` | Workflow metrics | Workflow definition name. | +| `version` | `workflow_input_size_bytes` | Workflow version as a string. Empty string when the version is absent. | +| `status` | Task time metrics | `SUCCESS` or `FAILURE`. For `http_api_client_request_seconds`, the HTTP status code as a string (e.g. `"200"`), or `"0"` on network failure. | +| `exception` | Error counters | Exception class name, such as `Faraday::TimeoutError`. | +| `method` | HTTP metrics | HTTP verb (`GET`, `POST`, etc.). | +| `uri` | HTTP metrics | API-relative path template (e.g. `/tasks/poll/batch/{taskType}`). Dynamic path segments retain `{placeholder}` tokens so label cardinality is bounded. | + +--- + +## Prometheus Integration + +### Setup + +Add the `prometheus-client` gem to your Gemfile: + +```ruby +gem 'prometheus-client', '~> 4.0' +``` + +### Basic Usage + +```ruby +require 'conductor' + +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) + +metrics_server = Conductor::Worker::Telemetry::MetricsServer.new(port: 9090) +metrics_server.start + +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [metrics] +) + +handler.start +handler.join + +metrics_server.stop +``` + +### Metrics Endpoints + +The MetricsServer exposes: + +- `GET /metrics` - Prometheus metrics in text format +- `GET /health` - Health check endpoint (`{"status":"healthy"}`) + +### Kubernetes Integration + +```yaml +# Pod annotations for Prometheus scraping +metadata: + annotations: + prometheus.io/scrape: "true" + prometheus.io/port: "9090" + prometheus.io/path: "/metrics" +``` + +### Custom Prometheus Registry + +```ruby +require 'prometheus/client' + +registry = Prometheus::Client::Registry.new +backend = Conductor::Worker::Telemetry::PrometheusBackend.new(registry: registry) +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: backend) +``` + +--- + +## Custom Metrics Backends + +You can create a custom backend by implementing three methods: + +```ruby +class DatadogBackend + def initialize(statsd_client) + @statsd = statsd_client + end + + def increment(name, labels: {}) + tags = labels.map { |k, v| "#{k}:#{v}" } + @statsd.increment(name, tags: tags) + end + + def observe(name, value, labels: {}) + tags = labels.map { |k, v| "#{k}:#{v}" } + @statsd.histogram(name, value, tags: tags) + end + + def set(name, value, labels: {}) + tags = labels.map { |k, v| "#{k}:#{v}" } + @statsd.gauge(name, value, tags: tags) + end +end + +# Use with MetricsCollector +require 'datadog/statsd' +statsd = Datadog::Statsd.new('localhost', 8125) +metrics = Conductor::Worker::Telemetry::MetricsCollector.create( + backend: DatadogBackend.new(statsd) +) +``` + +--- + +## Troubleshooting + +### Metrics Are Empty + +- Verify that `MetricsCollector.create` is called and the collector is passed + to `TaskHandler` via `event_listeners:`. +- Verify workers have polled or executed tasks. Metrics are created lazily when + the relevant event occurs. +- Confirm the scrape endpoint is reachable at the expected host and port. + +### Missing HTTP or Workflow Metrics + +- `http_api_client_request_seconds` requires the collector to be subscribed to + `GlobalDispatcher` for `HttpApiRequest` events from the HTTP layer. This + happens automatically when `subscribe_global_http: true` (the default). +- `workflow_input_size_bytes` and `workflow_start_error_total` only record when + the corresponding `WorkflowExecutor` events fire. + +### High Cardinality + +- The `uri` label on `http_api_client_request_seconds` uses path templates + (e.g. `/workflow/{workflowId}`) to keep cardinality bounded. If you see + fully-resolved paths in your metrics, verify that HTTP requests are going + through the SDK's `ApiClient` rather than a standalone `RestClient`. +- The `exception` label uses exception class names to keep cardinality bounded. +- Avoid embedding user identifiers or unbounded values in task type, workflow + type, or other label values. + +--- + +## Interceptor System + +The Conductor Ruby SDK provides an event-driven interceptor system that allows +you to: + +- **Monitor performance** - Track polling times, execution durations, error rates +- **Implement custom logging** - Add structured logging for task execution +- **Track errors** - Send failures to error tracking services (Sentry, Bugsnag, etc.) +- **Collect metrics** - Export to Prometheus, Datadog, or custom backends +- **Build alerting** - Monitor SLAs and trigger alerts on violations + +``` +TaskRunner + │ + │ publishes events + ▼ +SyncEventDispatcher ──────► Listener 1 (MetricsCollector) + ──────► Listener 2 (LoggingInterceptor) + ──────► Listener 3 (SentryInterceptor) +``` + +When a worker polls for tasks, executes them, or encounters errors, events are +published to all registered listeners. Listeners can then process these events +independently. + +--- + +## Event Types + +The SDK publishes events during worker execution. Any object that responds to +the corresponding `on_*` method can listen for these events. + +### Poll Events + +| Event | When Published | Key Attributes | +|---|---|---| +| `PollStarted` | Before polling for tasks | `task_type`, `worker_id`, `poll_count` | +| `PollCompleted` | After successful poll | `task_type`, `duration_ms`, `tasks_received` | +| `PollFailure` | When poll fails | `task_type`, `duration_ms`, `cause` | + +### Execution Events + +| Event | When Published | Key Attributes | +|---|---|---| +| `TaskExecutionStarted` | Before task execution | `task_type`, `task_id`, `worker_id`, `workflow_instance_id` | +| `TaskExecutionCompleted` | After successful execution | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `duration_ms`, `output_size_bytes` | +| `TaskExecutionFailure` | When execution fails | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `duration_ms`, `cause`, `is_retryable` | + +### Update Events + +| Event | When Published | Key Attributes | +|---|---|---| +| `TaskUpdateCompleted` | After successful result update | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `duration_ms` | +| `TaskUpdateFailure` | When result update fails after all retries | `task_type`, `task_id`, `worker_id`, `workflow_instance_id`, `cause`, `retry_count`, `task_result`, `duration_ms` | + +### Worker State Events + +| Event | When Published | Key Attributes | +|---|---|---| +| `TaskPaused` | When a paused worker skips a poll | `task_type` | +| `ThreadUncaughtException` | When a worker thread raises an uncaught exception | `cause`, `task_type` | +| `ActiveWorkersChanged` | When the active worker count changes | `task_type`, `count` | + +### Workflow Events + +| Event | When Published | Key Attributes | +|---|---|---| +| `WorkflowStartError` | When starting a workflow fails client-side | `workflow_type`, `version`, `cause` | +| `WorkflowInputSize` | When a workflow is started | `workflow_type`, `version`, `size_bytes` | + +### HTTP Events + +| Event | When Published | Key Attributes | +|---|---|---| +| `HttpApiRequest` | After every HTTP API client request | `method`, `uri`, `status`, `duration_ms` | + +**Important**: `TaskUpdateFailure` is a critical event indicating that a task +result was lost. You should always handle this event to prevent silent data +loss. + +--- + +## Creating Custom Interceptors + +### Basic Structure + +An interceptor is any object that responds to one or more `on_*` methods: + +```ruby +class MyInterceptor + def on_poll_started(event) + # Called before each poll + end + + def on_poll_completed(event) + # Called after successful poll + end + + def on_poll_failure(event) + # Called when poll fails + end + + def on_task_execution_started(event) + # Called before task execution + end + + def on_task_execution_completed(event) + # Called after successful execution + end + + def on_task_execution_failure(event) + # Called when execution fails + end + + def on_task_update_failure(event) + # Called when result update fails (CRITICAL) + end +end +``` + +### Error Tracking Interceptor (Sentry) + +```ruby +require 'sentry-ruby' + +class SentryInterceptor + def on_task_execution_failure(event) + Sentry.capture_exception(event.cause, extra: { + task_id: event.task_id, + task_type: event.task_type, + workflow_instance_id: event.workflow_instance_id, + duration_ms: event.duration_ms, + is_retryable: event.is_retryable + }) + end + + def on_task_update_failure(event) + Sentry.capture_message( + "CRITICAL: Task result lost after #{event.retry_count} retries", + level: :fatal, + extra: { + task_id: event.task_id, + task_type: event.task_type, + workflow_instance_id: event.workflow_instance_id + } + ) + end +end +``` + +### Structured Logging Interceptor + +```ruby +require 'json' + +class StructuredLoggingInterceptor + def initialize(output: $stdout) + @output = output + end + + def on_task_execution_started(event) + log('task_started', event) + end + + def on_task_execution_completed(event) + log('task_completed', event, duration_ms: event.duration_ms) + end + + def on_task_execution_failure(event) + log('task_failed', event, + duration_ms: event.duration_ms, + error: event.cause.class.name, + message: event.cause.message, + retryable: event.is_retryable) + end + + private + + def log(action, event, extra = {}) + entry = { + timestamp: event.timestamp.iso8601(3), + action: action, + task_type: event.task_type, + task_id: event.task_id, + worker_id: event.worker_id, + workflow_instance_id: event.workflow_instance_id, + **extra + } + @output.puts(entry.to_json) + end +end +``` + +--- + +## Advanced Use Cases + +### SLA Monitor + +Alert when tasks exceed duration thresholds: + +```ruby +class SLAMonitor + def initialize(thresholds:, alerter:) + @thresholds = thresholds # { 'task_type' => max_ms } + @alerter = alerter + end + + def on_task_execution_completed(event) + threshold = @thresholds[event.task_type] + return unless threshold && event.duration_ms > threshold + + @alerter.alert( + type: :sla_violation, + task_type: event.task_type, + task_id: event.task_id, + duration_ms: event.duration_ms, + threshold_ms: threshold + ) + end +end + +# Usage +sla_monitor = SLAMonitor.new( + thresholds: { + 'process_order' => 5000, + 'send_email' => 2000 + }, + alerter: SlackAlerter.new(webhook_url: ENV['SLACK_WEBHOOK']) +) +``` + +### Cost Tracker + +Track compute costs per task type: + +```ruby +class CostTracker + def initialize(cost_per_ms: {}) + @cost_per_ms = cost_per_ms + @costs = Hash.new(0.0) + @mutex = Mutex.new + end + + def on_task_execution_completed(event) + rate = @cost_per_ms[event.task_type] || 0.0001 + cost = rate * event.duration_ms + @mutex.synchronize { @costs[event.task_type] += cost } + end + + def report + @mutex.synchronize { @costs.dup } + end +end +``` + +### Retry Tracker + +Monitor retry patterns: + +```ruby +class RetryTracker + def initialize + @retries = Hash.new { |h, k| h[k] = [] } + @mutex = Mutex.new + end + + def on_task_execution_failure(event) + return unless event.is_retryable + + @mutex.synchronize do + @retries[event.task_type] << { + task_id: event.task_id, + error: event.cause.class.name, + timestamp: event.timestamp + } + end + end + + def retry_rate(task_type, window_seconds: 300) + @mutex.synchronize do + cutoff = Time.now - window_seconds + recent = @retries[task_type].select { |r| r[:timestamp] > cutoff } + recent.size + end + end +end +``` + +--- + +## Best Practices + +### 1. Keep Interceptors Fast + +Interceptors run synchronously in the worker thread. Keep processing fast to +avoid impacting task execution: + +```ruby +# BAD: Slow synchronous HTTP call +def on_task_execution_completed(event) + HTTParty.post('https://analytics.example.com', body: event.to_h.to_json) +end + +# GOOD: Queue for background processing +def on_task_execution_completed(event) + @queue << event.to_h +end +``` + +### 2. Handle Errors in Interceptors + +Errors in interceptors are caught and logged but don't affect other +interceptors: + +```ruby +def on_task_execution_completed(event) + external_service.track(event) +rescue => e + @logger.warn("Failed to track: #{e.message}") +end +``` + +### 3. Always Handle TaskUpdateFailure + +This is a critical event -- task results are lost: + +```ruby +def on_task_update_failure(event) + @logger.fatal("Task result lost: #{event.task_id}") + PagerDuty.trigger(severity: :critical, summary: "Task result lost: #{event.task_id}") + FailedTaskStore.save(event.task_result) +end +``` + +### 4. Use Multiple Interceptors + +Separate concerns into different interceptors: + +```ruby +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [ + Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus), + StructuredLoggingInterceptor.new, + SentryInterceptor.new, + SLAMonitor.new(thresholds: sla_config), + ] +) +``` + +### 5. Test Your Interceptors + +```ruby +RSpec.describe SentryInterceptor do + let(:interceptor) { described_class.new } + + describe '#on_task_execution_failure' do + it 'captures exception in Sentry' do + event = Conductor::Worker::Events::TaskExecutionFailure.new( + task_type: 'my_task', + task_id: 'task-123', + worker_id: 'worker-1', + workflow_instance_id: 'workflow-456', + duration_ms: 100, + cause: StandardError.new('Test error'), + is_retryable: true + ) + + expect(Sentry).to receive(:capture_exception) + interceptor.on_task_execution_failure(event) + end + end +end +``` + +--- + +## Reference + +### Telemetry Classes + +- `Conductor::Worker::Telemetry::MetricsCollector` - Canonical metric collector; `.create` is a convenience factory +- `Conductor::Worker::Telemetry::NullBackend` - No-op metrics backend +- `Conductor::Worker::Telemetry::PrometheusBackend` - Prometheus backend with canonical label schemas +- `Conductor::Worker::Telemetry::MetricsServer` - WEBrick HTTP server for `/metrics` and `/health` endpoints + +### Event Classes + +All events are in the `Conductor::Worker::Events` namespace: + +- `PollStarted`, `PollCompleted`, `PollFailure` +- `TaskExecutionStarted`, `TaskExecutionCompleted`, `TaskExecutionFailure` +- `TaskUpdateCompleted`, `TaskUpdateFailure` +- `TaskPaused`, `ThreadUncaughtException`, `ActiveWorkersChanged` +- `WorkflowStartError`, `WorkflowInputSize` +- `HttpApiRequest` + +### Registration Methods + +Register interceptors via `TaskHandler`: + +```ruby +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [listener1, listener2] +) +``` + +Or register manually with the event dispatcher: + +```ruby +dispatcher = handler.event_dispatcher +dispatcher.register(Conductor::Worker::Events::PollStarted, ->(event) { puts event }) +``` + +--- + +## Detailed Technical Notes -- [Unreleased] + +### Path template `uri` label (metric_uri) + +The `uri` label on `http_api_client_request_seconds` now carries the **path +template** (e.g. `/workflow/{workflowId}`) rather than the fully-resolved +request path (e.g. `/api/workflow/abc-123-def`). This keeps metric label +cardinality bounded regardless of how many unique workflow IDs, task types, or +other dynamic path segments pass through the SDK. + +**Data flow:** + +1. Every API resource method calls `ApiClient#call_api` with a `resource_path` + that contains `{placeholder}` tokens (e.g. `/tasks/poll/batch/{taskType}`). +2. `ApiClient#call_api_no_retry` saves `resource_path` as `metric_uri` *before* + substituting path parameters. +3. The substituted path is concatenated with `server_url` to form the HTTP URL. + `metric_uri` is passed alongside as a keyword argument to + `RestClient#request`. +4. `RestClient#emit_http_event` prefers `metric_uri` when present; it only + falls back to extracting the path from the full URL when `metric_uri` is + `nil` (e.g. for direct `RestClient` calls outside of `ApiClient`). +5. The `HttpApiRequest` event carries the template string as its `uri` field. + `MetricsCollector#on_http_api_request` records it as-is into the histogram. + +This approach mirrors the Python SDK (`metric_uri` parameter), the Java SDK +(`PathTemplateTag` on the OkHttp request), and the Go SDK (`WithPathTemplate` +context value). The base-URL path prefix (e.g. `/api`) is never included +because `metric_uri` is always the raw API-relative resource path. + +### MetricsCollector factory + +`MetricsCollector.create(backend:)` is a convenience constructor that returns +a `MetricsCollector` instance. It emits the harmonized cross-SDK catalog with +`camelCase` domain labels, Prometheus histograms with explicit bucket +boundaries, and the `exception` label derived from the Ruby exception class +name. + +### Event system + +The following event types are used by the collector: + +- `PollStarted`, `PollCompleted`, `PollFailure`, `TaskExecutionStarted`, + `TaskExecutionCompleted`, `TaskExecutionFailure`, `TaskUpdateCompleted`, + `TaskUpdateFailure`, `TaskPaused`, `ActiveWorkersChanged` -- emitted by + `TaskRunner` (and `FiberTaskRunner`). `RactorTaskRunner` constructs these + same events internally but **does not deliver them** to the + `SyncEventDispatcher` because the Ractor-to-main-thread event bridge is + not yet implemented (see + [Ractor Runner Limitations](#ractor-runner-limitations-work-in-progress----untested)). +- `WorkflowStartError`, `WorkflowInputSize` -- emitted by + `WorkflowExecutor`. +- `HttpApiRequest` -- emitted by `RestClient` via the process-wide + `GlobalDispatcher`, but only when at least one `HttpApiRequest` listener + is subscribed (i.e. a `MetricsCollector` is active). With no collector, + `RestClient` skips all timing overhead. +- `ThreadUncaughtException` -- event class and collector handler exist for + API completeness but are not currently emitted by any runner (see + [thread_uncaught_exceptions_total](#thread_uncaught_exceptions_total)). + +All events flow through the `SyncEventDispatcher` -> listener registry -> +collector pattern. The `GlobalDispatcher` singleton provides a secondary +channel so that `RestClient` (which has no direct reference to the task +handler's dispatcher) can still emit HTTP events. diff --git a/docs/design/EVENT_INTERCEPTOR_SYSTEM.md b/docs/design/EVENT_INTERCEPTOR_SYSTEM.md new file mode 100644 index 0000000..3ea5885 --- /dev/null +++ b/docs/design/EVENT_INTERCEPTOR_SYSTEM.md @@ -0,0 +1,907 @@ +# Event-Driven Interceptor System - Design Document + +## Table of Contents + +1. [Overview](#overview) +2. [Architecture](#architecture) +3. [Core Components](#core-components) +4. [Event Hierarchy](#event-hierarchy) +5. [Event Dispatcher](#event-dispatcher) +6. [Listener Protocol](#listener-protocol) +7. [Listener Registration](#listener-registration) +8. [Metrics Collection](#metrics-collection) +9. [Prometheus Integration](#prometheus-integration) +10. [Usage Examples](#usage-examples) +11. [Advanced Use Cases](#advanced-use-cases) +12. [Performance Considerations](#performance-considerations) +13. [File Structure](#file-structure) + +--- + +## Overview + +### Purpose + +The Event-Driven Interceptor System provides a decoupled, extensible mechanism for observing and reacting to task execution lifecycle events in the Conductor Ruby SDK. This enables: + +- **Metrics Collection** - Track poll times, execution durations, error rates +- **Custom Interceptors** - Add logging, tracing, auditing without modifying core code +- **SLA Monitoring** - Alert on tasks exceeding thresholds +- **Cost Tracking** - Monitor compute costs per task type +- **Error Tracking** - Send failures to external services (Sentry, Bugsnag, etc.) + +### Design Goals + +| Goal | Description | +|------|-------------| +| **Decoupled** | Event publishing is separate from event handling | +| **Thread-Safe** | Safe for concurrent task execution | +| **Extensible** | Add listeners without modifying SDK code | +| **Non-Blocking** | Listener failures never block worker execution | +| **Type-Safe** | Clear event contracts with documented attributes | +| **Pluggable** | Multiple metrics backends (null, Prometheus, custom) | + +### Non-Goals + +- **Distributed Tracing** - OpenTelemetry integration is a separate concern +- **Built-in Dashboards** - Users provide their own visualization +- **Async Dispatch** - Events are dispatched synchronously for simplicity + +--- + +## Architecture + +### High-Level Overview + +``` +┌─────────────────────────────────────────────────────────────────────────┐ +│ Task Execution Layer │ +│ ┌──────────────────┐ ┌──────────────────┐ │ +│ │ TaskRunner │ │ TaskHandler │ │ +│ │ (polling loop) │ │ (orchestrator) │ │ +│ └────────┬─────────┘ └────────┬─────────┘ │ +│ │ publish() │ register() │ +└───────────┼──────────────────────────────┼──────────────────────────────┘ + │ │ + ▼ ▼ +┌─────────────────────────────────────────────────────────────────────────┐ +│ Event Dispatch Layer │ +│ ┌────────────────────────────────────────────────────────────────────┐ │ +│ │ SyncEventDispatcher │ │ +│ │ • Thread-safe listener registration (Mutex) │ │ +│ │ • Synchronous event dispatch │ │ +│ │ • Error isolation (listener failures logged, not propagated) │ │ +│ │ • Type-based routing (event.class → listeners) │ │ +│ └──────────────────────────┬─────────────────────────────────────────┘ │ +│ │ dispatch │ +└──────────────────────────────┼──────────────────────────────────────────┘ + │ +┌──────────────────────────────▼──────────────────────────────────────────┐ +│ Listener/Consumer Layer │ +│ ┌────────────────┐ ┌────────────────┐ ┌─────────────────────────┐ │ +│ │MetricsCollector│ │ CustomListener │ │ SLA Monitor │ │ +│ │ (Prometheus) │ │ (Logging) │ │ (Alerting) │ │ +│ └────────────────┘ └────────────────┘ └─────────────────────────┘ │ +│ ┌────────────────┐ ┌────────────────┐ ┌─────────────────────────┐ │ +│ │ Audit Logger │ │ Cost Tracker │ │ Error Reporter │ │ +│ │ (Compliance) │ │ (FinOps) │ │ (Sentry/Bugsnag) │ │ +│ └────────────────┘ └────────────────┘ └─────────────────────────┘ │ +└─────────────────────────────────────────────────────────────────────────┘ +``` + +### Component Interaction Flow + +``` +TaskRunner SyncEventDispatcher Listeners + │ │ │ + │ publish(PollStarted) │ │ + │─────────────────────────────>│ │ + │ │ call(event) ──────────────>│ MetricsCollector + │ │ call(event) ──────────────>│ CustomListener + │ │<───────────────────────────│ + │<─────────────────────────────│ │ + │ │ │ + │ (execute task) │ │ + │ │ │ + │ publish(TaskExecutionCompleted) │ + │─────────────────────────────>│ │ + │ │ call(event) ──────────────>│ MetricsCollector + │ │ call(event) ──────────────>│ CustomListener + │ │<───────────────────────────│ + │<─────────────────────────────│ │ +``` + +--- + +## Core Components + +### Component Summary + +| Component | Location | Purpose | +|-----------|----------|---------| +| `ConductorEvent` | `events/conductor_event.rb` | Base event class with timestamp | +| `TaskRunnerEvent` | `events/conductor_event.rb` | Base for task runner events | +| `PollStarted`, etc. | `events/task_runner_events.rb` | Specific event types | +| `SyncEventDispatcher` | `events/sync_event_dispatcher.rb` | Thread-safe event router | +| `TaskRunnerEventsListener` | `events/listeners.rb` | Listener protocol (duck typing) | +| `ListenerRegistry` | `events/listener_registry.rb` | Bulk listener registration | +| `MetricsCollector` | `telemetry/metrics_collector.rb` | Canonical metric collector | +| `NullBackend` | `telemetry/metrics_collector.rb` | No-op backend | +| `PrometheusBackend` | `telemetry/prometheus_backend.rb` | Prometheus backend with canonical label schemas | +| `MetricsServer` | `telemetry/prometheus_backend.rb` | WEBrick HTTP server for `/metrics` | + +--- + +## Event Hierarchy + +### Class Hierarchy + +``` +ConductorEvent # Base - provides timestamp +└── TaskRunnerEvent # Base for task runner - adds task_type + ├── PollStarted # Polling started + ├── PollCompleted # Polling completed successfully + ├── PollFailure # Polling failed + ├── TaskExecutionStarted # Task execution started + ├── TaskExecutionCompleted # Task execution completed + ├── TaskExecutionFailure # Task execution failed + └── TaskUpdateFailure # Task result update failed (CRITICAL) +``` + +### Event Attributes + +#### ConductorEvent (Base) + +```ruby +class ConductorEvent + attr_reader :timestamp # Time - UTC timestamp when event was created + + def to_h + { timestamp: @timestamp.iso8601(3) } + end +end +``` + +#### TaskRunnerEvent (Base) + +```ruby +class TaskRunnerEvent < ConductorEvent + attr_reader :task_type # String - Task definition name +end +``` + +#### PollStarted + +Published when polling starts for a task type. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `worker_id` | String | Unique worker identifier | +| `poll_count` | Integer | Number of polls performed so far | + +#### PollCompleted + +Published when polling completes successfully. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `duration_ms` | Float | Duration of poll in milliseconds | +| `tasks_received` | Integer | Number of tasks received | + +#### PollFailure + +Published when polling fails. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `duration_ms` | Float | Duration of poll in milliseconds | +| `cause` | Exception | The exception that caused the failure | + +#### TaskExecutionStarted + +Published when task execution starts. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `task_id` | String | Unique task identifier | +| `worker_id` | String | Unique worker identifier | +| `workflow_instance_id` | String | Workflow instance identifier | + +#### TaskExecutionCompleted + +Published when task execution completes successfully. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `task_id` | String | Unique task identifier | +| `worker_id` | String | Unique worker identifier | +| `workflow_instance_id` | String | Workflow instance identifier | +| `duration_ms` | Float | Duration of execution in milliseconds | +| `output_size_bytes` | Integer | Size of output data in bytes (optional) | + +#### TaskExecutionFailure + +Published when task execution fails. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `task_id` | String | Unique task identifier | +| `worker_id` | String | Unique worker identifier | +| `workflow_instance_id` | String | Workflow instance identifier | +| `duration_ms` | Float | Duration of execution in milliseconds | +| `cause` | Exception | The exception that caused the failure | +| `is_retryable` | Boolean | Whether the error is retryable | + +#### TaskUpdateFailure (CRITICAL) + +Published when task result update fails after all retries. This is a **critical** event - the task result is lost. + +| Attribute | Type | Description | +|-----------|------|-------------| +| `task_type` | String | Task definition name | +| `task_id` | String | Unique task identifier | +| `worker_id` | String | Unique worker identifier | +| `workflow_instance_id` | String | Workflow instance identifier | +| `cause` | Exception | The exception that caused the failure | +| `retry_count` | Integer | Number of retry attempts made | +| `task_result` | TaskResult | The task result that failed to update (for recovery) | + +--- + +## Event Dispatcher + +### SyncEventDispatcher + +The `SyncEventDispatcher` is a thread-safe, synchronous event dispatcher that routes events to registered listeners. + +```ruby +module Conductor::Worker::Events + class SyncEventDispatcher + def initialize + @listeners = Hash.new { |h, k| h[k] = [] } + @mutex = Mutex.new + end + + # Register a listener for an event type + # @param event_type [Class] Event class to listen for + # @param listener [Proc, #call] Callable to invoke when event is published + # @return [self] + def register(event_type, listener) + @mutex.synchronize do + @listeners[event_type] << listener unless @listeners[event_type].include?(listener) + end + self + end + + # Unregister a listener for an event type + # @param event_type [Class] Event class + # @param listener [Proc, #call] Listener to remove + # @return [self] + def unregister(event_type, listener) + @mutex.synchronize do + @listeners[event_type].delete(listener) + end + self + end + + # Publish an event to all registered listeners + # @param event [ConductorEvent] Event to publish + # @return [self] + def publish(event) + listeners = @mutex.synchronize { @listeners[event.class].dup } + + listeners.each do |listener| + listener.call(event) + rescue StandardError => e + # Listener failure is isolated - never breaks the worker + warn "[Conductor] Event listener error for #{event.class}: #{e.message}" + end + + self + end + + # Check if there are listeners registered for an event type + def has_listeners?(event_type) + @mutex.synchronize { @listeners[event_type].any? } + end + + # Get the number of listeners for an event type + def listener_count(event_type) + @mutex.synchronize { @listeners[event_type].size } + end + + # Clear all listeners + def clear + @mutex.synchronize { @listeners.clear } + self + end + end +end +``` + +### Key Design Decisions + +| Decision | Rationale | +|----------|-----------| +| **Synchronous dispatch** | Simpler than async, avoids ordering issues | +| **Mutex for thread safety** | Protects listener list during registration and iteration | +| **Copy listeners before dispatch** | Allows modification during dispatch without deadlock | +| **Error isolation** | Listener exceptions are logged but don't propagate | +| **Type-based routing** | Events routed by class, not inheritance hierarchy | + +### Thread Safety Guarantees + +1. **Registration is thread-safe** - Multiple threads can register listeners concurrently +2. **Publishing is thread-safe** - Multiple threads can publish events concurrently +3. **Listeners are called sequentially** - Within a single publish call +4. **Listener exceptions are isolated** - One listener failure doesn't affect others + +--- + +## Listener Protocol + +### TaskRunnerEventsListener + +The listener protocol uses duck typing - implement only the methods you need: + +```ruby +module Conductor::Worker::Events + # Listener protocol for task runner events + # Include this module to document the expected interface + # All methods are optional - implement only the ones you need + module TaskRunnerEventsListener + # Called when polling starts + # @param event [PollStarted] + def on_poll_started(event); end + + # Called when polling completes successfully + # @param event [PollCompleted] + def on_poll_completed(event); end + + # Called when polling fails + # @param event [PollFailure] + def on_poll_failure(event); end + + # Called when task execution starts + # @param event [TaskExecutionStarted] + def on_task_execution_started(event); end + + # Called when task execution completes successfully + # @param event [TaskExecutionCompleted] + def on_task_execution_completed(event); end + + # Called when task execution fails + # @param event [TaskExecutionFailure] + def on_task_execution_failure(event); end + + # Called when task update fails after all retries (CRITICAL) + # @param event [TaskUpdateFailure] + def on_task_update_failure(event); end + end +end +``` + +### Implementation Example + +```ruby +class MyListener + # Only implement the methods you care about + def on_task_execution_completed(event) + puts "Task #{event.task_id} completed in #{event.duration_ms}ms" + end + + def on_task_execution_failure(event) + puts "Task #{event.task_id} FAILED: #{event.cause.message}" + end +end +``` + +--- + +## Listener Registration + +### ListenerRegistry + +The `ListenerRegistry` provides bulk registration of listener objects: + +```ruby +module Conductor::Worker::Events + class ListenerRegistry + # Mapping of event classes to listener method names + EVENT_METHOD_MAP = { + PollStarted => :on_poll_started, + PollCompleted => :on_poll_completed, + PollFailure => :on_poll_failure, + TaskExecutionStarted => :on_task_execution_started, + TaskExecutionCompleted => :on_task_execution_completed, + TaskExecutionFailure => :on_task_execution_failure, + TaskUpdateFailure => :on_task_update_failure + }.freeze + + # Register a listener object with the dispatcher + # Auto-detects implemented methods via respond_to? + # @param listener [Object] Object implementing TaskRunnerEventsListener methods + # @param dispatcher [SyncEventDispatcher] Event dispatcher + def self.register_task_runner_listener(listener, dispatcher) + EVENT_METHOD_MAP.each do |event_class, method_name| + if listener.respond_to?(method_name) + dispatcher.register(event_class, ->(event) { listener.send(method_name, event) }) + end + end + end + + # Register multiple listeners with the dispatcher + # @param listeners [Array] Array of listener objects + # @param dispatcher [SyncEventDispatcher] Event dispatcher + def self.register_all(listeners, dispatcher) + listeners.each do |listener| + register_task_runner_listener(listener, dispatcher) + end + end + end +end +``` + +### Usage in TaskHandler + +```ruby +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [MyListener.new, AnotherListener.new] +) +``` + +--- + +## Metrics Collection + +### MetricsCollector + +`MetricsCollector.create` returns a collector that emits the canonical +(harmonized) metric surface: + +```ruby +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) +``` + +See [docs/METRICS_AND_INTERCEPTORS.md](../METRICS_AND_INTERCEPTORS.md) for the +full metrics catalog and label reference. + +### Backend Protocol + +Metrics backends must implement these methods: + +```ruby +# Increment a counter +# @param name [String] Metric name +# @param labels [Hash] Metric labels +def increment(name, labels: {}) +end + +# Observe a value (histogram) +# @param name [String] Metric name +# @param value [Numeric] Value to observe +# @param labels [Hash] Metric labels +def observe(name, value, labels: {}) +end + +# Set a gauge value +# @param name [String] Metric name +# @param value [Numeric] Value to set +# @param labels [Hash] Metric labels +def set(name, value, labels: {}) +end +``` + +### NullBackend + +A no-op backend for when metrics are disabled: + +```ruby +class NullBackend + def increment(name, labels: {}); end + def observe(name, value, labels: {}); end + def set(name, value, labels: {}); end +end +``` + +--- + +## Prometheus Integration + +### Prometheus Backends + +The SDK ships `PrometheusBackend` with canonical metric registrations using +`taskType` labels, `status` on time histograms, and canonical bucket +boundaries. It implements `increment`, `observe`, and `set` and integrates +with the `prometheus-client` gem. See +[docs/METRICS_AND_INTERCEPTORS.md](../METRICS_AND_INTERCEPTORS.md) +for the full metric catalog emitted by each backend. + +### MetricsServer + +An optional HTTP server for exposing Prometheus metrics: + +```ruby +module Conductor::Worker::Telemetry + class MetricsServer + DEFAULT_PORT = 9090 + + def initialize(port: DEFAULT_PORT, registry: nil) + @port = port + @registry = registry || Prometheus::Client.registry + end + + def start + require 'webrick' + @server = WEBrick::HTTPServer.new(Port: @port, Logger: WEBrick::Log.new('/dev/null')) + + @server.mount_proc '/metrics' do |_req, res| + res.content_type = 'text/plain; version=0.0.4' + res.body = Prometheus::Client::Formats::Text.marshal(@registry) + end + + @server.mount_proc '/health' do |_req, res| + res.body = '{"status":"healthy"}' + end + + @thread = Thread.new { @server.start } + end + + def stop + @server&.shutdown + @thread&.join(5) + end + end +end +``` + +--- + +## Usage Examples + +### Basic Metrics Collection + +```ruby +require 'conductor' + +# Create configuration +config = Conductor::Configuration.new( + server_api_url: 'https://conductor.example.com/api', + key_id: 'key', + key_secret: 'secret' +) + +# Create metrics collector with Prometheus backend +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) + +# Start metrics server +metrics_server = Conductor::Worker::Telemetry::MetricsServer.new(port: 9090) +metrics_server.start + +# Create task handler with metrics +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [metrics] +) + +# Start workers +handler.start +handler.join +``` + +### Custom Logging Interceptor + +```ruby +class LoggingInterceptor + def initialize(logger = Logger.new($stdout)) + @logger = logger + end + + def on_poll_started(event) + @logger.debug("Polling for #{event.task_type}...") + end + + def on_poll_completed(event) + @logger.debug("Poll for #{event.task_type}: #{event.tasks_received} tasks in #{event.duration_ms}ms") + end + + def on_task_execution_started(event) + @logger.info("Starting task #{event.task_id} (#{event.task_type})") + end + + def on_task_execution_completed(event) + @logger.info("Completed task #{event.task_id} in #{event.duration_ms}ms") + end + + def on_task_execution_failure(event) + @logger.error("Task #{event.task_id} FAILED: #{event.cause.message}") + @logger.error(event.cause.backtrace.first(5).join("\n")) + end + + def on_task_update_failure(event) + @logger.fatal("CRITICAL: Task #{event.task_id} result LOST after #{event.retry_count} retries!") + end +end + +# Use with TaskHandler +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [LoggingInterceptor.new] +) +``` + +### Error Tracking (Sentry Integration) + +```ruby +class SentryInterceptor + def on_task_execution_failure(event) + Sentry.capture_exception(event.cause, extra: { + task_id: event.task_id, + task_type: event.task_type, + workflow_instance_id: event.workflow_instance_id, + duration_ms: event.duration_ms, + is_retryable: event.is_retryable + }) + end + + def on_task_update_failure(event) + Sentry.capture_message( + "CRITICAL: Task result lost", + level: :fatal, + extra: { + task_id: event.task_id, + task_type: event.task_type, + retry_count: event.retry_count + } + ) + end +end +``` + +### Multiple Listeners + +```ruby +# Combine metrics, logging, and error tracking +handler = Conductor::Worker::TaskHandler.new( + configuration: config, + event_listeners: [ + Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus), + LoggingInterceptor.new, + SentryInterceptor.new + ] +) +``` + +--- + +## Advanced Use Cases + +### SLA Monitor + +Monitor task execution times and alert on SLA violations: + +```ruby +class SLAMonitor + def initialize(thresholds:, alerter:) + @thresholds = thresholds # { 'task_type' => max_duration_ms } + @alerter = alerter + end + + def on_task_execution_completed(event) + threshold = @thresholds[event.task_type] + return unless threshold && event.duration_ms > threshold + + @alerter.alert( + type: :sla_violation, + task_type: event.task_type, + task_id: event.task_id, + duration_ms: event.duration_ms, + threshold_ms: threshold + ) + end +end + +# Usage +sla_monitor = SLAMonitor.new( + thresholds: { + 'process_order' => 5000, # 5 seconds + 'send_email' => 2000, # 2 seconds + 'generate_report' => 30000 # 30 seconds + }, + alerter: SlackAlerter.new(webhook_url: ENV['SLACK_WEBHOOK']) +) +``` + +### Cost Tracker + +Track compute costs per task type: + +```ruby +class CostTracker + def initialize(cost_per_ms:, reporting_interval: 60) + @cost_per_ms = cost_per_ms # { 'task_type' => cost_per_ms } + @costs = Hash.new(0.0) + @mutex = Mutex.new + @reporting_interval = reporting_interval + start_reporting_thread + end + + def on_task_execution_completed(event) + cost = (@cost_per_ms[event.task_type] || 0.0001) * event.duration_ms + @mutex.synchronize { @costs[event.task_type] += cost } + end + + private + + def start_reporting_thread + Thread.new do + loop do + sleep @reporting_interval + report_costs + end + end + end + + def report_costs + @mutex.synchronize do + total = @costs.values.sum + puts "Cost Report: Total=$#{format('%.4f', total)}" + @costs.each { |task_type, cost| puts " #{task_type}: $#{format('%.4f', cost)}" } + @costs.clear + end + end +end +``` + +### Audit Logger + +Log all task executions for compliance: + +```ruby +class AuditLogger + def initialize(log_file:) + @logger = Logger.new(log_file) + end + + def on_task_execution_started(event) + log_entry('STARTED', event) + end + + def on_task_execution_completed(event) + log_entry('COMPLETED', event, duration_ms: event.duration_ms) + end + + def on_task_execution_failure(event) + log_entry('FAILED', event, + duration_ms: event.duration_ms, + error: event.cause.class.name, + message: event.cause.message, + retryable: event.is_retryable) + end + + private + + def log_entry(status, event, extra = {}) + @logger.info({ + timestamp: event.timestamp.iso8601(3), + status: status, + task_type: event.task_type, + task_id: event.task_id, + worker_id: event.worker_id, + workflow_instance_id: event.workflow_instance_id, + **extra + }.to_json) + end +end +``` + +--- + +## Performance Considerations + +### Event Publishing Overhead + +Event publishing is synchronous but lightweight: + +1. **Mutex acquisition** - ~100ns on uncontended lock +2. **List copy** - O(n) where n = number of listeners (typically 1-5) +3. **Listener calls** - Dependent on listener implementation + +**Typical overhead**: < 1ms per event with 3 listeners + +### Recommendations + +| Concern | Recommendation | +|---------|----------------| +| **Many listeners** | Keep listener count low (< 10) | +| **Slow listeners** | Offload heavy work to background threads | +| **High-frequency events** | Consider sampling in custom listeners | +| **Logging** | Use async logging (Logger with queue) | +| **Metrics** | Prometheus client is thread-safe and efficient | + +### Thread Pool Sizing + +The event system doesn't use a separate thread pool. Events are processed in the TaskRunner thread. This means: + +- **Listener execution time** directly impacts polling interval +- **Blocking operations** in listeners will block task polling +- **Keep listeners fast** (< 10ms) or offload to background + +### Error Isolation Example + +```ruby +# If listener A fails, listener B still runs +class FailingListener + def on_task_execution_completed(event) + raise "Intentional failure" # This is caught and logged + end +end + +class WorkingListener + def on_task_execution_completed(event) + puts "Still runs!" # This executes even if FailingListener fails + end +end +``` + +--- + +## File Structure + +``` +lib/conductor/worker/ +├── events/ +│ ├── conductor_event.rb # Base event class + TaskRunnerEvent +│ ├── task_runner_events.rb # All task runner event types +│ ├── sync_event_dispatcher.rb # Thread-safe event dispatcher +│ ├── listeners.rb # TaskRunnerEventsListener protocol +│ └── listener_registry.rb # Bulk listener registration helper +├── telemetry/ +│ ├── metrics_collector.rb # MetricsCollector class + NullBackend +│ └── prometheus_backend.rb # PrometheusBackend + MetricsServer +├── task_runner.rb # Publishes events during polling/execution +└── task_handler.rb # Creates dispatcher, registers listeners + +spec/conductor/worker/ +├── events/ +│ ├── conductor_event_spec.rb +│ ├── task_runner_events_spec.rb +│ ├── sync_event_dispatcher_spec.rb +│ └── listener_registry_spec.rb +└── telemetry/ + ├── metrics_collector_spec.rb + └── prometheus_backend_spec.rb +``` + +--- + +## Comparison to Python SDK + +| Aspect | Python SDK | Ruby SDK | +|--------|------------|----------| +| **Dispatch model** | Async (asyncio.create_task) | Sync (same thread) | +| **Thread safety** | asyncio.Lock | Mutex | +| **Listener protocol** | typing.Protocol | Duck typing (respond_to?) | +| **Event classes** | @dataclass(frozen=True) | attr_reader + to_h | +| **Metrics backend** | Prometheus multiprocess | Prometheus single process | +| **Error isolation** | ✅ Caught and logged | ✅ Caught and logged | +| **Event types** | Same 7 event types | Same 7 event types | + +### Why Synchronous in Ruby? + +The Python SDK uses async dispatch because: +1. Python uses asyncio for workers +2. Async dispatch avoids blocking the event loop + +The Ruby SDK uses synchronous dispatch because: +1. Ruby workers use threads (GVL releases on I/O) +2. Simpler implementation with predictable ordering +3. Listeners typically complete in < 1ms +4. Thread-per-worker model already provides isolation diff --git a/docs/design/WORKER_DESIGN.md b/docs/design/WORKER_DESIGN.md new file mode 100644 index 0000000..1830ba2 --- /dev/null +++ b/docs/design/WORKER_DESIGN.md @@ -0,0 +1,1781 @@ +# Conductor Ruby SDK - Worker Infrastructure Design + +## Table of Contents + +1. [Overview & Goals](#overview--goals) +2. [Architecture Overview](#architecture-overview) +3. [Ruby vs Python Concurrency](#ruby-vs-python-concurrency) +4. [Core Components](#core-components) +5. [Three Runner Models](#three-runner-models) +6. [Worker Definition Patterns](#worker-definition-patterns) +7. [Task Context System](#task-context-system) +8. [Behavioral Algorithms](#behavioral-algorithms) +9. [Event System](#event-system) +10. [Configuration System](#configuration-system) +11. [Task Definition Auto-Registration](#task-definition-auto-registration) +12. [File Structure](#file-structure) +13. [Implementation Phases](#implementation-phases) + +--- + +## Overview & Goals + +### Purpose + +This document specifies the design for the Conductor Ruby SDK's worker infrastructure - the system that polls for tasks from a Conductor server, executes them using user-defined workers, and reports results back. + +### Goals + +1. **Full parity with Python SDK** - Support all features from the Python worker SDK including batch polling, adaptive backoff, capacity management, events, and metrics +2. **Ruby-idiomatic API** - Use Ruby conventions (blocks, mixins, snake_case) while maintaining the same capabilities +3. **Production-grade** - Handle edge cases, failures, and high-throughput scenarios reliably +4. **Extensible** - Support custom event listeners, metrics backends, and execution models +5. **Multiple concurrency models** - Support threads (default), Ractors (opt-in), and fibers (opt-in) + +### Non-Goals + +- Workflow definition DSL (covered in separate design) +- HTTP client implementation (already exists) +- Model serialization (already exists) + +--- + +## Architecture Overview + +### Component Hierarchy + +``` +┌─────────────────────────────────────────────────────────────────────┐ +│ User Code │ +│ (Worker classes, @worker_task methods, Worker.define blocks) │ +└─────────────────────────────────────────────────────────────────────┘ + │ + ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ TaskHandler │ +│ • Discovers workers (registry + auto-scan) │ +│ • Resolves configuration (3-tier hierarchy) │ +│ • Creates one Thread/Ractor per worker │ +│ • Manages lifecycle (start/stop/join) │ +│ • Aggregates events/metrics │ +└─────────────────────────────────────────────────────────────────────┘ + │ + ┌─────────────┼─────────────┐ + ▼ ▼ ▼ +┌─────────────────────┐ ┌─────────────────────┐ ┌─────────────────────┐ +│ TaskRunner │ │ TaskRunner │ │ RactorTaskRunner │ +│ (Thread-based) │ │ (Thread-based) │ │ (Ractor-based) │ +│ │ │ │ │ │ +│ • ThreadPoolExecutor│ │ • FiberExecutor │ │ • Ractor isolation │ +│ • Batch polling │ │ (async gem) │ │ • Own HTTP client │ +│ • Capacity mgmt │ │ • Batch polling │ │ • Message passing │ +│ • Event publishing │ │ • Capacity mgmt │ │ • Event publishing │ +└─────────────────────┘ └─────────────────────┘ └─────────────────────┘ + │ │ │ + ▼ ▼ ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ TaskResourceApi │ +│ • poll_task / batch_poll │ +│ • update_task │ +│ • HTTP communication via ApiClient │ +└─────────────────────────────────────────────────────────────────────┘ + │ + ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ Conductor Server │ +└─────────────────────────────────────────────────────────────────────┘ +``` + +### Comparison to Python SDK + +| Aspect | Python SDK | Ruby SDK | +|--------|-----------|----------| +| **Worker isolation** | One OS process per worker (`multiprocessing.Process`) | One Thread per worker (default), Ractor opt-in | +| **Task concurrency** | `ThreadPoolExecutor` (sync) or `asyncio` (async) | `Concurrent::ThreadPoolExecutor` (default), Fiber opt-in | +| **Why different** | Python GIL blocks ALL threads during CPU work | Ruby GVL releases during I/O (HTTP, sleep) | +| **Async model** | `asyncio` event loop, `async/await` | `async` gem with Fibers (opt-in) | +| **Event dispatch** | `SyncEventDispatcher` with `threading.Lock` | `SyncEventDispatcher` with `Mutex` | +| **Config resolution** | 3-tier: worker env → global env → code | Same | +| **HTTP client** | `requests` (sync), `httpx` (async) | `Faraday` with `net_http_persistent` | + +--- + +## Ruby vs Python Concurrency + +### Why Threads Work for Ruby Workers + +Python uses processes because the GIL (Global Interpreter Lock) prevents true thread parallelism for **any** Python code. Ruby's GVL (Global VM Lock) is similar but crucially different: + +**Ruby GVL releases during:** +- Network I/O (HTTP requests, socket operations) +- File I/O +- `sleep` calls +- C extension calls that release the GVL + +**Worker operations are I/O-bound:** +1. Poll HTTP endpoint (GVL released) +2. Execute worker (may be CPU-bound, but typically I/O) +3. Update HTTP endpoint (GVL released) + +This means Ruby threads provide **real concurrency** for typical worker workloads, making process-per-worker unnecessary overhead for most use cases. + +### When to Use Ractors + +Ractors provide true parallelism (no GVL sharing) but with restrictions: +- No shared mutable state between Ractors +- Limited gem compatibility (many gems use global state) +- Requires Ruby 3.1+ + +**Use Ractors when:** +- Worker performs CPU-intensive computation +- Worker doesn't need shared state +- All dependencies are Ractor-safe + +### When to Use Fibers + +Fibers provide lightweight cooperative concurrency within a single thread: +- ~400 bytes per fiber vs ~8KB per thread +- Can handle thousands of concurrent I/O operations +- Requires non-blocking I/O throughout + +**Use Fibers when:** +- Extremely high concurrency (hundreds of concurrent tasks) +- All operations are non-blocking (no blocking gem calls) +- Memory is constrained + +--- + +## Core Components + +### 1. TaskHandler + +The top-level orchestrator that manages all workers. + +```ruby +module Conductor + module Worker + class TaskHandler + # Initialize with optional workers and configuration + # @param workers [Array] Pre-created worker instances + # @param configuration [Configuration] Conductor configuration + # @param scan_for_annotated_workers [Boolean] Auto-discover @worker_task methods + # @param import_modules [Array] Ruby files/modules to require (triggers registration) + # @param event_listeners [Array] Custom event listeners + # @param metrics_settings [MetricsSettings] Metrics configuration + def initialize( + workers: nil, + configuration: nil, + scan_for_annotated_workers: true, + import_modules: nil, + event_listeners: nil, + metrics_settings: nil + ) + end + + # Start all worker threads/ractors + # @return [self] + def start + end + + # Stop all workers gracefully + # @param timeout [Integer] Seconds to wait before force-killing (default: 5) + # @return [self] + def stop(timeout: 5) + end + + # Wait for all workers to complete (blocking) + # @return [self] + def join + end + + # Check if handler is running + # @return [Boolean] + def running? + end + + # Get list of registered workers + # @return [Array] + def workers + end + end + end +end +``` + +**Responsibilities:** +1. Discover workers from registry + auto-scan +2. Resolve configuration for each worker (3-tier hierarchy) +3. Create appropriate runner (TaskRunner or RactorTaskRunner) based on config +4. Create one Thread (or Ractor) per worker +5. Manage lifecycle (start/stop/join) +6. Create shared EventDispatcher and register listeners +7. Optionally start MetricsProvider + +**Context Manager Pattern:** +```ruby +Conductor::Worker::TaskHandler.new(configuration: config) do |handler| + handler.start + handler.join +end +# Automatically calls stop on block exit +``` + +### 2. TaskRunner (Thread-based) + +The polling loop that runs in a dedicated Thread. + +```ruby +module Conductor + module Worker + class TaskRunner + # Initialize runner for a specific worker + # @param worker [Worker] The worker to run + # @param configuration [Configuration] Conductor configuration + # @param event_dispatcher [SyncEventDispatcher] Shared event dispatcher + # @param executor [Symbol] :thread_pool (default) or :fiber + def initialize(worker, configuration:, event_dispatcher:, executor: :thread_pool) + end + + # Main polling loop (runs until stopped) + def run + end + + # Single iteration of the polling loop + def run_once + end + + # Signal the runner to stop + def shutdown + end + + # Check if runner is running + # @return [Boolean] + def running? + end + end + end +end +``` + +**Internal State:** +```ruby +@worker # Worker instance +@configuration # Conductor configuration +@task_client # TaskClient for HTTP operations +@event_dispatcher # SyncEventDispatcher for publishing events +@executor # Concurrent::ThreadPoolExecutor or FiberExecutor +@running_tasks # Set of running futures/fibers +@consecutive_empty_polls # Counter for adaptive backoff +@auth_failures # Counter for auth failure backoff +@shutdown # AtomicBoolean for graceful shutdown +@last_poll_time # Time of last poll (for backoff calculation) +``` + +### 3. RactorTaskRunner + +The Ractor-based runner for CPU-bound workers requiring true parallelism. + +```ruby +module Conductor + module Worker + class RactorTaskRunner + # Initialize runner for a specific worker (runs inside Ractor) + # @param worker [Worker] The worker to run (must be Ractor-safe) + # @param configuration [Configuration] Conductor configuration (serializable parts only) + def initialize(worker, configuration:) + end + + # Main polling loop (creates HTTP client inside Ractor) + def run + end + + # Called by TaskHandler to receive events from Ractor + # @return [Array] Events from this poll cycle + def drain_events + end + end + end +end +``` + +**Key Differences from TaskRunner:** +1. Creates `TaskClient` **inside** `run()` (Ractors can't share objects) +2. Uses Ractor-local storage for TaskContext (not `Thread.current`) +3. Events are collected and sent to main Ractor via `Ractor.yield` for aggregation +4. No ThreadPoolExecutor - sequential execution within the Ractor (parallelism comes from multiple Ractors) + +### 4. Worker + +The user-facing worker definition that wraps an execute function. + +```ruby +module Conductor + module Worker + class Worker + attr_reader :task_definition_name, :execute_function, :config + attr_accessor :domain, :poll_interval, :thread_count, :worker_id, + :register_task_def, :overwrite_task_def, :strict_schema, + :paused, :poll_timeout, :isolation, :executor + + # Initialize a worker + # @param task_definition_name [String] Task type name in Conductor + # @param execute_function [Proc, Method] Function to execute tasks + # @param options [Hash] Worker configuration options + def initialize(task_definition_name, execute_function = nil, **options, &block) + end + + # Execute a task + # @param task [Task] The task to execute + # @return [TaskResult, TaskInProgress, Hash] Execution result + def execute(task) + end + + # Get polling interval in seconds + # @return [Float] + def polling_interval_seconds + end + + # Check if worker is async (for auto-detection, not used in Ruby) + # @return [Boolean] + def async? + end + end + end +end +``` + +**Execute Function Return Type Handling:** + +| Return Type | Behavior | +|-------------|----------| +| `TaskResult` | Use directly (set task_id, workflow_instance_id) | +| `TaskInProgress` | Create `IN_PROGRESS` result with `callback_after_seconds` | +| `Hash` | Wrap in `COMPLETED` TaskResult as output_data | +| `true` | `COMPLETED` with empty output | +| `false` | `FAILED` with empty output | +| `nil` | `COMPLETED` with empty output | +| Any other object | `COMPLETED` with `{ result: object }` output | +| Raises `NonRetryableError` | `FAILED_WITH_TERMINAL_ERROR` | +| Raises any `StandardError` | `FAILED` with error message | + +### 5. WorkerConfig + +Configuration resolver with 3-tier hierarchy. + +```ruby +module Conductor + module Worker + class WorkerConfig + # Resolve configuration for a worker + # @param worker_name [String] Task definition name + # @param defaults [Hash] Code-level defaults from worker definition + # @return [Hash] Resolved configuration + def self.resolve(worker_name, defaults = {}) + end + + # Configuration properties with types and defaults + PROPERTIES = { + poll_interval: { type: :integer, default: 100 }, # milliseconds + thread_count: { type: :integer, default: 1 }, + domain: { type: :string, default: nil }, + worker_id: { type: :string, default: -> { generate_worker_id } }, + poll_timeout: { type: :integer, default: 100 }, # milliseconds + register_task_def: { type: :boolean, default: false }, + overwrite_task_def: { type: :boolean, default: true }, + strict_schema: { type: :boolean, default: false }, + paused: { type: :boolean, default: false }, + isolation: { type: :symbol, default: :thread }, # :thread or :ractor + executor: { type: :symbol, default: :thread_pool } # :thread_pool or :fiber + }.freeze + end + end +end +``` + +**Resolution Priority (highest to lowest):** + +1. **Worker-specific environment variable:** + - `conductor.worker.{task_name}.{property}` (dotted) + - `CONDUCTOR_WORKER_{TASK_NAME}_{PROPERTY}` (uppercase) + +2. **Global worker environment variable:** + - `conductor.worker.all.{property}` (dotted) + - `CONDUCTOR_WORKER_ALL_{PROPERTY}` (uppercase) + +3. **Legacy environment variable:** + - `CONDUCTOR_WORKER_{PROPERTY}` (old format) + +4. **Code-level default:** + - Value passed to `worker_task` or `Worker.new` + +**Boolean Parsing:** Accepts `true/1/yes` and `false/0/no` (case-insensitive). + +--- + +## Three Runner Models + +### Model 1: TaskRunner with ThreadPoolExecutor (Default) + +``` +┌─────────────────────────────────────────────────────────────┐ +│ Worker Thread │ +│ ┌─────────────────────────────────────────────────────┐ │ +│ │ TaskRunner.run() │ │ +│ │ ┌────────────────────────────────────────────────┐ │ │ +│ │ │ Polling Loop │ │ │ +│ │ │ 1. Check capacity │ │ │ +│ │ │ 2. Adaptive backoff │ │ │ +│ │ │ 3. Batch poll │ │ │ +│ │ │ 4. Submit tasks to executor ─────────────────┐ │ │ │ +│ │ │ 5. Loop │ │ │ │ +│ │ └───────────────────────────────────────────────┘ │ │ │ +│ └──────────────────────────────────────────────────│─┘ │ +│ │ │ +│ ┌──────────────────────────────────────────────────▼─┐ │ +│ │ Concurrent::ThreadPoolExecutor │ │ +│ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │ +│ │ │ Thread 1│ │ Thread 2│ │ Thread 3│ │ Thread N│ │ │ +│ │ │ Task A │ │ Task B │ │ Task C │ │ idle │ │ │ +│ │ └─────────┘ └─────────┘ └─────────┘ └─────────┘ │ │ +│ │ (thread_count = N) │ │ +│ └────────────────────────────────────────────────────┘ │ +└─────────────────────────────────────────────────────────────┘ +``` + +**Configuration:** +```ruby +worker_task 'my_task', thread_count: 5, poll_interval: 100 +# or +Worker.new('my_task', thread_count: 5, executor: :thread_pool) +``` + +**Characteristics:** +- One dedicated thread for the polling loop +- ThreadPoolExecutor with `thread_count` threads for task execution +- GVL released during HTTP I/O, so threads provide real concurrency +- Best for: Most workloads (I/O-bound or mixed) + +### Model 2: TaskRunner with FiberExecutor (Opt-in) + +``` +┌─────────────────────────────────────────────────────────────┐ +│ Worker Thread │ +│ ┌─────────────────────────────────────────────────────┐ │ +│ │ TaskRunner.run() │ │ +│ │ ┌────────────────────────────────────────────────┐ │ │ +│ │ │ Async Event Loop (via async gem) │ │ │ +│ │ │ ┌───────┐ ┌───────┐ ┌───────┐ ┌───────┐ │ │ │ +│ │ │ │Fiber 1│ │Fiber 2│ │Fiber 3│ │Fiber N│ │ │ │ +│ │ │ │Task A │ │Task B │ │Task C │ │ poll │ │ │ │ +│ │ │ └───────┘ └───────┘ └───────┘ └───────┘ │ │ │ +│ │ │ (cooperative scheduling, single thread) │ │ │ +│ │ │ (thread_count = concurrency limit via Semaphore) │ │ +│ │ └────────────────────────────────────────────────┘ │ │ +│ └─────────────────────────────────────────────────────┘ │ +└─────────────────────────────────────────────────────────────┘ +``` + +**Configuration:** +```ruby +worker_task 'my_task', thread_count: 100, executor: :fiber +# Requires: gem 'async' in Gemfile +``` + +**Characteristics:** +- Single thread with fiber-based cooperative concurrency +- Requires `async` gem (optional dependency, loaded lazily) +- `thread_count` becomes fiber concurrency limit (semaphore) +- All I/O must be non-blocking (async gem provides non-blocking HTTP) +- Best for: Very high concurrency I/O-bound tasks (hundreds/thousands) + +**Lazy Loading:** +```ruby +# In fiber_executor.rb +def self.load_async_gem + require 'async' + require 'async/http' +rescue LoadError + raise Conductor::ConfigurationError, + "The 'async' gem is required for fiber executor. Add `gem 'async'` to your Gemfile." +end +``` + +### Model 3: RactorTaskRunner (Opt-in) + +``` +┌────────────────────────────────────────────────────────────────────────┐ +│ Main Thread (TaskHandler) │ +│ ┌──────────────────────────────────────────────────────────────────┐ │ +│ │ Event Aggregation Loop │ │ +│ │ • Receives events from Ractors via Ractor.receive │ │ +│ │ • Dispatches to shared EventDispatcher │ │ +│ └──────────────────────────────────────────────────────────────────┘ │ +│ ▲ ▲ ▲ │ +│ │ events │ events │ events │ +└───────────┼────────────────────┼────────────────────┼────────────────────┘ + │ │ │ +┌───────────┴──────┐ ┌─────────┴────────┐ ┌───────┴──────────┐ +│ Ractor 1 │ │ Ractor 2 │ │ Ractor 3 │ +│ ┌──────────────┐ │ │ ┌──────────────┐ │ │ ┌──────────────┐ │ +│ │RactorTaskRunner│ │ │ │RactorTaskRunner│ │ │ │RactorTaskRunner│ │ +│ │ Worker A │ │ │ │ Worker B │ │ │ │ Worker C │ │ +│ │ │ │ │ │ │ │ │ │ │ │ +│ │ Own HTTP │ │ │ │ Own HTTP │ │ │ │ Own HTTP │ │ +│ │ client │ │ │ │ client │ │ │ │ client │ │ +│ │ │ │ │ │ │ │ │ │ │ │ +│ │ Sequential │ │ │ │ Sequential │ │ │ │ Sequential │ │ +│ │ task exec │ │ │ │ task exec │ │ │ │ task exec │ │ +│ └──────────────┘ │ │ └──────────────┘ │ │ └──────────────┘ │ +│ (no GVL sharing) │ │ (no GVL sharing) │ │ (no GVL sharing) │ +└──────────────────┘ └──────────────────┘ └──────────────────┘ + True parallel execution across Ractors +``` + +**Configuration:** +```ruby +worker_task 'cpu_intensive_task', isolation: :ractor, thread_count: 4 +# Creates 4 Ractors, each running the same worker +``` + +**Characteristics:** +- True parallelism (each Ractor has its own GVL) +- HTTP client created inside Ractor (can't be shared) +- Events sent to main thread via Ractor messaging +- `thread_count` = number of Ractors (each processes one task at a time) +- Worker must be Ractor-safe (no shared mutable state) +- Requires Ruby 3.1+ +- Best for: CPU-intensive workers + +**Ractor Constraints:** +```ruby +# These will NOT work in Ractor-based workers: +- Global variables (@@, $) +- Class instance variables +- Mutating shared objects +- Many gems that use global state + +# These WILL work: +- Pure functions +- Immutable data +- Ractor-local state via Ractor.current[:key] +``` + +--- + +## Worker Definition Patterns + +### Pattern 1: Class-based with Module Mixin + +```ruby +class OrderProcessor + include Conductor::Worker::WorkerMixin + + worker_task 'process_order', + poll_interval: 200, + thread_count: 5, + domain: 'orders' + + def execute(task) + order_id = task.input_data['order_id'] + amount = task.input_data['amount'] + + # Process the order... + result = process(order_id, amount) + + # Return hash (auto-wrapped in COMPLETED TaskResult) + { processed: true, order_id: order_id, total: result.total } + end +end + +# Usage +handler = TaskHandler.new(workers: [OrderProcessor.new]) +``` + +### Pattern 2: Block-based with `Worker.define` + +```ruby +Conductor::Worker.define('send_notification', poll_interval: 100, thread_count: 3) do |task| + recipient = task.input_data['recipient'] + message = task.input_data['message'] + + NotificationService.send(to: recipient, body: message) + + { sent: true, recipient: recipient } +end + +# Workers registered automatically, discovered by TaskHandler +handler = TaskHandler.new(scan_for_annotated_workers: true) +``` + +### Pattern 3: Method Annotation with `worker_task` + +```ruby +module MyWorkers + extend Conductor::Worker::Annotatable + + worker_task 'greet_user', poll_interval: 50 + def self.greet(name:, greeting: 'Hello') + "#{greeting}, #{name}!" + end + + worker_task 'calculate_total', thread_count: 10 + def self.calculate(items:) + total = items.sum { |item| item['price'] * item['quantity'] } + { total: total, item_count: items.size } + end +end + +# Usage +handler = TaskHandler.new( + scan_for_annotated_workers: true, + import_modules: ['./lib/my_workers'] +) +``` + +### Pattern 4: Keyword Argument Mapping + +When the execute function has keyword arguments, they are automatically mapped from `task.input_data`: + +```ruby +# Worker definition with keyword args +worker_task 'process_payment' +def process_payment(order_id:, amount:, currency: 'USD') + # order_id, amount extracted from task.input_data + # currency uses default if not in input_data + PaymentGateway.charge(order_id, amount, currency) +end + +# Task input_data: { "order_id" => "123", "amount" => 99.99 } +# Mapped to: process_payment(order_id: "123", amount: 99.99, currency: 'USD') +``` + +### Pattern 5: Full Task Access + +```ruby +worker_task 'audit_task' +def audit(task) + # Full access to task object + puts "Task ID: #{task.task_id}" + puts "Workflow: #{task.workflow_instance_id}" + puts "Retry count: #{task.retry_count}" + puts "Poll count: #{task.poll_count}" + + # Access input + data = task.input_data + + # Return result + { audited: true } +end +``` + +--- + +## Task Context System + +TaskContext provides execution context accessible from anywhere in the worker code. + +### Thread-local Storage (TaskRunner) + +```ruby +module Conductor + module Worker + class TaskContext + # Get current context (thread-local) + # @return [TaskContext, nil] + def self.current + Thread.current[:conductor_task_context] + end + + # Set current context (internal use) + def self.current=(context) + Thread.current[:conductor_task_context] = context + end + + # Clear current context (internal use) + def self.clear + Thread.current[:conductor_task_context] = nil + end + + attr_reader :task, :task_result + + def initialize(task, task_result) + @task = task + @task_result = task_result + end + + # Convenience accessors + def task_id + @task.task_id + end + + def workflow_instance_id + @task.workflow_instance_id + end + + def retry_count + @task.retry_count || 0 + end + + def poll_count + @task.poll_count || 0 + end + + def input + @task.input_data || {} + end + + def task_def_name + @task.task_def_name + end + + # Mutable context methods + def add_log(message) + @task_result.log(message) + end + + def set_callback_after(seconds) + @task_result.callback_after_seconds = seconds + end + + def set_output(output_data) + @task_result.output_data = output_data + end + + def callback_after_seconds + @task_result.callback_after_seconds + end + end + end +end +``` + +### Fiber Storage (FiberExecutor) + +```ruby +# When using fiber executor, context stored in Fiber.current.storage +def self.current + if defined?(Fiber.current.storage) + Fiber.current.storage[:conductor_task_context] + else + Thread.current[:conductor_task_context] + end +end +``` + +### Ractor Storage (RactorTaskRunner) + +```ruby +# Ractors use Ractor.current for isolation +def self.current + Ractor.current[:conductor_task_context] +rescue + Thread.current[:conductor_task_context] +end +``` + +### Usage in Worker Code + +```ruby +worker_task 'my_task' +def execute(task) + ctx = Conductor::Worker::TaskContext.current + + # Log progress + ctx.add_log("Starting processing for #{ctx.task_id}") + + # Check retry count to avoid infinite loops + if ctx.retry_count > 3 + raise NonRetryableError, "Too many retries" + end + + # Long-running task - set callback + if will_take_long? + ctx.set_callback_after(60) # Check back in 60 seconds + return TaskInProgress.new(output: { status: 'processing' }) + end + + # Process... + result = do_work(ctx.input) + + ctx.add_log("Completed processing") + result +end +``` + +--- + +## Behavioral Algorithms + +### Algorithm 1: Main Polling Loop (`run_once`) + +```ruby +def run_once + # 1. Cleanup completed tasks (removes done futures from tracking set) + cleanup_completed_tasks + + # 2. Check capacity + current_capacity = @running_tasks.size + if current_capacity >= @max_workers + sleep(0.001) # 1ms sleep to prevent busy-waiting + return + end + + available_slots = @max_workers - current_capacity + + # 3. Adaptive backoff for empty polls + if @consecutive_empty_polls > 0 + backoff_ms = [1 * (2 ** [@consecutive_empty_polls, 10].min), @poll_interval].min + elapsed_ms = (Time.now - @last_poll_time) * 1000 + if elapsed_ms < backoff_ms + sleep((backoff_ms - elapsed_ms) / 1000.0) + return + end + end + + # 4. Batch poll for tasks + @last_poll_time = Time.now + tasks = batch_poll(available_slots) + + # 5. Submit tasks for execution + if tasks.empty? + @consecutive_empty_polls += 1 + else + @consecutive_empty_polls = 0 + tasks.each do |task| + future = @executor.post { execute_and_update(task) } + @running_tasks << future + end + end +end +``` + +### Algorithm 2: Batch Poll with Auth Backoff + +```ruby +def batch_poll(count) + # Skip if worker is paused + return [] if @worker.paused + + # Auth failure exponential backoff (capped at 60 seconds) + if @auth_failures > 0 + backoff_seconds = [2 ** @auth_failures, 60].min + elapsed = Time.now - @last_auth_failure_time + if elapsed < backoff_seconds + return [] + end + end + + # Publish PollStarted event + @event_dispatcher.publish(Events::PollStarted.new( + task_type: @worker.task_definition_name, + worker_id: @worker_id, + poll_count: @poll_count + )) + + start_time = Time.now + + begin + # HTTP batch poll + tasks = @task_client.batch_poll( + @worker.task_definition_name, + count: count, + timeout: @worker.poll_timeout, + worker_id: @worker_id, + domain: @worker.domain.presence # nil if empty string + ) + + duration_ms = (Time.now - start_time) * 1000 + @poll_count += 1 + + # Publish PollCompleted event + @event_dispatcher.publish(Events::PollCompleted.new( + task_type: @worker.task_definition_name, + duration_ms: duration_ms, + tasks_received: tasks.size + )) + + # Reset auth failures on success + @auth_failures = 0 + + tasks + rescue AuthorizationError => e + @auth_failures += 1 + @last_auth_failure_time = Time.now + duration_ms = (Time.now - start_time) * 1000 + + @event_dispatcher.publish(Events::PollFailure.new( + task_type: @worker.task_definition_name, + duration_ms: duration_ms, + cause: e + )) + + @logger.warn("Auth failure ##{@auth_failures}, backing off #{[2 ** @auth_failures, 60].min}s") + [] + rescue StandardError => e + duration_ms = (Time.now - start_time) * 1000 + + @event_dispatcher.publish(Events::PollFailure.new( + task_type: @worker.task_definition_name, + duration_ms: duration_ms, + cause: e + )) + + @logger.error("Poll failed: #{e.message}") + [] + end +end +``` + +### Algorithm 3: Task Execution + +```ruby +def execute_and_update(task) + task_result = execute_task(task) + + # Skip update for TaskInProgress (task stays in IN_PROGRESS state) + return if task_result.nil? || task_result.is_a?(TaskInProgress) + + update_task_with_retry(task_result) +end + +def execute_task(task) + # Create initial TaskResult for context + initial_result = TaskResult.new + initial_result.task_id = task.task_id + initial_result.workflow_instance_id = task.workflow_instance_id + initial_result.worker_id = @worker_id + + # Set task context (thread-local) + TaskContext.current = TaskContext.new(task, initial_result) + + start_time = Time.now + + # Publish TaskExecutionStarted + @event_dispatcher.publish(Events::TaskExecutionStarted.new( + task_type: @worker.task_definition_name, + task_id: task.task_id, + worker_id: @worker_id, + workflow_instance_id: task.workflow_instance_id + )) + + begin + # Execute worker + output = @worker.execute(task) + + duration_ms = (Time.now - start_time) * 1000 + + # Handle different return types + task_result = case output + when TaskResult + output + when TaskInProgress + result = TaskResult.in_progress + result.callback_after_seconds = output.callback_after_seconds + result.output_data = output.output + result + when Hash + result = TaskResult.complete + result.output_data = output + result + when true + TaskResult.complete + when false + TaskResult.failed('Worker returned false') + when nil + TaskResult.complete + else + result = TaskResult.complete + result.output_data = { 'result' => output } + result + end + + # Set IDs and merge context modifications + task_result.task_id = task.task_id + task_result.workflow_instance_id = task.workflow_instance_id + task_result.worker_id = @worker_id + + # Merge logs and callback_after from context + ctx = TaskContext.current + task_result.logs ||= [] + task_result.logs.concat(ctx.task_result.logs || []) + task_result.callback_after_seconds ||= ctx.callback_after_seconds + + output_size = task_result.output_data.to_json.bytesize rescue 0 + + # Publish TaskExecutionCompleted + @event_dispatcher.publish(Events::TaskExecutionCompleted.new( + task_type: @worker.task_definition_name, + task_id: task.task_id, + worker_id: @worker_id, + workflow_instance_id: task.workflow_instance_id, + duration_ms: duration_ms, + output_size_bytes: output_size + )) + + task_result + + rescue NonRetryableError => e + duration_ms = (Time.now - start_time) * 1000 + task_result = TaskResult.failed_with_terminal_error(e.message) + task_result.task_id = task.task_id + task_result.workflow_instance_id = task.workflow_instance_id + task_result.log("NonRetryableError: #{e.class}: #{e.message}") + + @event_dispatcher.publish(Events::TaskExecutionFailure.new( + task_type: @worker.task_definition_name, + task_id: task.task_id, + worker_id: @worker_id, + workflow_instance_id: task.workflow_instance_id, + duration_ms: duration_ms, + cause: e, + is_retryable: false + )) + + task_result + + rescue StandardError => e + duration_ms = (Time.now - start_time) * 1000 + task_result = TaskResult.failed(e.message) + task_result.task_id = task.task_id + task_result.workflow_instance_id = task.workflow_instance_id + task_result.log("Error: #{e.class}: #{e.message}\n#{e.backtrace.first(5).join("\n")}") + + @event_dispatcher.publish(Events::TaskExecutionFailure.new( + task_type: @worker.task_definition_name, + task_id: task.task_id, + worker_id: @worker_id, + workflow_instance_id: task.workflow_instance_id, + duration_ms: duration_ms, + cause: e, + is_retryable: true + )) + + task_result + + ensure + TaskContext.clear + end +end +``` + +### Algorithm 4: Task Update with Retry + +```ruby +RETRY_BACKOFFS = [0, 10, 20, 30].freeze # seconds + +def update_task_with_retry(task_result) + RETRY_BACKOFFS.each_with_index do |backoff, attempt| + sleep(backoff) if backoff > 0 + + start_time = Time.now + begin + @task_client.update_task(task_result) + duration_ms = (Time.now - start_time) * 1000 + + publish_task_update_completed(task_result, duration_ms) + return # Success + rescue StandardError => e + duration_ms = (Time.now - start_time) * 1000 + @logger.error("Task update failed (attempt #{attempt + 1}/#{RETRY_BACKOFFS.size}): #{e.message}") + + if attempt == RETRY_BACKOFFS.size - 1 + # All retries exhausted - CRITICAL: task result is lost + @logger.fatal("CRITICAL: Task update failed after #{RETRY_BACKOFFS.size} attempts. " \ + "Task #{task_result.task_id} result is LOST.") + publish_task_update_failure(task_result, e, duration_ms) + end + end + end +end +``` + +### Algorithm 5: Capacity Management + +The semaphore/capacity is held during BOTH execute AND update to prevent over-polling: + +```ruby +# In ThreadPoolExecutor model: +def execute_and_update(task) + # This entire method runs in a thread pool thread + # The "slot" is occupied from the moment the task is submitted + # until this method returns (after execute + update) + + task_result = execute_task(task) + return if task_result.nil? || task_result.is_a?(TaskInProgress) + + # Still occupying the slot during update retries + update_task_with_retry(task_result) + + # Slot released when this method returns and future completes +end + +# In FiberExecutor model: +def execute_and_update_fiber(task) + @semaphore.acquire # Acquire slot + + begin + task_result = execute_task(task) + return if task_result.nil? || task_result.is_a?(TaskInProgress) + + update_task_with_retry(task_result) + ensure + @semaphore.release # Release slot only after update + end +end +``` + +--- + +## Event System + +### Event Hierarchy + +```ruby +module Conductor + module Worker + module Events + # Base event with timestamp + class ConductorEvent + attr_reader :timestamp + + def initialize + @timestamp = Time.now.utc + end + end + + # Base for task runner events + class TaskRunnerEvent < ConductorEvent + attr_reader :task_type + + def initialize(task_type:) + super() + @task_type = task_type + end + end + + # Poll started + class PollStarted < TaskRunnerEvent + attr_reader :worker_id, :poll_count + + def initialize(task_type:, worker_id:, poll_count:) + super(task_type: task_type) + @worker_id = worker_id + @poll_count = poll_count + end + end + + # Poll completed successfully + class PollCompleted < TaskRunnerEvent + attr_reader :duration_ms, :tasks_received + + def initialize(task_type:, duration_ms:, tasks_received:) + super(task_type: task_type) + @duration_ms = duration_ms + @tasks_received = tasks_received + end + end + + # Poll failed + class PollFailure < TaskRunnerEvent + attr_reader :duration_ms, :cause + + def initialize(task_type:, duration_ms:, cause:) + super(task_type: task_type) + @duration_ms = duration_ms + @cause = cause + end + end + + # Task execution started + class TaskExecutionStarted < TaskRunnerEvent + attr_reader :task_id, :worker_id, :workflow_instance_id + + def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:) + super(task_type: task_type) + @task_id = task_id + @worker_id = worker_id + @workflow_instance_id = workflow_instance_id + end + end + + # Task execution completed + class TaskExecutionCompleted < TaskRunnerEvent + attr_reader :task_id, :worker_id, :workflow_instance_id, + :duration_ms, :output_size_bytes + + def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:, + duration_ms:, output_size_bytes:) + super(task_type: task_type) + @task_id = task_id + @worker_id = worker_id + @workflow_instance_id = workflow_instance_id + @duration_ms = duration_ms + @output_size_bytes = output_size_bytes + end + end + + # Task execution failed + class TaskExecutionFailure < TaskRunnerEvent + attr_reader :task_id, :worker_id, :workflow_instance_id, + :duration_ms, :cause, :is_retryable + + def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:, + duration_ms:, cause:, is_retryable: true) + super(task_type: task_type) + @task_id = task_id + @worker_id = worker_id + @workflow_instance_id = workflow_instance_id + @duration_ms = duration_ms + @cause = cause + @is_retryable = is_retryable + end + end + + # Task update failed (CRITICAL - result lost) + class TaskUpdateFailure < TaskRunnerEvent + attr_reader :task_id, :worker_id, :workflow_instance_id, + :cause, :retry_count, :task_result + + def initialize(task_type:, task_id:, worker_id:, workflow_instance_id:, + cause:, retry_count:, task_result:) + super(task_type: task_type) + @task_id = task_id + @worker_id = worker_id + @workflow_instance_id = workflow_instance_id + @cause = cause + @retry_count = retry_count + @task_result = task_result # For recovery + end + end + end + end +end +``` + +### SyncEventDispatcher + +```ruby +module Conductor + module Worker + class SyncEventDispatcher + def initialize + @listeners = Hash.new { |h, k| h[k] = [] } + @mutex = Mutex.new + end + + # Register a listener for an event type + # @param event_type [Class] Event class to listen for + # @param listener [Proc, #call] Callable to invoke + def register(event_type, listener) + @mutex.synchronize do + @listeners[event_type] << listener unless @listeners[event_type].include?(listener) + end + end + + # Unregister a listener + def unregister(event_type, listener) + @mutex.synchronize do + @listeners[event_type].delete(listener) + end + end + + # Publish an event to all registered listeners + # @param event [ConductorEvent] Event to publish + def publish(event) + listeners = @mutex.synchronize { @listeners[event.class].dup } + + listeners.each do |listener| + begin + listener.call(event) + rescue StandardError => e + # Listener failure is isolated - never breaks the worker + warn "Event listener error for #{event.class}: #{e.message}" + end + end + end + + # Check if there are listeners for an event type + def has_listeners?(event_type) + @mutex.synchronize { @listeners[event_type].any? } + end + + # Clear all listeners (for testing) + def clear + @mutex.synchronize { @listeners.clear } + end + end + end +end +``` + +### Listener Protocol (Duck Typing) + +```ruby +module Conductor + module Worker + # Listener protocol - implement any/all of these methods + # Methods are optional - only implemented methods are called + module TaskRunnerEventsListener + # @param event [PollStarted] + def on_poll_started(event); end + + # @param event [PollCompleted] + def on_poll_completed(event); end + + # @param event [PollFailure] + def on_poll_failure(event); end + + # @param event [TaskExecutionStarted] + def on_task_execution_started(event); end + + # @param event [TaskExecutionCompleted] + def on_task_execution_completed(event); end + + # @param event [TaskExecutionFailure] + def on_task_execution_failure(event); end + + # @param event [TaskUpdateFailure] + def on_task_update_failure(event); end + end + end +end +``` + +### Listener Registration Helper + +```ruby +module Conductor + module Worker + class ListenerRegistry + # Register a listener object with the dispatcher + # Auto-detects implemented methods via respond_to? + # @param listener [Object] Object implementing TaskRunnerEventsListener methods + # @param dispatcher [SyncEventDispatcher] Event dispatcher + def self.register_task_runner_listener(listener, dispatcher) + EVENT_METHOD_MAP.each do |event_class, method_name| + if listener.respond_to?(method_name) + dispatcher.register(event_class, ->(event) { listener.send(method_name, event) }) + end + end + end + + EVENT_METHOD_MAP = { + Events::PollStarted => :on_poll_started, + Events::PollCompleted => :on_poll_completed, + Events::PollFailure => :on_poll_failure, + Events::TaskExecutionStarted => :on_task_execution_started, + Events::TaskExecutionCompleted => :on_task_execution_completed, + Events::TaskExecutionFailure => :on_task_execution_failure, + Events::TaskUpdateFailure => :on_task_update_failure + }.freeze + end + end +end +``` + +### MetricsCollector + +`MetricsCollector.create` returns a collector that emits the canonical +(harmonized) metric surface: + +```ruby +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) +``` + +See [docs/METRICS_AND_INTERCEPTORS.md](../METRICS_AND_INTERCEPTORS.md) for the +full legacy and canonical metrics catalogs, label reference, and migration +guide. + +--- + +## Configuration System + +### Environment Variable Formats + +For a worker named `process_order` and property `poll_interval`: + +**Worker-specific (highest priority):** +```bash +# Dotted format (preferred) +conductor.worker.process_order.poll_interval=200 + +# Uppercase format +CONDUCTOR_WORKER_PROCESS_ORDER_POLL_INTERVAL=200 +``` + +**Global (applies to all workers):** +```bash +# Dotted format +conductor.worker.all.poll_interval=100 + +# Uppercase format +CONDUCTOR_WORKER_ALL_POLL_INTERVAL=100 +``` + +**Legacy (lowest priority):** +```bash +CONDUCTOR_WORKER_POLL_INTERVAL=100 +``` + +### Configuration Properties + +| Property | Type | Default | Env Var Suffix | Description | +|----------|------|---------|----------------|-------------| +| `poll_interval` | Integer | 100 | `POLL_INTERVAL` | Polling interval in milliseconds | +| `thread_count` | Integer | 1 | `THREAD_COUNT` | Max concurrent tasks (or Ractor count) | +| `domain` | String | nil | `DOMAIN` | Task domain for isolation | +| `worker_id` | String | auto | `WORKER_ID` | Unique worker identifier | +| `poll_timeout` | Integer | 100 | `POLL_TIMEOUT` | Server-side long poll timeout (ms) | +| `register_task_def` | Boolean | false | `REGISTER_TASK_DEF` | Auto-register task definition | +| `overwrite_task_def` | Boolean | true | `OVERWRITE_TASK_DEF` | Overwrite existing task defs | +| `strict_schema` | Boolean | false | `STRICT_SCHEMA` | Enforce strict JSON schema | +| `paused` | Boolean | false | `PAUSED` | Pause worker (stop polling) | +| `isolation` | Symbol | :thread | `ISOLATION` | `:thread` or `:ractor` | +| `executor` | Symbol | :thread_pool | `EXECUTOR` | `:thread_pool` or `:fiber` | + +### Auto-generated Worker ID + +```ruby +def self.generate_worker_id + hostname = Socket.gethostname rescue 'unknown' + pid = Process.pid + thread_id = Thread.current.object_id.to_s(16) + "#{hostname}-#{pid}-#{thread_id}" +end +``` + +--- + +## Task Definition Auto-Registration + +When `register_task_def: true`, the worker automatically registers its task definition on startup. + +### Registration Flow + +```ruby +def register_task_definition + return unless @worker.register_task_def + + # Build TaskDef from worker config or template + task_def = @worker.task_def_template&.dup || TaskDef.new + task_def.name = @worker.task_definition_name + + # Generate JSON schemas if possible + if @worker.execute_function.respond_to?(:parameters) + input_schema = generate_input_schema(@worker.execute_function) + output_schema = generate_output_schema(@worker.execute_function) + + register_schemas(input_schema, output_schema) if input_schema || output_schema + end + + # Register or update task definition + if @worker.overwrite_task_def + begin + @metadata_client.update_task_def(task_def) + rescue ApiError => e + if e.status == 404 + @metadata_client.register_task_def([task_def]) + else + raise + end + end + else + # Check if exists first + begin + @metadata_client.get_task_def(@worker.task_definition_name) + @logger.info("Task definition '#{@worker.task_definition_name}' already exists, skipping registration") + rescue ApiError => e + if e.status == 404 + @metadata_client.register_task_def([task_def]) + else + raise + end + end + end + + @logger.info("Registered task definition: #{@worker.task_definition_name}") +rescue StandardError => e + # Graceful degradation - worker still starts + @logger.warn("Failed to register task definition: #{e.message}") +end +``` + +### JSON Schema Generation + +```ruby +def generate_input_schema(func) + return nil unless func.respond_to?(:parameters) + + properties = {} + required = [] + + func.parameters.each do |type, name| + next if name == :task # Skip if taking full task object + + properties[name.to_s] = { 'type' => 'string' } # Default type + + case type + when :keyreq # Required keyword arg + required << name.to_s + when :key # Optional keyword arg + # Not required + end + end + + return nil if properties.empty? + + schema = { + '$schema' => 'http://json-schema.org/draft-07/schema#', + 'type' => 'object', + 'properties' => properties + } + schema['required'] = required unless required.empty? + schema['additionalProperties'] = !@worker.strict_schema + + schema +end +``` + +--- + +## File Structure + +``` +lib/conductor/ +├── worker/ +│ ├── worker.rb # Worker class +│ ├── worker_mixin.rb # WorkerMixin module for class-based workers +│ ├── worker_config.rb # Configuration resolver +│ ├── worker_registry.rb # Global worker registry +│ ├── task_handler.rb # Top-level orchestrator +│ ├── task_runner.rb # Thread-based runner +│ ├── ractor_task_runner.rb # Ractor-based runner (Phase 2) +│ ├── fiber_executor.rb # Fiber executor (Phase 3) +│ ├── task_context.rb # Execution context +│ ├── task_in_progress.rb # TaskInProgress return type +│ └── exceptions.rb # Worker-specific exceptions +├── worker/events/ +│ ├── conductor_event.rb # Base event class +│ ├── task_runner_events.rb # All 7 event classes +│ ├── sync_event_dispatcher.rb # Thread-safe event dispatcher +│ ├── listener_registry.rb # Listener registration helper +│ └── listeners.rb # Listener protocol module +├── worker/telemetry/ +│ ├── metrics_collector.rb # MetricsCollector class + NullBackend +│ └── prometheus_backend.rb # PrometheusBackend + MetricsServer +└── exceptions.rb # Add NonRetryableError +``` + +--- + +## Implementation Phases + +### Phase 1: Core Thread-based Runner (MVP) + +**Goal:** Production-ready thread-based workers with all Python SDK behavioral parity. + +**Components:** +1. `Worker` class with execute function handling +2. `WorkerMixin` module for class-based workers +3. `worker_task` DSL method and global registry +4. `WorkerConfig` with 3-tier resolution +5. `TaskRunner` with ThreadPoolExecutor +6. `TaskHandler` orchestrator +7. `TaskContext` (thread-local) +8. `TaskInProgress` return type +9. All 7 event classes +10. `SyncEventDispatcher` +11. `ListenerRegistry` +12. `MetricsCollector` with null backend + +**Algorithms:** +- Batch polling with dynamic sizing +- Adaptive backoff for empty polls +- Auth failure exponential backoff +- Task update with 4 retries +- Capacity management (semaphore during execute + update) + +**Tests:** +- Unit tests for each component +- Integration tests against local Conductor server +- All Python SDK test scenarios ported + +### Phase 2: Ractor-based Runner (Work-in-Progress) + +**Goal:** True parallelism for CPU-bound workers. + +**Status:** Partially implemented. The `RactorTaskRunner` can poll and execute +tasks inside Ractors, but the event bridge to the main thread is not yet +wired. This means metrics, interceptors, and custom event listeners receive +**no events** from Ractor workers. See +[Ractor Runner Limitations](../METRICS_AND_INTERCEPTORS.md#ractor-runner-limitations-work-in-progress----untested) +for details. + +**Components:** +1. `RactorTaskRunner` -- implemented, untested end-to-end +2. Ractor-local TaskContext storage -- implemented +3. Event aggregation via Ractor messaging -- **not yet implemented** +4. `isolation: :ractor` configuration -- implemented + +**Constraints:** +- Requires Ruby 3.1+ +- Worker must be Ractor-safe +- HTTP client created inside Ractor + +### Phase 3: Fiber Executor + +**Goal:** High-concurrency I/O-bound workers. + +**Components:** +1. `FiberExecutor` using `async` gem +2. Fiber-local TaskContext storage +3. `executor: :fiber` configuration +4. Async-compatible HTTP client + +**Constraints:** +- Requires `async` gem (optional dependency) +- All I/O must be non-blocking +- Worker must not use blocking operations + +### Phase 4: Metrics & Observability + +**Goal:** Production observability. + +**Components:** +1. `PrometheusBackend` with full metric set +2. HTTP metrics endpoint +3. Health check endpoint +4. Datadog backend (optional) + +### Phase 5: Advanced Features + +**Goal:** Feature parity with all Python SDK capabilities. + +**Components:** +1. Task definition auto-registration with JSON schemas +2. Worker auto-discovery (directory scanning) +3. Graceful shutdown with drain +4. Dynamic worker pause/resume + +--- + +## Appendix A: Complete Example + +```ruby +require 'conductor' + +# Configure +config = Conductor::Configuration.new( + server_api_url: 'http://localhost:8080/api' +) + +# Define workers + +# Class-based worker +class OrderProcessor + include Conductor::Worker::WorkerMixin + + worker_task 'process_order', + poll_interval: 200, + thread_count: 5 + + def execute(task) + order = task.input_data + ctx = Conductor::Worker::TaskContext.current + + ctx.add_log("Processing order #{order['id']}") + + # Simulate processing + result = process_order(order) + + ctx.add_log("Order processed successfully") + + { status: 'completed', total: result.total } + end + + private + + def process_order(order) + # Business logic here + OpenStruct.new(total: order['amount'] * 1.1) + end +end + +# Block-based worker +Conductor::Worker.define('send_notification', thread_count: 3) do |task| + recipient = task.input_data['recipient'] + message = task.input_data['message'] + + # Send notification + NotificationService.send(to: recipient, body: message) + + { sent: true } +end + +# Method-based worker with keyword args +module PaymentWorkers + extend Conductor::Worker::Annotatable + + worker_task 'charge_card', poll_interval: 100 + def self.charge_card(card_token:, amount:, currency: 'USD') + result = PaymentGateway.charge(card_token, amount, currency) + { transaction_id: result.id, status: result.status } + end +end + +# Custom event listener +class AuditLogger + def on_task_execution_completed(event) + puts "[AUDIT] Task #{event.task_id} completed in #{event.duration_ms}ms" + end + + def on_task_execution_failure(event) + puts "[AUDIT] Task #{event.task_id} FAILED: #{event.cause.message}" + end +end + +# Start workers +Conductor::Worker::TaskHandler.new( + configuration: config, + workers: [OrderProcessor.new], + scan_for_annotated_workers: true, + event_listeners: [AuditLogger.new] +) do |handler| + puts "Starting workers..." + handler.start + + # Handle shutdown gracefully + trap('INT') { handler.stop } + trap('TERM') { handler.stop } + + handler.join + puts "Workers stopped." +end +``` + +--- + +## Appendix B: Migration from Current Implementation + +The current `lib/conductor/worker/` has basic implementations that need to be replaced: + +| Current | New | +|---------|-----| +| `worker.rb` (WorkerModule) | `worker.rb` + `worker_mixin.rb` | +| `task_runner.rb` (simple) | `task_runner.rb` (full algorithms) | +| N/A | `task_handler.rb` | +| N/A | `worker_config.rb` | +| N/A | `task_context.rb` | +| N/A | `events/*` | +| N/A | `telemetry/*` | + +**Migration Strategy:** +1. Keep existing files during development +2. Build new implementation alongside +3. Update `lib/conductor.rb` to use new implementation +4. Remove old files after testing + +--- + +*Last Updated: February 2026* +*Status: Design Complete - Ready for Implementation* diff --git a/docs/design/WORKFLOW_DSL.md b/docs/design/WORKFLOW_DSL.md new file mode 100644 index 0000000..d8fb8e3 --- /dev/null +++ b/docs/design/WORKFLOW_DSL.md @@ -0,0 +1,561 @@ +# Conductor Ruby SDK - Workflow DSL Design + +## Table of Contents + +1. [Overview](#overview) +2. [Design Goals](#design-goals) +3. [Architecture](#architecture) +4. [Core Components](#core-components) +5. [DSL Syntax Reference](#dsl-syntax-reference) +6. [Reference Types](#reference-types) +7. [Control Flow](#control-flow) +8. [Task Types](#task-types) +9. [Implementation Details](#implementation-details) +10. [Examples](#examples) + +--- + +## Overview + +The Conductor Ruby SDK provides a clean, Ruby-idiomatic DSL for defining workflows. The DSL uses blocks, method chaining, and Ruby's dynamic features to create a natural syntax for workflow definition. + +### Basic Example + +```ruby +workflow = Conductor.workflow :order_processing, version: 1, executor: executor do + # Access workflow inputs with wf[:param] + user = simple :get_user, user_id: wf[:user_id] + + # Reference task outputs with task[:field] + order = simple :create_order, user_email: user[:email] + + # Parallel execution + parallel do + simple :ship_order, order_id: order[:id] + simple :send_confirmation, email: user[:email] + end + + # Set workflow output + output order_id: order[:id], status: 'completed' +end + +# Register and execute +workflow.register(overwrite: true) +result = workflow.execute(input: { user_id: 123 }, wait_for_seconds: 60) +``` + +--- + +## Design Goals + +1. **Ruby-idiomatic** - Use Ruby conventions (blocks, symbols, keyword arguments) +2. **Type-safe references** - `OutputRef` and `InputRef` for compile-time safety +3. **Minimal boilerplate** - Auto-generate reference names, convert types automatically +4. **Composable** - Nested blocks for control flow (`parallel`, `decide`, `loop_over`) +5. **Full feature coverage** - Support all Conductor task types including LLM tasks + +--- + +## Architecture + +``` +┌─────────────────────────────────────────────────────────────┐ +│ User Code │ +│ Conductor.workflow :name do ... end │ +└─────────────────────────────────────────────────────────────┘ + │ + ▼ instance_eval(&block) +┌─────────────────────────────────────────────────────────────┐ +│ WorkflowBuilder │ +│ • Holds workflow metadata (name, version, description) │ +│ • Contains task method implementations │ +│ • Collects TaskRefs during DSL evaluation │ +│ • Resolves OutputRef/InputRef to expression strings │ +└─────────────────────────────────────────────────────────────┘ + │ + ▼ returns WorkflowDefinition +┌─────────────────────────────────────────────────────────────┐ +│ WorkflowDefinition │ +│ • Wraps WorkflowBuilder │ +│ • Provides .register(), .execute(), .call() methods │ +│ • Delegates to WorkflowExecutor for execution │ +└─────────────────────────────────────────────────────────────┘ + │ + ▼ to_workflow_def +┌─────────────────────────────────────────────────────────────┐ +│ Conductor::Http::Models::WorkflowDef │ +│ • Serializable workflow definition │ +│ • Contains WorkflowTask array │ +│ • Ready for API submission │ +└─────────────────────────────────────────────────────────────┘ +``` + +--- + +## Core Components + +### 1. WorkflowBuilder (`lib/conductor/workflow/dsl/workflow_builder.rb`) + +The core DSL engine containing: + +- **Workflow metadata methods**: `description`, `timeout`, `owner_email`, `restartable`, `failure_workflow`, `output` +- **Task methods**: `simple`, `http`, `wait`, `terminate`, `sub_workflow`, etc. +- **Control flow methods**: `parallel`, `decide`, `when_true`, `when_false`, `loop_over`, `loop_while` +- **LLM methods**: `llm_chat`, `llm_embed`, `generate_image`, etc. +- **Value resolution**: `resolve_value`, `resolve_hash` for OutputRef/InputRef conversion + +**Key responsibility**: During `instance_eval`, task methods create `TaskRef` objects and return them, allowing chaining like `user[:email]`. + +### 2. WorkflowDefinition (`lib/conductor/workflow/dsl/workflow_definition.rb`) + +Wrapper class returned by `Conductor.workflow`: + +```ruby +class WorkflowDefinition + def initialize(builder, executor: nil) + @builder = builder + @executor = executor + end + + def register(overwrite: false) + # Register workflow via executor + end + + def execute(input: {}, wait_for_seconds: nil, ...) + # Execute workflow via executor + end + + def call(...) + execute(...) + end + + def to_workflow_def + @builder.to_workflow_def + end +end +``` + +### 3. TaskRef (`lib/conductor/workflow/dsl/task_ref.rb`) + +Stores task metadata during DSL evaluation: + +```ruby +class TaskRef + attr_reader :ref_name, :task_name, :task_type, :input_parameters, :options + + def [](field) + OutputRef.new(ref_name, field.to_s) + end + + def to_workflow_task + # Convert to Conductor::Http::Models::WorkflowTask + end +end +``` + +### 4. OutputRef (`lib/conductor/workflow/dsl/output_ref.rb`) + +Enables `task[:field]` syntax: + +```ruby +class OutputRef + def initialize(task_ref, path) + @task_ref = task_ref + @path = path + end + + def [](field) + OutputRef.new(@task_ref, "#{@path}.#{field}") + end + + def to_s + "${#{@task_ref}.output.#{@path}}" + end +end +``` + +### 5. InputRef (`lib/conductor/workflow/dsl/input_ref.rb`) + +Enables `wf[:param]` syntax: + +```ruby +class InputRef + def [](field) + InputFieldRef.new("workflow.input.#{field}") + end + + def var(name) + InputFieldRef.new("workflow.variables.#{name}") + end +end + +class InputFieldRef + def [](field) + InputFieldRef.new("#{@path}.#{field}") + end + + def to_s + "${#{@path}}" + end +end +``` + +### 6. Control Flow Builders + +**ParallelBuilder** - Collects branches for `parallel do...end`: + +```ruby +class ParallelBuilder + def initialize(parent_builder) + @parent = parent_builder + @branches = [[]] + end + + def method_missing(name, *args, **kwargs, &block) + # Delegate to parent builder, collect tasks into current branch + end + + def finalize + @branches # Array of task arrays + end +end +``` + +**SwitchBuilder** - Handles `decide expr do...end`: + +```ruby +class SwitchBuilder + def on(value, &block) + # Add case branch + end + + def otherwise(&block) + # Add default branch + end +end +``` + +--- + +## DSL Syntax Reference + +### Workflow Definition + +```ruby +Conductor.workflow :name, version: 1, executor: executor do + description 'Workflow description' + timeout 3600 # Timeout in seconds + owner_email 'owner@example.com' + restartable true + failure_workflow 'failure_handler' + + # ... tasks ... + + output key: value # Workflow output parameters +end +``` + +### Input/Output References + +```ruby +# Workflow inputs +wf[:user_id] # "${workflow.input.user_id}" +wf[:data][:items] # "${workflow.input.data.items}" +wf.var(:counter) # "${workflow.variables.counter}" + +# Task outputs +task[:result] # "${task_ref.output.result}" +task[:data][:nested][:field] # "${task_ref.output.data.nested.field}" +``` + +--- + +## Control Flow + +### Parallel Execution + +```ruby +parallel do + simple :task_a + simple :task_b + simple :task_c +end +``` + +Generates FORK_JOIN → tasks → JOIN structure. + +### Conditional Branching + +```ruby +decide user[:tier] do + on 'gold' do + simple :apply_gold_discount + end + on 'silver' do + simple :apply_silver_discount + end + otherwise do + simple :no_discount + end +end +``` + +### Conditional Shortcuts + +```ruby +when_true order[:is_premium] do + simple :apply_discount +end + +when_false order[:validated] do + terminate :failed, 'Validation failed' +end +``` + +### Loops + +```ruby +# Loop N times +loop_times 3 do + simple :process_batch +end + +# Loop with condition +loop_while '$.has_more == true' do + simple :fetch_page +end + +# Loop over items +loop_over users[:list] do + simple :process_user, user: iteration[:item] +end +``` + +--- + +## Task Types + +### Basic Tasks + +| Method | Type | Description | +|--------|------|-------------| +| `simple :name, **inputs` | SIMPLE | Worker task execution | +| `http :name, url:, method:, body:, headers:` | HTTP | HTTP request | +| `javascript :name, script:, **bindings` | INLINE | Inline JavaScript | +| `jq :name, query:, **inputs` | JSON_JQ_TRANSFORM | JQ transformation | +| `set var: value` | SET_VARIABLE | Set workflow variables | +| `human :name, assignee:, display_name:` | HUMAN | Human/manual task | + +### Wait and Events + +| Method | Type | Description | +|--------|------|-------------| +| `wait seconds` | WAIT | Wait for duration | +| `wait until_time: 'ISO8601'` | WAIT | Wait until time | +| `event :name, sink:, **payload` | EVENT | Publish event | +| `wait_for_webhook :name, matches: {}` | WAIT_FOR_WEBHOOK | Wait for callback | + +### Workflow Control + +| Method | Type | Description | +|--------|------|-------------| +| `terminate :status, 'reason'` | TERMINATE | End workflow | +| `sub_workflow :name, workflow:, version:` | SUB_WORKFLOW | Call workflow | +| `start_workflow :name, workflow:, **inputs` | START_WORKFLOW | Fire-and-forget | +| `inline_workflow :name do...end` | SUB_WORKFLOW | Inline sub-workflow | + +### Dynamic Tasks + +| Method | Type | Description | +|--------|------|-------------| +| `dynamic :name, dynamic_task_param:` | DYNAMIC | Runtime task name | +| `dynamic_fork :name, tasks_param:, tasks_input_param:` | FORK_JOIN_DYNAMIC | Dynamic parallel | +| `http_poll :name, url:, termination_condition:` | HTTP_POLL | Poll until condition | + +### LLM/AI Tasks + +| Method | Type | Description | +|--------|------|-------------| +| `llm_chat :name, provider:, model:, messages:` | LLM_CHAT_COMPLETE | Chat completion | +| `llm_complete :name, provider:, model:, prompt:` | LLM_TEXT_COMPLETE | Text completion | +| `llm_embed :name, provider:, model:, text:` | LLM_GENERATE_EMBEDDINGS | Generate embeddings | +| `llm_store_embeddings :name, vector_db:, index:, embeddings:` | LLM_STORE_EMBEDDINGS | Store in vector DB | +| `llm_search_embeddings :name, vector_db:, index:, embeddings:` | LLM_SEARCH_EMBEDDINGS | Search vector DB | +| `generate_image :name, provider:, model:, prompt:` | GENERATE_IMAGE | Image generation | +| `generate_audio :name, provider:, model:, text:, voice:` | GENERATE_AUDIO | Text-to-speech | +| `list_mcp_tools :name, mcp_server:` | LIST_MCP_TOOLS | List MCP tools | +| `call_mcp_tool :name, mcp_server:, method:, arguments:` | CALL_MCP_TOOL | Call MCP tool | +| `get_document :name, url:, media_type:` | GET_DOCUMENT | Retrieve document | + +--- + +## Implementation Details + +### Reference Resolution + +The DSL automatically converts references to Conductor expression strings: + +```ruby +def resolve_value(value) + case value + when OutputRef, InputRef + value.to_s # "${task_ref.output.field}" + when Hash + resolve_hash(value) # Recursively resolve + when Array + value.map { |v| resolve_value(v) } + else + value # Literals pass through + end +end +``` + +### Task Reference Name Generation + +```ruby +def generate_ref_name(task_name) + base = "#{task_name}_ref" + @ref_counter[base] += 1 + @ref_counter[base] == 1 ? base : "#{base}_#{@ref_counter[base]}" +end +``` + +This ensures unique reference names: +- First `simple :get_user` → `get_user_ref` +- Second `simple :get_user` → `get_user_ref_2` + +### Workflow Task Conversion + +`TaskRef#to_workflow_task` converts DSL representation to `WorkflowTask` model: + +```ruby +def to_workflow_task + wf_task = Conductor::Http::Models::WorkflowTask.new( + name: @task_name, + task_reference_name: @ref_name, + type: @task_type, + input_parameters: @input_parameters + ) + + # Handle special task types (FORK_JOIN, SWITCH, DO_WHILE, etc.) + case @task_type + when TaskType::FORK_JOIN + wf_task.fork_tasks = convert_branches(@options[:fork_branches]) + when TaskType::SWITCH + wf_task.expression = @options[:expression] + wf_task.decision_cases = @options[:decision_cases] + wf_task.default_case = @options[:default_case] + # ... etc + end + + wf_task +end +``` + +--- + +## Examples + +### E-commerce Order Processing + +```ruby +workflow = Conductor.workflow :order_processing, version: 1, executor: executor do + description 'Process customer orders' + timeout 3600 + + # Validate order + validation = simple :validate_order, + order_id: wf[:order_id], + customer_id: wf[:customer_id] + + # Check inventory + inventory = simple :check_inventory, + items: validation[:items] + + # Conditional based on stock + decide inventory[:in_stock] do + on 'true' do + # Process payment + payment = simple :process_payment, + amount: validation[:total], + customer_id: wf[:customer_id] + + # Parallel fulfillment + parallel do + simple :ship_order, order_id: wf[:order_id] + simple :send_confirmation, email: validation[:customer_email] + simple :update_analytics, order_data: validation[:result] + end + end + + otherwise do + simple :notify_backorder, items: inventory[:missing_items] + terminate :failed, 'Items out of stock' + end + end + + output order_id: wf[:order_id], status: 'completed' +end +``` + +### AI Document Processing + +```ruby +workflow = Conductor.workflow :document_analysis, version: 1, executor: executor do + # Fetch document + doc = get_document :fetch_doc, + url: wf[:document_url], + media_type: wf[:media_type] + + # Generate embeddings + embeddings = llm_embed :embed_doc, + provider: 'openai', + model: 'text-embedding-3-small', + text: doc[:content] + + # Store in vector database + llm_store_embeddings :store_vectors, + vector_db: 'pinecone', + index: 'documents', + embeddings: embeddings[:embeddings], + id: wf[:doc_id], + metadata: { source: wf[:document_url] } + + # Analyze with LLM + analysis = llm_chat :analyze, + provider: 'openai', + model: 'gpt-4', + messages: [ + { role: :system, message: 'Analyze the following document and extract key insights.' }, + { role: :user, message: doc[:content] } + ], + temperature: 0.3 + + output analysis: analysis[:content], doc_id: wf[:doc_id] +end +``` + +--- + +## Testing + +DSL tests are in `spec/conductor/workflow/dsl/workflow_builder_spec.rb`: + +```bash +bundle exec rspec spec/conductor/workflow/dsl/ --format documentation +``` + +Current coverage: 52 test cases covering: +- All basic task methods +- All LLM task methods +- Control flow (parallel, decide, loops) +- Reference resolution (OutputRef, InputRef) +- WorkflowDef conversion + +--- + +## Related Documents + +- [DESIGN.md](../../DESIGN.md) - High-level architecture +- [WORKER_DESIGN.md](WORKER_DESIGN.md) - Worker infrastructure +- [README.md](../../README.md) - User documentation diff --git a/harness/README.md b/harness/README.md new file mode 100644 index 0000000..b50a56f --- /dev/null +++ b/harness/README.md @@ -0,0 +1,50 @@ +# Ruby SDK Docker Harness + +A long-running worker harness built from the root `Dockerfile`. + +## Worker Harness + +A self-feeding worker that runs indefinitely. On startup it registers five simulated tasks (`ruby_worker_0` through `ruby_worker_4`) and the `ruby_simulated_tasks_workflow`, then runs two background services: + +- **WorkflowGovernor** -- starts a configurable number of `ruby_simulated_tasks_workflow` instances per second (default 2), indefinitely. +- **SimulatedTaskWorkers** -- five task handlers, each with a codename and a default sleep duration. Each worker supports configurable delay types, failure simulation, and output generation via task input parameters. The workflow chains them in sequence: quickpulse (1s) → whisperlink (2s) → shadowfetch (3s) → ironforge (4s) → deepcrawl (5s). + +```bash +docker build --target harness -t ruby-sdk-harness . + +docker run -d \ + -e CONDUCTOR_SERVER_URL=https://your-cluster.example.com/api \ + -e CONDUCTOR_AUTH_KEY=$CONDUCTOR_AUTH_KEY \ + -e CONDUCTOR_AUTH_SECRET=$CONDUCTOR_AUTH_SECRET \ + -e HARNESS_WORKFLOWS_PER_SEC=4 \ + ruby-sdk-harness +``` + +You can also run the harness locally without Docker: + +```bash +export CONDUCTOR_SERVER_URL=https://your-cluster.example.com/api +export CONDUCTOR_AUTH_KEY=$CONDUCTOR_AUTH_KEY +export CONDUCTOR_AUTH_SECRET=$CONDUCTOR_AUTH_SECRET + +ruby harness/main.rb +``` + +Override defaults with environment variables as needed: + +```bash +HARNESS_WORKFLOWS_PER_SEC=4 HARNESS_BATCH_SIZE=10 ruby harness/main.rb +``` + +All resource names use a `ruby_` prefix so multiple SDK harnesses (Python, Java, Go, C#, etc.) can coexist on the same cluster. + +### Environment Variables + +| Variable | Required | Default | Description | +|---|---|---|---| +| `CONDUCTOR_SERVER_URL` | yes | -- | Conductor API base URL | +| `CONDUCTOR_AUTH_KEY` | no | -- | Orkes auth key | +| `CONDUCTOR_AUTH_SECRET` | no | -- | Orkes auth secret | +| `HARNESS_WORKFLOWS_PER_SEC` | no | 2 | Workflows to start per second | +| `HARNESS_BATCH_SIZE` | no | 20 | Number of tasks each worker polls per batch | +| `HARNESS_POLL_INTERVAL_MS` | no | 100 | Milliseconds between poll cycles | diff --git a/harness/manifests/README.md b/harness/manifests/README.md new file mode 100644 index 0000000..3b3d943 --- /dev/null +++ b/harness/manifests/README.md @@ -0,0 +1,132 @@ +# Kubernetes Manifests + +This directory contains Kubernetes manifests for deploying the Ruby SDK harness worker to the certification clusters. + +## Prerequisites + +**Set your namespace environment variable:** +```bash +export NS=your-namespace-here +``` + +All kubectl commands below use `-n $NS` to specify the namespace. The manifests intentionally do not include hardcoded namespaces. + +**Note:** The harness worker images are published as public packages on GHCR and do not require authentication to pull. No image pull secrets are needed. + +## Files + +| File | Description | +|---|---| +| `deployment.yaml` | Deployment (single file, works on all clusters) | +| `configmap-aws.yaml` | Conductor URL + auth key for certification-aws | +| `configmap-azure.yaml` | Conductor URL + auth key for certification-az | +| `configmap-gcp.yaml` | Conductor URL + auth key for certification-gcp | +| `secret-conductor.yaml` | Conductor auth secret (placeholder template) | + +## Quick Start + +### 1. Create the Conductor Auth Secret + +The `CONDUCTOR_AUTH_SECRET` must be created as a Kubernetes secret before deploying. + +```bash +kubectl create secret generic conductor-credentials \ + --from-literal=auth-secret=YOUR_AUTH_SECRET \ + -n $NS +``` + +If the `conductor-credentials` secret already exists in the namespace (e.g. from the e2e-testrunner-worker), it can be reused as-is. + +See `secret-conductor.yaml` for more details. + +### 2. Apply the ConfigMap for Your Cluster + +```bash +# AWS +kubectl apply -f manifests/configmap-aws.yaml -n $NS + +# Azure +kubectl apply -f manifests/configmap-azure.yaml -n $NS + +# GCP +kubectl apply -f manifests/configmap-gcp.yaml -n $NS +``` + +### 3. Deploy + +```bash +kubectl apply -f manifests/deployment.yaml -n $NS +``` + +### 4. Verify + +```bash +# Check pod status +kubectl get pods -n $NS -l app=ruby-sdk-harness-worker + +# Watch logs +kubectl logs -n $NS -l app=ruby-sdk-harness-worker -f +``` + +## Building and Pushing the Image + +From the repository root: + +```bash +# Build the harness target and push to GHCR +docker buildx build \ + --platform linux/amd64,linux/arm64 \ + --target harness \ + -t ghcr.io/conductor-oss/ruby-sdk/harness-worker:latest \ + --push . +``` + +After pushing a new image with the same tag, restart the deployment to pull it: + +```bash +kubectl rollout restart deployment/ruby-sdk-harness-worker -n $NS +kubectl rollout status deployment/ruby-sdk-harness-worker -n $NS +``` + +## Tuning + +The harness worker accepts these optional environment variables (set in `deployment.yaml`): + +| Variable | Default | Description | +|---|---|---| +| `HARNESS_WORKFLOWS_PER_SEC` | 2 | Workflows to start per second | +| `HARNESS_BATCH_SIZE` | 20 | Tasks each worker polls per batch | +| `HARNESS_POLL_INTERVAL_MS` | 100 | Milliseconds between poll cycles | + +Edit `deployment.yaml` to change these, then re-apply: + +```bash +kubectl apply -f manifests/deployment.yaml -n $NS +``` + +## Troubleshooting + +### Pod not starting + +```bash +kubectl describe pod -n $NS -l app=ruby-sdk-harness-worker +kubectl logs -n $NS -l app=ruby-sdk-harness-worker --tail=100 +``` + +### Secret not found + +```bash +kubectl get secret conductor-credentials -n $NS +``` + +## Resource Limits + +Default resource allocation: +- **Memory**: 256Mi (request) / 512Mi (limit) +- **CPU**: 100m (request) / 500m (limit) + +Adjust in `deployment.yaml` based on workload. Higher `HARNESS_WORKFLOWS_PER_SEC` values may need more CPU/memory. + +## Service + +The harness worker does **not** need a Service or Ingress. It connects to Conductor via outbound HTTP polling. All communication is outbound. From e1cc9cffc32ad1a0158687514ab7d29ab97fd54b Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 09:51:25 -0700 Subject: [PATCH 12/20] Make secret integration tests handle read-only OSS backends --- spec/integration/orkes_spec.rb | 77 ++++++++++++++++------------------ 1 file changed, 37 insertions(+), 40 deletions(-) diff --git a/spec/integration/orkes_spec.rb b/spec/integration/orkes_spec.rb index 570c872..7a6e4d8 100644 --- a/spec/integration/orkes_spec.rb +++ b/spec/integration/orkes_spec.rb @@ -57,13 +57,7 @@ def skip_if_limit_reached(error) let(:secret_key) { "#{test_id}_secret" } let(:secret_value) { "test_secret_value_#{SecureRandom.hex(8)}" } - # OSS Conductor registers a full secrets CRUD controller by default (the - # `agentspan` module's `conductor.integrations.ai.enabled=true` default), - # but only ships read-only SecretsDAO backends: writes (put/delete) return - # a real 501 "read-only backend" rather than succeeding. Reads work - # against an env-backed secret seeded via - # CONDUCTOR_SECRET_RUBY_SDK_INTEGRATION_TEST in scripts/docker-compose-oss.yaml - # -- keep these two constants in sync with that file. + # The Docker integration stack supplies this fixture for the read-only env backend. OSS_SEEDED_SECRET_NAME = 'RUBY_SDK_INTEGRATION_TEST' OSS_SEEDED_SECRET_VALUE = 'ruby-sdk-oss-secret-value' @@ -76,46 +70,49 @@ def skip_if_limit_reached(error) end it 'performs CRUD operations on secrets' do - if IntegrationHelper.oss? - # Verify reads work against the pre-seeded env-backed secret, and that - # writes fail with a real 501 (read-only backend) rather than silently - # succeeding or failing for some other reason. - expect(secret_client.get_secret(OSS_SEEDED_SECRET_NAME)).to eq(OSS_SEEDED_SECRET_VALUE) - expect(secret_client.secret_exists(OSS_SEEDED_SECRET_NAME)).to be true - expect(secret_client.list_all_secret_names).to include(OSS_SEEDED_SECRET_NAME) - - begin - secret_client.put_secret(secret_key, secret_value) - # A future OSS release might ship a writable backend; if so, clean up. - secret_client.delete_secret(secret_key) - rescue Conductor::ApiError => e - raise unless e.status == 501 - end - else - # Create + begin secret_client.put_secret(secret_key, secret_value) + rescue Conductor::ApiError => e + raise unless IntegrationHelper.oss? && e.status == 501 - # Verify it exists - exists = secret_client.secret_exists(secret_key) - expect(exists).to be true + skip 'Secret creation is unsupported by the server’s read-only secrets backend' + end - # List secrets should include our key - secrets = secret_client.list_all_secret_names - expect(secrets).to include(secret_key) + expect(secret_client.secret_exists(secret_key)).to be true + expect(secret_client.list_all_secret_names).to include(secret_key) + # Orkes may mask the value depending on permissions. + expect(secret_client.get_secret(secret_key)).not_to be_nil - # Get secret (note: Orkes may return masked value or the actual value depending on permissions) - retrieved = secret_client.get_secret(secret_key) - expect(retrieved).not_to be_nil + secret_client.delete_secret(secret_key) + expect(secret_client.secret_exists(secret_key)).to be false + rescue Conductor::ApiError => e + skip_if_limit_reached(e) + end - # Delete - secret_client.delete_secret(secret_key) + it 'reports unsupported mutations and missing secrets on a read-only OSS backend' do + skip 'Only applies to OSS secrets backends' unless IntegrationHelper.oss? - # Verify deleted - exists_after = secret_client.secret_exists(secret_key) - expect(exists_after).to be false + begin + secret_client.put_secret(secret_key, secret_value) + skip 'The configured secrets backend supports writes; covered by the CRUD test' + rescue Conductor::ApiError => e + expect(e.status).to eq(501) end - rescue Conductor::ApiError => e - skip_if_limit_reached(e) + + expect(secret_client.secret_exists(secret_key)).to be false + expect(secret_client.list_all_secret_names).not_to include(secret_key) + expect { secret_client.get_secret(secret_key) }.to raise_error(Conductor::ApiError) { |error| expect(error.status).to eq(404) } + expect { secret_client.delete_secret(secret_key) }.to raise_error(Conductor::ApiError) { |error| expect(error.status).to eq(501) } + end + + it 'reads the environment-backed secret supplied by the OSS integration stack' do + skip 'Only applies to the OSS integration stack' unless IntegrationHelper.oss? + unless secret_client.secret_exists(OSS_SEEDED_SECRET_NAME) + skip 'Server fixture is absent; scripts/run-integration-oss.sh seeds it when starting Conductor' + end + + expect(secret_client.get_secret(OSS_SEEDED_SECRET_NAME)).to eq(OSS_SEEDED_SECRET_VALUE) + expect(secret_client.list_all_secret_names).to include(OSS_SEEDED_SECRET_NAME) end it 'handles secret tags' do From f62c6a2e46b9e5cce46720df23493eebe3acccb6 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 13:18:34 -0700 Subject: [PATCH 13/20] Reuse shared Conductor startup for SDK agent playback --- .github/scripts/start-agent-playback.sh | 42 ------------- .github/scripts/start-agent-services.sh | 70 +++++++++++++++++++++ .github/workflows/agents-playback.yml | 29 ++++----- CONTRIBUTING.md | 14 +++-- scripts/run-agents-playback.sh | 81 +++++++++++++++++++++++++ 5 files changed, 176 insertions(+), 60 deletions(-) delete mode 100644 .github/scripts/start-agent-playback.sh create mode 100755 .github/scripts/start-agent-services.sh create mode 100755 scripts/run-agents-playback.sh diff --git a/.github/scripts/start-agent-playback.sh b/.github/scripts/start-agent-playback.sh deleted file mode 100644 index fba3d0d..0000000 --- a/.github/scripts/start-agent-playback.sh +++ /dev/null @@ -1,42 +0,0 @@ -#!/usr/bin/env bash -# Start only a dedicated test server. The checkout, action and recordings must -# all use the same Conductor feature branch from agents-playback.yml. -set -euo pipefail -conductor_dir=$(cd "${1:?Usage: start-agent-playback.sh CONDUCTOR_CHECKOUT}" && pwd) -playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-"$PWD/tmp/agent-playback"} -mkdir -p "$playback_dir" -playback_dir=$(cd "$playback_dir" && pwd) -if [[ -e "$playback_dir/server.pid" || -e "$playback_dir/playback.db" ]]; then - echo "Use a fresh CONDUCTOR_PLAYBACK_WORK_DIR for each playback run" >&2 - exit 1 -fi -export CONDUCTOR_RECORDINGS_DIR="$conductor_dir/llm-recordings" -export CONDUCTOR_SECRET_GITHUB_TOKEN=playback-test-key -export CONDUCTOR_SECRET_HTTP_TEST_API_KEY=playback-test-key -export CONDUCTOR_SECRET_MCP_TEST_API_KEY=playback-test-key -python3 spec/support/agents/http_fixture.py > "$playback_dir/http.log" 2>&1 & -echo $! > "$playback_dir/http.pid" -mcp-testkit --transport http --auth playback-test-key > "$playback_dir/mcp.log" 2>&1 & -echo $! > "$playback_dir/mcp.pid" -java -Xmx2g -jar "$conductor_dir"/server/build/libs/*-boot.jar \ - --server.port=18080 \ - --spring.datasource.url="jdbc:sqlite:$playback_dir/playback.db" \ - --conductor.ai.enable-llm-mocks=true \ - --conductor.ai.recordings-directory="$CONDUCTOR_RECORDINGS_DIR" \ - --conductor.ai.outbound.allowed-origins=http://localhost:3001,http://localhost:3002 \ - --conductor.ai.outbound.allow-private-networks=true \ - > "$playback_dir/server.log" 2>&1 & -echo $! > "$playback_dir/server.pid" -for attempt in $(seq 1 90); do - if curl -fsS http://localhost:18080/health > /dev/null 2>&1; then - exit 0 - fi - if ! kill -0 "$(cat "$playback_dir/server.pid")" 2>/dev/null; then - cat "$playback_dir/server.log" - exit 1 - fi - sleep 2 -done -cat "$playback_dir/server.log" -echo 'Conductor did not become ready' >&2 -exit 1 diff --git a/.github/scripts/start-agent-services.sh b/.github/scripts/start-agent-services.sh new file mode 100755 index 0000000..f778d94 --- /dev/null +++ b/.github/scripts/start-agent-services.sh @@ -0,0 +1,70 @@ +#!/usr/bin/env bash +# Start the authenticated HTTP/MCP fixtures; --check only checks existing services. +set -euo pipefail +repo_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd) +conductor_dir=$(cd "${1:?Usage: start-agent-services.sh CONDUCTOR_CHECKOUT [--check]}" && pwd) +cd "$repo_dir" +export CONDUCTOR_RECORDINGS_DIR="$conductor_dir/llm-recordings" +playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-"$repo_dir/tmp/agent-playback"} + +check_services() { + curl --fail --silent --show-error --max-time 3 \ + -H 'Authorization: Bearer playback-test-key' -H 'Content-Type: application/json' \ + --data '{"text":"hello world"}' http://localhost:3001/api/string/reverse > /dev/null && + curl --fail --silent --show-error --max-time 3 \ + -H 'Authorization: Bearer playback-test-key' -H 'Content-Type: application/json' \ + -H 'Accept: application/json, text/event-stream' \ + --data '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"ruby-sdk-playback","version":"1"}}}' \ + http://localhost:3001/mcp > /dev/null && + curl --fail --silent --show-error --max-time 3 \ + -H 'Authorization: Bearer playback-test-key' \ + 'http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated' > /dev/null +} + +if [[ "${2:-}" == --check ]]; then + check_services + exit +fi +if [[ -n "${2:-}" ]]; then + echo "Unknown option: $2" >&2 + exit 2 +fi +command -v mcp-testkit > /dev/null +# Never replace or stop an existing service. +python3 - <<'PY' +import socket +for port in (3001, 3002): + with socket.socket() as sock: + sock.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) + sock.bind(('127.0.0.1', port)) +PY +mkdir -p "$playback_dir" +pids=() +cleanup_on_error() { + result=$? + if ((result != 0)); then + for pid in "${pids[@]}"; do kill "$pid" 2>/dev/null || true; done + fi +} +trap cleanup_on_error EXIT +nohup python3 spec/support/agents/http_fixture.py < /dev/null > "$playback_dir/http.log" 2>&1 & +pids+=("$!") +echo "$!" > "$playback_dir/http.pid" +nohup mcp-testkit --transport http --host 127.0.0.1 --port 3001 --auth playback-test-key < /dev/null > "$playback_dir/mcp.log" 2>&1 & +pids+=("$!") +echo "$!" > "$playback_dir/mcp.pid" +for attempt in $(seq 1 30); do + for pid in "${pids[@]}"; do + if ! kill -0 "$pid" 2>/dev/null; then + cat "$playback_dir/http.log" "$playback_dir/mcp.log" + exit 1 + fi + done + if check_services > "$playback_dir/services-check.log" 2>&1; then + echo 'Authenticated HTTP and MCP fixtures are ready on ports 3001 and 3002.' + exit 0 + fi + sleep 1 +done +cat "$playback_dir/services-check.log" +exit 1 diff --git a/.github/workflows/agents-playback.yml b/.github/workflows/agents-playback.yml index 809261c..f20916c 100644 --- a/.github/workflows/agents-playback.yml +++ b/.github/workflows/agents-playback.yml @@ -16,6 +16,10 @@ jobs: CONDUCTOR_SERVER_URL: http://localhost:18080/api CONDUCTOR_AGENT_LLM_MODEL: mock/mockLLM CONDUCTOR_AGENTS_PLAYBACK: 'true' + CONDUCTOR_PLAYBACK_WORK_DIR: tmp/agent-playback + CONDUCTOR_SECRET_GITHUB_TOKEN: playback-test-key + CONDUCTOR_SECRET_HTTP_TEST_API_KEY: playback-test-key + CONDUCTOR_SECRET_MCP_TEST_API_KEY: playback-test-key GITHUB_REPOS_URL: http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated steps: - uses: actions/checkout@v4 @@ -42,26 +46,23 @@ jobs: python-version: '3.12' - name: Install MCP test service run: pip install mcp-testkit==1.0.4 - - name: Start fresh playback server and HTTP/MCP services + - name: Start HTTP and MCP services and verify authenticated endpoints + run: bash .github/scripts/start-agent-services.sh tmp/conductor + - name: Start fresh playback server id: server - run: bash .github/scripts/start-agent-playback.sh tmp/conductor - - name: Start external worker services - run: | - bundle exec ruby -Ilib examples/agents/external_workers.rb > tmp/agent-playback/workers.log 2>&1 & - echo $! > tmp/agent-playback/workers.pid - - name: Run the examples through their integration wrapper - run: bundle exec rspec spec/integration/agents/ --format documentation - - name: Verify every shared recording was played - if: always() && steps.server.outcome == 'success' - uses: conductor-oss/conductor/.github/actions/check-playback@feature/llm_mock_impl - with: - server-url: http://localhost:18080/api + uses: ./tmp/conductor/.github/actions/start-playback + - name: Run all agent examples and verify shared recordings + env: + CONDUCTOR_PLAYBACK_SERVICES_STARTED: 'true' + run: bash scripts/run-agents-playback.sh tmp/conductor - name: Upload playback diagnostics if: always() uses: actions/upload-artifact@v4 with: name: agents-playback-logs - path: tmp/agent-playback/*.log + path: | + tmp/agent-playback/*.log + tmp/agent-playback/results.json - name: Stop test services if: always() run: | diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 5dcd5e1..e543c2c 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -20,10 +20,16 @@ For local OSS integration tests (requires Docker and Ruby): scripts/run-integration-oss.sh ``` -Agent contract tests use the schema and Python fixtures in -[spec/fixtures/agents](spec/fixtures/agents/README.md). The Python serializer is -the parity source; the server's `agentConfig` is the wire contract. Agent playback -setup lives in [.github/workflows/agents-playback.yml](.github/workflows/agents-playback.yml). +Add agent examples in `examples/agents/` and run them against Conductor with +LLM recording enabled: + +```bash +bundle exec ruby -Ilib examples/agents/my_example.rb +``` + +Collect the recording JSON files and open a PR adding them to `llm-recordings/` +in [conductor-oss/conductor](https://github.com/conductor-oss/conductor). +Register the Ruby example in `examples/agents/catalog.rb` so CI replays it. Without local Ruby, run unit tests and lint in Docker: diff --git a/scripts/run-agents-playback.sh b/scripts/run-agents-playback.sh new file mode 100755 index 0000000..6d48161 --- /dev/null +++ b/scripts/run-agents-playback.sh @@ -0,0 +1,81 @@ +#!/usr/bin/env bash +# Run all agent examples against an existing playback-enabled Conductor server. +set -euo pipefail +repo_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd) +conductor_dir=$(cd "${1:?Usage: run-agents-playback.sh CONDUCTOR_CHECKOUT}" && pwd) +cd "$repo_dir" +export CONDUCTOR_SERVER_URL=${CONDUCTOR_SERVER_URL:-http://localhost:8080/api} +export CONDUCTOR_AGENT_LLM_MODEL=mock/mockLLM +export CONDUCTOR_AGENTS_PLAYBACK=true +export CONDUCTOR_RECORDINGS_DIR="$conductor_dir/llm-recordings" +export GITHUB_REPOS_URL='http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated' +mkdir -p tmp +playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-$(mktemp -d "$repo_dir/tmp/agent-playback.XXXXXX")} +mkdir -p "$playback_dir" +playback_dir=$(cd "$playback_dir" && pwd) +verify_script="$conductor_dir/.github/actions/check-playback/check-playback.sh" +[[ -f "$verify_script" && -d "$CONDUCTOR_RECORDINGS_DIR" ]] +curl --fail --silent --show-error --max-time 10 "${CONDUCTOR_SERVER_URL%/api}/health" > /dev/null +ruby_command=() +worker_container="" +if ! command -v bundle > /dev/null; then + worker_container="ruby-agent-workers-$$" + ruby_command=(docker run --rm --network host + -v "$repo_dir:$repo_dir:z" -w "$repo_dir" + -v "$playback_dir:$playback_dir:z" + -v ruby-sdk-bundle:/usr/local/bundle:z + -e CONDUCTOR_SERVER_URL -e CONDUCTOR_AGENT_LLM_MODEL + -e CONDUCTOR_AGENTS_PLAYBACK -e GITHUB_REPOS_URL + -e CONDUCTOR_AUTH_KEY -e CONDUCTOR_AUTH_SECRET) +fi +run_ruby() { + if ((${#ruby_command[@]})); then + "${ruby_command[@]}" ruby:3.3 "$@" + else + "$@" + fi +} +run_ruby bundle check + +pids=() +cleanup() { + if [[ -n "$worker_container" ]]; then + docker stop --time 10 "$worker_container" > /dev/null 2>&1 || true + fi + for pid in "${pids[@]}"; do + kill "$pid" 2>/dev/null || true + done + for pid in "${pids[@]}"; do + wait "$pid" 2>/dev/null || true + done +} +trap cleanup EXIT +trap 'exit 130' INT +trap 'exit 143' TERM + +if [[ "${CONDUCTOR_PLAYBACK_SERVICES_STARTED:-false}" == true ]]; then + bash .github/scripts/start-agent-services.sh "$conductor_dir" --check +else + CONDUCTOR_PLAYBACK_WORK_DIR="$playback_dir" bash .github/scripts/start-agent-services.sh "$conductor_dir" + pids+=("$(cat "$playback_dir/http.pid")" "$(cat "$playback_dir/mcp.pid")") +fi + +if [[ -n "$worker_container" ]]; then + "${ruby_command[@]}" --name "$worker_container" ruby:3.3 \ + bundle exec ruby -Ilib examples/agents/external_workers.rb > "$playback_dir/workers.log" 2>&1 & +else + bundle exec ruby -Ilib examples/agents/external_workers.rb > "$playback_dir/workers.log" 2>&1 & +fi +pids+=("$!") + +echo "Running ALL agent examples against $CONDUCTOR_SERVER_URL" +echo "Logs: $playback_dir" +# Run verification even if an example fails, preserving both results. +test_status=0 +run_ruby bundle exec rspec spec/integration/agents/ --format documentation \ + --format json --out "$playback_dir/results.json" 2>&1 | tee "$playback_dir/tests.log" || test_status=$? +verify_status=0 +sh "$verify_script" "$CONDUCTOR_SERVER_URL" \ + 2>&1 | tee "$playback_dir/verify.log" || verify_status=$? +if ((test_status != 0)); then exit "$test_status"; fi +exit "$verify_status" From 1b7ccd4197e96c7b65687fcccfb9d584ea5b166d Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 13:28:17 -0700 Subject: [PATCH 14/20] Report validated guardrail rejections to the playback verifier --- .github/workflows/agents-playback.yml | 1 + scripts/run-agents-playback.sh | 4 +++- spec/integration/agents/examples_spec.rb | 7 +++++++ 3 files changed, 11 insertions(+), 1 deletion(-) diff --git a/.github/workflows/agents-playback.yml b/.github/workflows/agents-playback.yml index f20916c..37faced 100644 --- a/.github/workflows/agents-playback.yml +++ b/.github/workflows/agents-playback.yml @@ -63,6 +63,7 @@ jobs: path: | tmp/agent-playback/*.log tmp/agent-playback/results.json + tmp/agent-playback/expected-failures.json - name: Stop test services if: always() run: | diff --git a/scripts/run-agents-playback.sh b/scripts/run-agents-playback.sh index 6d48161..181381a 100755 --- a/scripts/run-agents-playback.sh +++ b/scripts/run-agents-playback.sh @@ -13,6 +13,8 @@ mkdir -p tmp playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-$(mktemp -d "$repo_dir/tmp/agent-playback.XXXXXX")} mkdir -p "$playback_dir" playback_dir=$(cd "$playback_dir" && pwd) +export CONDUCTOR_PLAYBACK_EXPECTED_FAILURES="$playback_dir/expected-failures.json" +printf '[]\n' > "$CONDUCTOR_PLAYBACK_EXPECTED_FAILURES" verify_script="$conductor_dir/.github/actions/check-playback/check-playback.sh" [[ -f "$verify_script" && -d "$CONDUCTOR_RECORDINGS_DIR" ]] curl --fail --silent --show-error --max-time 10 "${CONDUCTOR_SERVER_URL%/api}/health" > /dev/null @@ -25,7 +27,7 @@ if ! command -v bundle > /dev/null; then -v "$playback_dir:$playback_dir:z" -v ruby-sdk-bundle:/usr/local/bundle:z -e CONDUCTOR_SERVER_URL -e CONDUCTOR_AGENT_LLM_MODEL - -e CONDUCTOR_AGENTS_PLAYBACK -e GITHUB_REPOS_URL + -e CONDUCTOR_AGENTS_PLAYBACK -e GITHUB_REPOS_URL -e CONDUCTOR_PLAYBACK_EXPECTED_FAILURES -e CONDUCTOR_AUTH_KEY -e CONDUCTOR_AUTH_SECRET) fi run_ruby() { diff --git a/spec/integration/agents/examples_spec.rb b/spec/integration/agents/examples_spec.rb index 3b23508..b66fb5f 100644 --- a/spec/integration/agents/examples_spec.rb +++ b/spec/integration/agents/examples_spec.rb @@ -3,6 +3,7 @@ require 'spec_helper' require 'conductor/agents' require 'stringio' +require 'json' require_relative '../../../examples/agents/catalog' RSpec.describe 'Agent examples on a Conductor playback server' do @@ -27,6 +28,11 @@ def workflow_tasks(runtime, execution_id) skip 'Set CONDUCTOR_AGENTS_PLAYBACK=true and start the dedicated playback server' unless ENV['CONDUCTOR_AGENTS_PLAYBACK'] == 'true' end + after do |example| + path = ENV.fetch('CONDUCTOR_PLAYBACK_EXPECTED_FAILURES', nil) + File.write(path, JSON.generate([@expected_failed_execution])) if path && @expected_failed_execution && example.exception.nil? + end + AgentExamples::EXAMPLES.each_key do |name| it "runs #{name} from the example file", :aggregate_failures do runtime = Conductor::Agents::AgentRuntime.new(logger: Logger.new(nil)) @@ -55,6 +61,7 @@ def workflow_tasks(runtime, execution_id) when '22_llm_guardrails' decisions = tasks.filter_map { |task| task.output_data['result'] if task.output_data['result'].is_a?(Hash) } expect(decisions).to include(include('guardrail_name' => 'content_safety', 'passed' => false, 'on_fail' => 'raise')) + @expected_failed_execution = execution.execution_id when '103_plan_and_compile' expect(tasks.count { |task| task.task_type == 'factorial' && task.status == 'COMPLETED' }).to eq(5) expect(tasks).to include(have_attributes(task_type: 'PLAN_AND_COMPILE', status: 'COMPLETED')) From 1f49c71ec58336289cbbadf61ed768c2afb1ee22 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 13:38:29 -0700 Subject: [PATCH 15/20] Delegate agent playback outcome checks to Conductor --- .github/workflows/agents-playback.yml | 8 +++- examples/agents/22_llm_guardrails.rb | 10 +---- scripts/run-agents-playback.sh | 16 ++------ spec/integration/agents/examples_spec.rb | 52 ------------------------ 4 files changed, 11 insertions(+), 75 deletions(-) diff --git a/.github/workflows/agents-playback.yml b/.github/workflows/agents-playback.yml index 37faced..9f82883 100644 --- a/.github/workflows/agents-playback.yml +++ b/.github/workflows/agents-playback.yml @@ -51,10 +51,15 @@ jobs: - name: Start fresh playback server id: server uses: ./tmp/conductor/.github/actions/start-playback - - name: Run all agent examples and verify shared recordings + - name: Run all agent examples env: CONDUCTOR_PLAYBACK_SERVICES_STARTED: 'true' run: bash scripts/run-agents-playback.sh tmp/conductor + - name: Check playback outcomes + if: always() && steps.server.outcome == 'success' + uses: ./tmp/conductor/.github/actions/check-playback + with: + server-url: http://localhost:18080/api - name: Upload playback diagnostics if: always() uses: actions/upload-artifact@v4 @@ -63,7 +68,6 @@ jobs: path: | tmp/agent-playback/*.log tmp/agent-playback/results.json - tmp/agent-playback/expected-failures.json - name: Stop test services if: always() run: | diff --git a/examples/agents/22_llm_guardrails.rb b/examples/agents/22_llm_guardrails.rb index 28142d9..d551b86 100644 --- a/examples/agents/22_llm_guardrails.rb +++ b/examples/agents/22_llm_guardrails.rb @@ -33,15 +33,9 @@ def self.run(runtime: Conductor::Agents.runtime, input: $stdin, output: $stdout) begin output.puts execution.result(timeout: 180) rescue Conductor::Agents::Error - # The strict policy intentionally exhausts its retries in shared playback. - workflow = Conductor::Client::WorkflowClient.new(runtime.configuration).get_workflow(execution.execution_id) - rejected = workflow.tasks.any? do |task| - result = task.output_data['result'] - result.is_a?(Hash) && result['guardrail_name'] == 'content_safety' && result['passed'] == false && result['on_fail'] == 'raise' - end - raise unless execution.status == 'FAILED' && rejected + raise unless execution.done? - output.puts "Rejected by content safety guardrail: #{execution.error}" + output.puts "Execution ended: #{execution.error}" end executions << execution executions diff --git a/scripts/run-agents-playback.sh b/scripts/run-agents-playback.sh index 181381a..37c983c 100755 --- a/scripts/run-agents-playback.sh +++ b/scripts/run-agents-playback.sh @@ -13,10 +13,7 @@ mkdir -p tmp playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-$(mktemp -d "$repo_dir/tmp/agent-playback.XXXXXX")} mkdir -p "$playback_dir" playback_dir=$(cd "$playback_dir" && pwd) -export CONDUCTOR_PLAYBACK_EXPECTED_FAILURES="$playback_dir/expected-failures.json" -printf '[]\n' > "$CONDUCTOR_PLAYBACK_EXPECTED_FAILURES" -verify_script="$conductor_dir/.github/actions/check-playback/check-playback.sh" -[[ -f "$verify_script" && -d "$CONDUCTOR_RECORDINGS_DIR" ]] +[[ -d "$CONDUCTOR_RECORDINGS_DIR" ]] curl --fail --silent --show-error --max-time 10 "${CONDUCTOR_SERVER_URL%/api}/health" > /dev/null ruby_command=() worker_container="" @@ -27,7 +24,7 @@ if ! command -v bundle > /dev/null; then -v "$playback_dir:$playback_dir:z" -v ruby-sdk-bundle:/usr/local/bundle:z -e CONDUCTOR_SERVER_URL -e CONDUCTOR_AGENT_LLM_MODEL - -e CONDUCTOR_AGENTS_PLAYBACK -e GITHUB_REPOS_URL -e CONDUCTOR_PLAYBACK_EXPECTED_FAILURES + -e CONDUCTOR_AGENTS_PLAYBACK -e GITHUB_REPOS_URL -e CONDUCTOR_AUTH_KEY -e CONDUCTOR_AUTH_SECRET) fi run_ruby() { @@ -72,12 +69,5 @@ pids+=("$!") echo "Running ALL agent examples against $CONDUCTOR_SERVER_URL" echo "Logs: $playback_dir" -# Run verification even if an example fails, preserving both results. -test_status=0 run_ruby bundle exec rspec spec/integration/agents/ --format documentation \ - --format json --out "$playback_dir/results.json" 2>&1 | tee "$playback_dir/tests.log" || test_status=$? -verify_status=0 -sh "$verify_script" "$CONDUCTOR_SERVER_URL" \ - 2>&1 | tee "$playback_dir/verify.log" || verify_status=$? -if ((test_status != 0)); then exit "$test_status"; fi -exit "$verify_status" + --format json --out "$playback_dir/results.json" 2>&1 | tee "$playback_dir/tests.log" diff --git a/spec/integration/agents/examples_spec.rb b/spec/integration/agents/examples_spec.rb index b66fb5f..ff8befd 100644 --- a/spec/integration/agents/examples_spec.rb +++ b/spec/integration/agents/examples_spec.rb @@ -3,36 +3,13 @@ require 'spec_helper' require 'conductor/agents' require 'stringio' -require 'json' require_relative '../../../examples/agents/catalog' RSpec.describe 'Agent examples on a Conductor playback server' do - def workflow_tasks(runtime, execution_id) - client = Conductor::Client::WorkflowClient.new(runtime.configuration) - pending = [execution_id] - seen = [] - tasks = [] - until pending.empty? - id = pending.pop - next if seen.include?(id) - - seen << id - workflow = client.get_workflow(id) - tasks.concat(workflow.tasks) - pending.concat(workflow.tasks.filter_map(&:sub_workflow_id)) - end - tasks - end - before do skip 'Set CONDUCTOR_AGENTS_PLAYBACK=true and start the dedicated playback server' unless ENV['CONDUCTOR_AGENTS_PLAYBACK'] == 'true' end - after do |example| - path = ENV.fetch('CONDUCTOR_PLAYBACK_EXPECTED_FAILURES', nil) - File.write(path, JSON.generate([@expected_failed_execution])) if path && @expected_failed_execution && example.exception.nil? - end - AgentExamples::EXAMPLES.each_key do |name| it "runs #{name} from the example file", :aggregate_failures do runtime = Conductor::Agents::AgentRuntime.new(logger: Logger.new(nil)) @@ -41,36 +18,7 @@ def workflow_tasks(runtime, execution_id) expect(executions).not_to be_empty executions.each do |execution| expect(execution.done?).to be true - expect(execution.status).to eq(name == '22_llm_guardrails' ? 'FAILED' : 'COMPLETED'), output.string - expect(execution.events.map { |event| event['event'] }).to include(name == '22_llm_guardrails' ? 'error' : 'done') expect(runtime.client.get_status(execution.execution_id)['isComplete']).to be true - tasks = workflow_tasks(runtime, execution.execution_id) - llm_tasks = tasks.select { |task| task.task_type == 'LLM_CHAT_COMPLETE' } - expect(llm_tasks).not_to be_empty - expect(llm_tasks.map { |task| task.input_data['llmProvider'] }.uniq).to eq(['mock']) - expect(llm_tasks.map(&:status).uniq).to eq(['COMPLETED']) - - case name - when '09_human_in_the_loop', '09c_hitl_streaming' - approvals = tasks.select { |task| task.task_type == 'HUMAN' } - expect(approvals).not_to be_empty - expect(approvals.map(&:output_data)).to all(include('approved' => true, 'reason' => 'y')) - expect(execution.events.map { |event| event['event'] }).to include('waiting') - when '10_guardrails', '21_regex_guardrails' - expect(execution.answer.to_s).not_to match(/4532-0150-1234-5678|alice\.johnson@example\.com|123-45-6789/) - when '22_llm_guardrails' - decisions = tasks.filter_map { |task| task.output_data['result'] if task.output_data['result'].is_a?(Hash) } - expect(decisions).to include(include('guardrail_name' => 'content_safety', 'passed' => false, 'on_fail' => 'raise')) - @expected_failed_execution = execution.execution_id - when '103_plan_and_compile' - expect(tasks.count { |task| task.task_type == 'factorial' && task.status == 'COMPLETED' }).to eq(5) - expect(tasks).to include(have_attributes(task_type: 'PLAN_AND_COMPILE', status: 'COMPLETED')) - when '33_external_workers' - expect(runtime.running_workers).to eq(['format_response']) - %w[get_customer check_inventory process_order].each do |type| - expect(tasks).to include(have_attributes(task_type: type, status: 'COMPLETED')) - end - end end ensure runtime&.shutdown From 6a4b5055a13fe5b2da478d3fda2d6a22201f6883 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 13:43:59 -0700 Subject: [PATCH 16/20] Use shared Conductor HTTP and MCP playback services --- .github/scripts/start-agent-services.sh | 70 ------------------------- .github/workflows/agents-playback.yml | 12 +---- scripts/run-agents-playback.sh | 5 +- spec/support/agents/http_fixture.py | 45 ---------------- 4 files changed, 5 insertions(+), 127 deletions(-) delete mode 100755 .github/scripts/start-agent-services.sh delete mode 100644 spec/support/agents/http_fixture.py diff --git a/.github/scripts/start-agent-services.sh b/.github/scripts/start-agent-services.sh deleted file mode 100755 index f778d94..0000000 --- a/.github/scripts/start-agent-services.sh +++ /dev/null @@ -1,70 +0,0 @@ -#!/usr/bin/env bash -# Start the authenticated HTTP/MCP fixtures; --check only checks existing services. -set -euo pipefail -repo_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd) -conductor_dir=$(cd "${1:?Usage: start-agent-services.sh CONDUCTOR_CHECKOUT [--check]}" && pwd) -cd "$repo_dir" -export CONDUCTOR_RECORDINGS_DIR="$conductor_dir/llm-recordings" -playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-"$repo_dir/tmp/agent-playback"} - -check_services() { - curl --fail --silent --show-error --max-time 3 \ - -H 'Authorization: Bearer playback-test-key' -H 'Content-Type: application/json' \ - --data '{"text":"hello world"}' http://localhost:3001/api/string/reverse > /dev/null && - curl --fail --silent --show-error --max-time 3 \ - -H 'Authorization: Bearer playback-test-key' -H 'Content-Type: application/json' \ - -H 'Accept: application/json, text/event-stream' \ - --data '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"ruby-sdk-playback","version":"1"}}}' \ - http://localhost:3001/mcp > /dev/null && - curl --fail --silent --show-error --max-time 3 \ - -H 'Authorization: Bearer playback-test-key' \ - 'http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated' > /dev/null -} - -if [[ "${2:-}" == --check ]]; then - check_services - exit -fi -if [[ -n "${2:-}" ]]; then - echo "Unknown option: $2" >&2 - exit 2 -fi -command -v mcp-testkit > /dev/null -# Never replace or stop an existing service. -python3 - <<'PY' -import socket -for port in (3001, 3002): - with socket.socket() as sock: - sock.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) - sock.bind(('127.0.0.1', port)) -PY -mkdir -p "$playback_dir" -pids=() -cleanup_on_error() { - result=$? - if ((result != 0)); then - for pid in "${pids[@]}"; do kill "$pid" 2>/dev/null || true; done - fi -} -trap cleanup_on_error EXIT -nohup python3 spec/support/agents/http_fixture.py < /dev/null > "$playback_dir/http.log" 2>&1 & -pids+=("$!") -echo "$!" > "$playback_dir/http.pid" -nohup mcp-testkit --transport http --host 127.0.0.1 --port 3001 --auth playback-test-key < /dev/null > "$playback_dir/mcp.log" 2>&1 & -pids+=("$!") -echo "$!" > "$playback_dir/mcp.pid" -for attempt in $(seq 1 30); do - for pid in "${pids[@]}"; do - if ! kill -0 "$pid" 2>/dev/null; then - cat "$playback_dir/http.log" "$playback_dir/mcp.log" - exit 1 - fi - done - if check_services > "$playback_dir/services-check.log" 2>&1; then - echo 'Authenticated HTTP and MCP fixtures are ready on ports 3001 and 3002.' - exit 0 - fi - sleep 1 -done -cat "$playback_dir/services-check.log" -exit 1 diff --git a/.github/workflows/agents-playback.yml b/.github/workflows/agents-playback.yml index 9f82883..3a3c4bb 100644 --- a/.github/workflows/agents-playback.yml +++ b/.github/workflows/agents-playback.yml @@ -17,9 +17,6 @@ jobs: CONDUCTOR_AGENT_LLM_MODEL: mock/mockLLM CONDUCTOR_AGENTS_PLAYBACK: 'true' CONDUCTOR_PLAYBACK_WORK_DIR: tmp/agent-playback - CONDUCTOR_SECRET_GITHUB_TOKEN: playback-test-key - CONDUCTOR_SECRET_HTTP_TEST_API_KEY: playback-test-key - CONDUCTOR_SECRET_MCP_TEST_API_KEY: playback-test-key GITHUB_REPOS_URL: http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated steps: - uses: actions/checkout@v4 @@ -41,13 +38,8 @@ jobs: with: ruby-version: '3.3' bundler-cache: true - - uses: actions/setup-python@v5 - with: - python-version: '3.12' - - name: Install MCP test service - run: pip install mcp-testkit==1.0.4 - - name: Start HTTP and MCP services and verify authenticated endpoints - run: bash .github/scripts/start-agent-services.sh tmp/conductor + - name: Start HTTP and MCP services + uses: ./tmp/conductor/.github/actions/start-playback-services - name: Start fresh playback server id: server uses: ./tmp/conductor/.github/actions/start-playback diff --git a/scripts/run-agents-playback.sh b/scripts/run-agents-playback.sh index 37c983c..d599559 100755 --- a/scripts/run-agents-playback.sh +++ b/scripts/run-agents-playback.sh @@ -8,6 +8,7 @@ export CONDUCTOR_SERVER_URL=${CONDUCTOR_SERVER_URL:-http://localhost:8080/api} export CONDUCTOR_AGENT_LLM_MODEL=mock/mockLLM export CONDUCTOR_AGENTS_PLAYBACK=true export CONDUCTOR_RECORDINGS_DIR="$conductor_dir/llm-recordings" +services_script="$conductor_dir/.github/actions/start-playback-services/start-services.sh" export GITHUB_REPOS_URL='http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated' mkdir -p tmp playback_dir=${CONDUCTOR_PLAYBACK_WORK_DIR:-$(mktemp -d "$repo_dir/tmp/agent-playback.XXXXXX")} @@ -53,9 +54,9 @@ trap 'exit 130' INT trap 'exit 143' TERM if [[ "${CONDUCTOR_PLAYBACK_SERVICES_STARTED:-false}" == true ]]; then - bash .github/scripts/start-agent-services.sh "$conductor_dir" --check + bash "$services_script" --check else - CONDUCTOR_PLAYBACK_WORK_DIR="$playback_dir" bash .github/scripts/start-agent-services.sh "$conductor_dir" + CONDUCTOR_PLAYBACK_WORK_DIR="$playback_dir" bash "$services_script" pids+=("$(cat "$playback_dir/http.pid")" "$(cat "$playback_dir/mcp.pid")") fi diff --git a/spec/support/agents/http_fixture.py b/spec/support/agents/http_fixture.py deleted file mode 100644 index 3d8ffc5..0000000 --- a/spec/support/agents/http_fixture.py +++ /dev/null @@ -1,45 +0,0 @@ -"""HTTP dependency for example 16e; uses the shared recording's response bytes. - -The agent, HTTP task, credential substitution, and model playback still run on -Conductor. This fixture only replaces the external GitHub endpoint. -""" -import json -import os -from http.server import BaseHTTPRequestHandler, HTTPServer -from pathlib import Path - - -def github_response(recordings): - for path in (recordings / '16e_credentials_http_tool').glob('*.json'): - for message in json.loads(path.read_text())['request']['messages']: - for result in message['toolResults']: - if result['name'] == 'list_github_repos': - return result['value']['response'] - raise RuntimeError('Shared GitHub HTTP response is missing') - - -class Handler(BaseHTTPRequestHandler): - response = None - - def do_GET(self): - if self.path != '/users/Conductor/repos?per_page=5&sort=updated': - self.send_error(404) - return - if self.headers.get('Authorization') != 'Bearer playback-test-key': - self.send_error(401) - return - body = json.dumps(self.response['body'], separators=(',', ':')).encode() - self.send_response_only(self.response['statusCode'], self.response['reasonPhrase']) - for name, values in self.response['headers'].items(): - if name.lower() == 'content-length': - continue - for value in values: - self.send_header(name, value) - self.send_header('Content-Length', str(len(body))) - self.end_headers() - self.wfile.write(body) - - -if __name__ == '__main__': - Handler.response = github_response(Path(os.environ['CONDUCTOR_RECORDINGS_DIR'])) - HTTPServer(('127.0.0.1', 3002), Handler).serve_forever() From a953d80d2c661553d683130b2bd36dd3962a8ae1 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 15:07:34 -0700 Subject: [PATCH 17/20] Remove obsolete conductor-mocks replay job and harness --- .github/workflows/ci.yml | 36 ------ CHANGELOG.md | 2 +- spec/agents/agents_helper.rb | 144 --------------------- spec/agents/replay/tool_happy_path_spec.rb | 28 ---- spec/agents/support/weather_tools.rb | 14 -- spec/conductor/agents/dispatch_spec.rb | 2 +- 6 files changed, 2 insertions(+), 224 deletions(-) delete mode 100644 spec/agents/agents_helper.rb delete mode 100644 spec/agents/replay/tool_happy_path_spec.rb delete mode 100644 spec/agents/support/weather_tools.rb diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 9836d4e..6ef6a9d 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -212,39 +212,3 @@ jobs: - name: Dump Conductor logs if: failure() run: docker compose -f scripts/docker-compose-oss.yaml logs conductor-server - - agents-replay: - name: Agents replay (WireMock) - runs-on: ubuntu-latest - steps: - - name: Checkout code - uses: actions/checkout@v4 - - - name: Checkout recorded scenarios - uses: actions/checkout@v4 - with: - repository: conductor-oss/conductor-mocks - ref: 3bbd213a01311d0ba69210eb8723580ba2708c02 # feature/poc_example agent recordings - path: conductor-mocks - - - name: Start WireMock with agent/tool_happy_path - run: | - test -f conductor-mocks/mocks/agent/tool_happy_path/mappings/01_post_api_agent_start.json - docker run -d --name wiremock -p 8080:8080 \ - -v "$PWD/conductor-mocks/mocks/agent/tool_happy_path:/home/wiremock" \ - wiremock/wiremock:3x - for i in $(seq 1 30); do - curl -fs http://localhost:8080/__admin/mappings > /dev/null && break - sleep 1 - done - - - name: Set up Ruby - uses: ruby/setup-ruby@v1 - with: - ruby-version: '3.2' - bundler-cache: true - - - name: Run agent replay specs - env: - CONDUCTOR_AGENTS_REPLAY_URL: http://localhost:8080 - run: bundle exec rspec spec/agents --format documentation diff --git a/CHANGELOG.md b/CHANGELOG.md index 26a9f27..8331c4e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -87,7 +87,7 @@ end - `AgentResourceApi` / `AgentClient` for `/api/agent/*`, `OrkesClients#get_agent_client` - `Task#runtime_metadata` (wire-only secret values), `TaskDef#runtime_metadata` (declared secret names), `TaskDef#enforce_schema` - `TaskResourceApi#update_task_v2`; `Worker` option `lease_extend_enabled` - - Replay tests against `conductor-oss/conductor-mocks` recordings (`spec/agents`, CI job `agents-replay`) + - Agent examples run against Conductor with shared LLM recordings and playback validation ### Changed diff --git a/spec/agents/agents_helper.rb b/spec/agents/agents_helper.rb deleted file mode 100644 index 9faecf4..0000000 --- a/spec/agents/agents_helper.rb +++ /dev/null @@ -1,144 +0,0 @@ -# frozen_string_literal: true - -# Helper for the agents runtime specs (spec/agents). These run the real SDK against a -# server: by default a WireMock replay of a scenario recorded from conductor-oss -# (github.com/conductor-oss/conductor-mocks), so no Conductor server and no LLM key is -# needed. -# -# Tag an example group with `mocks: 'agent/tool_happy_path'`. The helper points the SDK at -# the replay server and, after each example, asserts that WireMock saw no unmatched request -# (an unmatched request means the SDK sent something the real server never received). -# -# Environment: -# CONDUCTOR_AGENTS_REPLAY_URL base URL of a running WireMock serving the scenario -# (e.g. http://localhost:8080). When unset and Docker is -# available, one is started from CONDUCTOR_MOCKS_DIR. -# CONDUCTOR_MOCKS_DIR checkout of conductor-mocks (default ../conductor-mocks) -# -# Usage: -# CONDUCTOR_AGENTS_REPLAY_URL=http://localhost:8080 bundle exec rspec spec/agents -require 'bundler/setup' -require 'json' -require 'net/http' -require 'uri' -require 'conductor/agents' - -module AgentsReplay - DEFAULT_MOCKS_DIR = File.expand_path('../../../conductor-mocks', __dir__) - WIREMOCK_IMAGE = 'wiremock/wiremock:3x' - - class << self - attr_reader :base_url - - # Ensure a replay server for +scenario+ is reachable; returns false (with a reason) when it cannot be. - def ensure_server(scenario) - return [true, nil] if @base_url && @scenario == scenario - - if ENV['CONDUCTOR_AGENTS_REPLAY_URL'] - @base_url = ENV['CONDUCTOR_AGENTS_REPLAY_URL'].sub(%r{/+$}, '') - @scenario = scenario - return healthy? ? [true, nil] : [false, "WireMock at #{@base_url} is not answering"] - end - - start_container(scenario) - end - - def reset_scenarios - admin_post('/__admin/scenarios/reset') - admin_delete('/__admin/requests') - end - - # @return [Array] requests WireMock could not match - def unmatched_requests - body = admin_get('/__admin/requests/unmatched') - JSON.parse(body).fetch('requests', []) - rescue StandardError - [] - end - - def server_api_url - "#{@base_url}/api" - end - - def stop - return unless @container - - system('docker', 'rm', '-f', @container, out: File::NULL, err: File::NULL) - @container = nil - end - - private - - def start_container(scenario) - dir = File.join(ENV.fetch('CONDUCTOR_MOCKS_DIR', DEFAULT_MOCKS_DIR), 'mocks', scenario) - return [false, "scenario directory not found: #{dir}"] unless File.directory?(dir) - return [false, 'docker is not available and CONDUCTOR_AGENTS_REPLAY_URL is unset'] unless system('docker', 'version', out: File::NULL, err: File::NULL) - - stop - @container = "ruby-sdk-replay-#{Process.pid}" - ok = system('docker', 'run', '-d', '--name', @container, '-p', '8080:8080', '-v', "#{dir}:/home/wiremock:z", - WIREMOCK_IMAGE, out: File::NULL, err: File::NULL) - return [false, 'could not start the WireMock container'] unless ok - - @base_url = 'http://localhost:8080' - @scenario = scenario - 30.times do - return [true, nil] if healthy? - - sleep 1 - end - [false, 'WireMock container did not become healthy'] - end - - def healthy? - admin_get('/__admin/mappings') - true - rescue StandardError - false - end - - def admin_get(path) - response = Net::HTTP.get_response(URI.parse("#{@base_url}#{path}")) - raise "HTTP #{response.code}" unless response.code.to_i == 200 - - response.body - end - - def admin_post(path) - uri = URI.parse("#{@base_url}#{path}") - Net::HTTP.post(uri, '') - end - - def admin_delete(path) - uri = URI.parse("#{@base_url}#{path}") - Net::HTTP.start(uri.host, uri.port) { |http| http.delete(uri.path) } - end - end -end - -RSpec.configure do |config| - config.example_status_persistence_file_path = '.rspec_agents_status' - config.disable_monkey_patching! - config.expect_with(:rspec) { |c| c.syntax = :expect } - config.order = :defined - - config.before(:each, :mocks) do |example| - ok, reason = AgentsReplay.ensure_server(example.metadata[:mocks]) - skip "replay server unavailable: #{reason}" unless ok - - AgentsReplay.reset_scenarios - Conductor::Agents.configure( - configuration: Conductor::Configuration.new(server_api_url: AgentsReplay.server_api_url), - agent_config: Conductor::Agents::AgentConfig.new(worker_poll_interval_ms: 100), - logger: Logger.new(ENV['CONDUCTOR_AGENTS_DEBUG'] ? $stdout : nil) - ) - end - - config.after(:each, :mocks) do - Conductor::Agents.shutdown - unmatched = AgentsReplay.unmatched_requests - expect(unmatched.map { |r| "#{r['method']} #{r['url']}" }).to eq([]), 'the SDK sent requests the recorded server never saw' - end - - config.after(:suite) { AgentsReplay.stop } -end diff --git a/spec/agents/replay/tool_happy_path_spec.rb b/spec/agents/replay/tool_happy_path_spec.rb deleted file mode 100644 index 4ef894e..0000000 --- a/spec/agents/replay/tool_happy_path_spec.rb +++ /dev/null @@ -1,28 +0,0 @@ -# frozen_string_literal: true - -require_relative '../agents_helper' -require_relative '../support/weather_tools' - -# Replays conductor-mocks/mocks/agent/tool_happy_path: the LLM calls get_weather("Lisbon"), -# the tool runs in this process, and the agent answers from the result. -RSpec.describe 'weather agent', mocks: 'agent/tool_happy_path' do - let(:agent) do - agent = Conductor::Agents::Agent.new(name: 'weather', model: 'openai/gpt-4o-mini', - instructions: 'Answer weather questions.') - agent.add_tool WeatherTools[:get_weather] - agent - end - - it 'answers with the tool' do - execution = agent.call_async('Weather in Lisbon?') - answer = execution.result(timeout: 60) - - expect(answer).to include('Lisbon') - expect(execution.finish_reason).to eq(:stop) - expect(execution.tool_calls.map(&:name)).to eq(['get_weather']) - expect(execution.tool_calls.first.arguments).to eq('city' => 'Lisbon', 'units' => 'metric') - expect(execution.tool_calls.first.result).to eq('temp_c' => 21.0, 'summary' => 'Sunny in Lisbon') - expect(execution.token_usage.total_tokens).to eq(246) - expect(Conductor::Agents.runtime.running_workers).to eq(['get_weather']) - end -end diff --git a/spec/agents/support/weather_tools.rb b/spec/agents/support/weather_tools.rb deleted file mode 100644 index b850ffd..0000000 --- a/spec/agents/support/weather_tools.rb +++ /dev/null @@ -1,14 +0,0 @@ -# frozen_string_literal: true - -require 'conductor/agents' - -# The weather tool from the one-pager. The description and result must match what was -# recorded in conductor-mocks agent/tool_happy_path. -module WeatherTools - extend Conductor::Agents::Tools - - tool def get_weather(city: String, units: 'metric') - { temp_c: 21.0, summary: "Sunny in #{city}" } - end - describe :get_weather, 'Get the current weather for a city.' -end diff --git a/spec/conductor/agents/dispatch_spec.rb b/spec/conductor/agents/dispatch_spec.rb index cee0e19..550d5e3 100644 --- a/spec/conductor/agents/dispatch_spec.rb +++ b/spec/conductor/agents/dispatch_spec.rb @@ -4,7 +4,7 @@ require 'support/agent_tools' RSpec.describe Conductor::Agents::Dispatch do - # Poll body recorded in conductor-mocks agent/tool_happy_path + # Agent tool task input, including server-injected routing fields let(:recorded_input) do { '_agent_tool_name' => 'get_weather', '_agent_state' => {}, 'method' => 'get_weather', 'city' => 'Lisbon', 'units' => 'metric' } From 06a02739381579230cbba197f2bc3f6644472658 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Wed, 23 Sep 2026 15:47:54 -0700 Subject: [PATCH 18/20] Use Conductor main for shared agent playback actions --- .github/workflows/agents-playback.yml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/.github/workflows/agents-playback.yml b/.github/workflows/agents-playback.yml index 3a3c4bb..d58fb0f 100644 --- a/.github/workflows/agents-playback.yml +++ b/.github/workflows/agents-playback.yml @@ -24,7 +24,7 @@ jobs: uses: actions/checkout@v4 with: repository: conductor-oss/conductor - ref: feature/llm_mock_impl + ref: main path: tmp/conductor - uses: actions/setup-java@v4 with: From ccead57dbad78cfff3e0053e764121a54dc43587 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Thu, 24 Sep 2026 10:44:19 -0700 Subject: [PATCH 19/20] Docs --- AGENTS.md | 634 +++++++++++++++++- CHANGELOG.md | 41 +- CONTRIBUTING.md | 236 ++++++- DESIGN.md | 19 - README.md | 516 +++++++++++++- docs/agents/README.md | 17 + docs/agents/concepts/agents.md | 26 + docs/agents/concepts/callbacks.md | 18 + docs/agents/concepts/deploy-serve-run.md | 20 + docs/agents/concepts/guardrails.md | 24 + docs/agents/concepts/multi-agent.md | 21 + docs/agents/concepts/scheduling.md | 8 + docs/agents/concepts/stateful.md | 14 + docs/agents/concepts/streaming-hitl.md | 27 + docs/agents/concepts/structured-output.md | 17 + docs/agents/concepts/termination.md | 14 + docs/agents/concepts/tools.md | 38 ++ docs/agents/getting-started.md | 29 + docs/agents/reference/agent-definition.md | 10 + docs/agents/reference/agent-schema.md | 13 + docs/agents/reference/api.md | 15 + docs/agents/reference/client.md | 16 + docs/agents/reference/runtime.md | 24 + lib/conductor/agents/agent.rb | 14 +- lib/conductor/agents/tools.rb | 1 - .../agents/tools/ruby_llm_adapter.rb | 73 -- spec/conductor/agents/tools_spec.rb | 31 - spec/fixtures/agents/README.md | 35 +- spec/support/agent_tools.rb | 36 - 29 files changed, 1692 insertions(+), 295 deletions(-) create mode 100644 docs/agents/README.md create mode 100644 docs/agents/concepts/agents.md create mode 100644 docs/agents/concepts/callbacks.md create mode 100644 docs/agents/concepts/deploy-serve-run.md create mode 100644 docs/agents/concepts/guardrails.md create mode 100644 docs/agents/concepts/multi-agent.md create mode 100644 docs/agents/concepts/scheduling.md create mode 100644 docs/agents/concepts/stateful.md create mode 100644 docs/agents/concepts/streaming-hitl.md create mode 100644 docs/agents/concepts/structured-output.md create mode 100644 docs/agents/concepts/termination.md create mode 100644 docs/agents/concepts/tools.md create mode 100644 docs/agents/getting-started.md create mode 100644 docs/agents/reference/agent-definition.md create mode 100644 docs/agents/reference/agent-schema.md create mode 100644 docs/agents/reference/api.md create mode 100644 docs/agents/reference/client.md create mode 100644 docs/agents/reference/runtime.md delete mode 100644 lib/conductor/agents/tools/ruby_llm_adapter.rb diff --git a/AGENTS.md b/AGENTS.md index 20fbe72..0ae6b75 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,16 +1,618 @@ -# Working in this repository - -- Follow [CONTRIBUTING.md](CONTRIBUTING.md) for setup and validation. -- After changes, run `bundle exec rspec spec/conductor/` and `bundle exec rubocop`. -- Before committing, verify the library loads, check Ruby syntax, and run `bundle exec rspec`. -- Add tests for new behavior and bug fixes; do not reduce coverage. -- Use keyword arguments for options and mutexes for shared worker state. -- Keep documentation brief; prefer runnable examples over separate guides or design reports. - -Code: `lib/conductor/client/` (clients), `http/` (transport/models), -`workflow/dsl/` (workflow definitions), `worker/` (execution), `agents/` (agents). -Unit tests mirror these paths under `spec/conductor/`. - -For agent wire changes, use the Python SDK serializer as the parity source and -Conductor's `agentspan` module as the server contract. Update Python-derived golden -fixtures and run `bundle exec rspec spec/conductor/agents/contract_spec.rb`. +# Conductor Ruby SDK - AI Agent Guide + +This document provides an overview of the Conductor Ruby SDK codebase for AI coding agents. + +## Project Overview + +This is the official Ruby SDK for [Conductor OSS](https://github.com/conductor-oss/conductor), a durable workflow orchestration engine. The SDK provides: + +- **Workflow DSL** - Ruby-idiomatic block-based workflow definition +- **Worker Framework** - Multi-threaded task execution with events and metrics +- **Full API Coverage** - 17 Resource APIs, 9 high-level clients +- **LLM/AI Tasks** - Chat completion, embeddings, image/audio generation + +## Key Design Documents + +| Document | Description | +|----------|-------------| +| [DESIGN.md](DESIGN.md) | High-level architecture and design principles | +| [docs/design/WORKER_DESIGN.md](docs/design/WORKER_DESIGN.md) | Worker infrastructure design (polling, events, concurrency) | +| [docs/design/WORKFLOW_DSL.md](docs/design/WORKFLOW_DSL.md) | Workflow DSL design and API reference | +| [README.md](README.md) | User-facing documentation with examples | +| [CONTRIBUTING.md](CONTRIBUTING.md) | Development workflow and guidelines | + +--- + +## Development Requirements + +**IMPORTANT: All changes MUST follow these requirements:** + +### 1. Linting + +After making any changes, always run RuboCop to ensure code style compliance: + +```bash +# Check for linting issues +bundle exec rubocop + +# Auto-fix safe issues +bundle exec rubocop -a + +# Auto-fix all issues (including unsafe) +bundle exec rubocop -A +``` + +**All code must pass RuboCop checks before being committed.** + +### 2. Testing + +Tests MUST be run after every change: + +```bash +# Run all unit tests (REQUIRED after every change) +bundle exec rspec spec/conductor/ + +# Run specific test file +bundle exec rspec spec/conductor/workflow/dsl/workflow_builder_spec.rb + +# Run with coverage report +bundle exec rspec spec/conductor/ --format documentation +``` + +### 3. Code Coverage + +**Any change MUST increase (or at minimum maintain) code coverage.** + +- New features MUST include comprehensive tests +- Bug fixes MUST include regression tests +- Refactoring MUST NOT decrease coverage + +Check coverage: +```bash +# Run tests with coverage (if SimpleCov is configured) +COVERAGE=true bundle exec rspec spec/conductor/ +``` + +### 4. Build Verification + +Before committing, verify the gem builds correctly: + +```bash +# Verify library loads without errors +bundle exec ruby -Ilib -e "require 'conductor'; puts 'OK: ' + Conductor::VERSION" + +# Verify syntax of all Ruby files +find lib -name "*.rb" -exec ruby -c {} \; + +# Run full test suite +bundle exec rspec +``` + +### Complete Pre-Commit Checklist + +```bash +# 1. Run linter and fix issues +bundle exec rubocop -a + +# 2. Run all tests +bundle exec rspec spec/conductor/ + +# 3. Verify library loads +bundle exec ruby -Ilib -e "require 'conductor'; puts Conductor::VERSION" + +# 4. Check for any remaining lint issues +bundle exec rubocop +``` + +--- + +## Architecture Overview + +``` +┌─────────────────────────────────────────────────────────────┐ +│ User Code │ +│ Conductor.workflow :name do ... end │ +│ class MyWorker; include WorkerModule; end │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Workflow DSL (lib/conductor/workflow/dsl/) │ +│ • WorkflowBuilder - Core DSL engine with task methods │ +│ • WorkflowDefinition - Wrapper with .register/.execute │ +│ • TaskRef, OutputRef, InputRef - Reference types │ +│ • ParallelBuilder, SwitchBuilder - Control flow helpers │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Worker Framework (lib/conductor/worker/) │ +│ • TaskRunner - Polling and execution │ +│ • TaskHandler - Worker orchestration │ +│ • WorkerModule - Mixin for class-based workers │ +│ • Events - Task lifecycle hooks │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ High-Level Clients (lib/conductor/client/) │ +│ • WorkflowClient, TaskClient, MetadataClient │ +│ • WorkflowExecutor - Synchronous execution │ +│ • OrkesClients - Factory for all clients │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ Resource APIs (lib/conductor/http/api/) │ +│ • WorkflowResourceApi, TaskResourceApi, etc. │ +│ • Direct mapping to Conductor REST endpoints │ +└─────────────────────────────────────────────────────────────┘ + │ +┌─────────────────────────────▼───────────────────────────────┐ +│ HTTP Transport (lib/conductor/http/) │ +│ • ApiClient - Auth, serialization, dispatch │ +│ • RestClient - Faraday-based HTTP client │ +│ • Models - 50+ request/response models │ +└─────────────────────────────────────────────────────────────┘ +``` + +--- + +## Worker Framework (Detailed) + +The worker framework provides multi-threaded task execution with a comprehensive event system for interceptors and metrics. + +### Component Hierarchy + +``` +┌─────────────────────────────────────────────────────────────────────┐ +│ User Code │ +│ (Worker classes, Worker.define blocks) │ +└─────────────────────────────────────────────────────────────────────┘ + │ + ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ TaskHandler │ +│ • Discovers workers (registry + auto-scan) │ +│ • Resolves configuration (3-tier hierarchy) │ +│ • Creates one Thread per worker type │ +│ • Manages lifecycle (start/stop/join) │ +│ • Aggregates events/metrics │ +└─────────────────────────────────────────────────────────────────────┘ + │ + ┌─────────────┼─────────────┐ + ▼ ▼ ▼ +┌─────────────────────┐ ┌─────────────────────┐ ┌─────────────────────┐ +│ TaskRunner │ │ TaskRunner │ │ TaskRunner │ +│ (Worker A) │ │ (Worker B) │ │ (Worker C) │ +│ │ │ │ │ │ +│ • ThreadPoolExecutor│ │ • ThreadPoolExecutor│ │ • ThreadPoolExecutor│ +│ • Batch polling │ │ • Batch polling │ │ • Batch polling │ +│ • Adaptive backoff │ │ • Adaptive backoff │ │ • Adaptive backoff │ +│ • Event publishing │ │ • Event publishing │ │ • Event publishing │ +└─────────────────────┘ └─────────────────────┘ └─────────────────────┘ +``` + +### Multi-Threading Model + +```ruby +# Each TaskRunner has its own ThreadPoolExecutor +class TaskRunner + def initialize(worker, configuration:, event_dispatcher:) + @worker = worker + @executor = Concurrent::ThreadPoolExecutor.new( + min_threads: 1, + max_threads: worker.thread_count, # Configurable per worker + max_queue: 0, # Synchronous handoff + fallback_policy: :caller_runs + ) + end + + def run + while @running + # 1. Check capacity + available_slots = @worker.thread_count - @running_tasks.size + + # 2. Batch poll for tasks + tasks = batch_poll(available_slots) + + # 3. Submit each task to thread pool + tasks.each do |task| + future = @executor.post { execute_and_update(task) } + @running_tasks << future + end + + # 4. Adaptive backoff for empty polls + apply_backoff if tasks.empty? + end + end +end +``` + +### Event System (Interceptors) + +The event system allows hooking into task lifecycle for logging, metrics, and custom behavior: + +```ruby +# Event Types +module Conductor::Worker::Events + PollStarted # Fired before polling + PollCompleted # Fired after successful poll + PollFailure # Fired on poll error + TaskExecutionStarted # Fired before task execution + TaskExecutionCompleted # Fired after successful execution + TaskExecutionFailure # Fired on execution error + TaskUpdateFailure # CRITICAL: Fired when result update fails +end + +# Custom Event Listener (Interceptor) +class MyInterceptor + def on_poll_started(event) + puts "Polling for #{event.task_type}..." + end + + def on_task_execution_started(event) + puts "Starting task #{event.task_id}" + end + + def on_task_execution_completed(event) + puts "Task #{event.task_id} completed in #{event.duration_ms}ms" + end + + def on_task_execution_failure(event) + puts "Task #{event.task_id} FAILED: #{event.cause.message}" + # Send to error tracking service + ErrorTracker.capture(event.cause, context: { task_id: event.task_id }) + end +end + +# Register listener +handler = TaskHandler.new( + configuration: config, + event_listeners: [MyInterceptor.new] +) +``` + +### Event Dispatcher (Thread-Safe) + +```ruby +class SyncEventDispatcher + def initialize + @listeners = Hash.new { |h, k| h[k] = [] } + @mutex = Mutex.new + end + + def register(event_type, listener) + @mutex.synchronize do + @listeners[event_type] << listener + end + end + + def publish(event) + listeners = @mutex.synchronize { @listeners[event.class].dup } + listeners.each do |listener| + begin + listener.call(event) + rescue StandardError => e + # Listener failure is isolated - never breaks the worker + warn "[Conductor] Event listener error: #{e.message}" + end + end + end +end +``` + +### Metrics Collection + +`MetricsCollector.create` returns a collector that emits the canonical +(harmonized) metric surface: + +```ruby +metrics = Conductor::Worker::Telemetry::MetricsCollector.create(backend: :prometheus) +``` + +See [docs/METRICS_AND_INTERCEPTORS.md](docs/METRICS_AND_INTERCEPTORS.md) for +the full metrics catalog and label reference. + +### Worker Configuration (3-Tier Hierarchy) + +Configuration is resolved in order of priority: + +1. **Worker-specific environment variable** (highest priority) + ```bash + CONDUCTOR_WORKER_PROCESS_ORDER_POLL_INTERVAL=200 + ``` + +2. **Global worker environment variable** + ```bash + CONDUCTOR_WORKER_ALL_POLL_INTERVAL=100 + ``` + +3. **Code-level default** (lowest priority) + ```ruby + worker_task 'process_order', poll_interval: 100 + ``` + +### Configuration Properties + +| Property | Type | Default | Description | +|----------|------|---------|-------------| +| `poll_interval` | Integer | 100 | Polling interval in milliseconds | +| `thread_count` | Integer | 1 | Max concurrent tasks per worker | +| `domain` | String | nil | Task domain for isolation | +| `worker_id` | String | auto | Unique worker identifier | +| `poll_timeout` | Integer | 100 | Server-side long poll timeout (ms) | +| `register_task_def` | Boolean | false | Auto-register task definition | +| `paused` | Boolean | false | Pause worker (stop polling) | + +### Task Context (Thread-Local) + +```ruby +# Access execution context from anywhere in worker code +def execute(task) + ctx = Conductor::Worker::TaskContext.current + + ctx.add_log("Processing task #{ctx.task_id}") + ctx.add_log("Retry count: #{ctx.retry_count}") + + # Long-running task - set callback + if will_take_long? + ctx.set_callback_after(60) # Check back in 60 seconds + return TaskInProgress.new(output: { status: 'processing' }) + end + + { result: 'success' } +end +``` + +--- + +## Directory Structure + +``` +lib/conductor/ +├── version.rb # VERSION constant +├── configuration.rb # Configuration class +├── exceptions.rb # Exception hierarchy +├── client/ # High-level client facades +│ ├── workflow_client.rb +│ ├── task_client.rb +│ ├── metadata_client.rb +│ └── ... +├── http/ +│ ├── api/ # Resource API classes (17) +│ │ ├── workflow_resource_api.rb +│ │ ├── task_resource_api.rb +│ │ └── ... +│ ├── models/ # HTTP models (50+) +│ │ ├── workflow_def.rb +│ │ ├── workflow_task.rb +│ │ ├── task.rb +│ │ └── ... +│ ├── api_client.rb # Auth + serialization +│ └── rest_client.rb # Faraday HTTP client +├── orkes/ # Orkes Cloud specific +│ ├── orkes_clients.rb # Main factory +│ └── models/ +├── worker/ # Worker framework +│ ├── task_runner.rb # Polling loop + ThreadPoolExecutor +│ ├── task_handler.rb # Worker management +│ ├── worker.rb # Worker module +│ ├── worker_config.rb # Configuration resolver +│ ├── worker_registry.rb # Global worker registry +│ ├── task_context.rb # Thread-local context +│ ├── task_in_progress.rb # Long-running task signal +│ ├── events/ # Event system +│ │ ├── conductor_event.rb # Base event class +│ │ ├── task_runner_events.rb # All event types +│ │ ├── sync_event_dispatcher.rb # Thread-safe dispatcher +│ │ ├── listeners.rb # Listener protocol +│ │ └── listener_registry.rb # Registration helper +│ └── telemetry/ # Metrics +│ ├── metrics_collector.rb # MetricsCollector class + NullBackend +│ └── prometheus_backend.rb # PrometheusBackend + MetricsServer +└── workflow/ + ├── dsl/ # Workflow DSL + │ ├── workflow_builder.rb # Core DSL engine (~1000 lines) + │ ├── workflow_definition.rb # Wrapper class + │ ├── task_ref.rb # Task reference + │ ├── output_ref.rb # Output reference (task[:field]) + │ ├── input_ref.rb # Input reference (wf[:param]) + │ ├── parallel_builder.rb # parallel do...end + │ └── switch_builder.rb # decide do...end + ├── llm/ # LLM helper classes + │ ├── chat_message.rb + │ ├── tool_call.rb + │ ├── tool_spec.rb + │ └── embedding_model.rb + ├── task_type.rb # Task type constants + ├── timeout_policy.rb + └── workflow_executor.rb +``` + +--- + +## Key Files to Understand + +### Workflow DSL (Most Important) + +1. **`lib/conductor/workflow/dsl/workflow_builder.rb`** (~1000 lines) + - Core DSL engine with all task methods + - `simple`, `http`, `wait`, `terminate`, `sub_workflow` + - `parallel`, `decide`, `loop_over`, `when_true/when_false` + - LLM tasks: `llm_chat`, `llm_embed`, `generate_image`, etc. + - Value resolution: `OutputRef`, `InputRef` → expression strings + +2. **`lib/conductor/workflow/dsl/workflow_definition.rb`** + - Wrapper class returned by `Conductor.workflow` + - Provides `.register()`, `.execute()`, `.call()` methods + - Delegates to WorkflowExecutor for execution + +3. **`lib/conductor/workflow/dsl/task_ref.rb`** + - Stores task metadata during DSL evaluation + - Converts to `WorkflowTask` model for serialization + - Supports `[]` operator for output references + +### Worker Framework + +1. **`lib/conductor/worker/task_runner.rb`** + - Main polling loop with adaptive backoff + - ThreadPoolExecutor for concurrent task execution + - Event publishing for lifecycle hooks + +2. **`lib/conductor/worker/worker.rb`** + - `WorkerModule` mixin for class-based workers + - `worker_task` class method for registration + - `Conductor::Worker.define` for block-based workers + +3. **`lib/conductor/worker/events/`** + - Event classes for task lifecycle + - SyncEventDispatcher for thread-safe event publishing + - ListenerRegistry for listener management + +### HTTP Layer + +1. **`lib/conductor/http/api_client.rb`** + - Token management with TTL-based refresh + - Request serialization, response deserialization + - Retry logic for 401/403 errors + +2. **`lib/conductor/http/models/workflow_def.rb`** + - WorkflowDef model with all workflow properties + - Used for registration and serialization + +--- + +## Common Patterns + +### Creating a Workflow + +```ruby +workflow = Conductor.workflow :my_workflow, version: 1, executor: executor do + user = simple :get_user, user_id: wf[:user_id] + simple :send_email, email: user[:email] + output result: user[:name] +end + +workflow.register(overwrite: true) +result = workflow.execute(input: { user_id: 123 }) +``` + +### Task Reference Flow + +``` +DSL Method Call TaskRef Created WorkflowTask Generated +───────────────────────────────────────────────────────────────────────── +simple :foo, x: wf[:y] → TaskRef(ref: 'foo_ref') → WorkflowTask( + task_name: 'foo' name: 'foo' + inputs: {...} type: 'SIMPLE' + input_parameters: {...} + ) +``` + +### Output References + +```ruby +task[:field] # → OutputRef → "${task_ref.output.field}" +task[:nested][:path] # → OutputRef → "${task_ref.output.nested.path}" +wf[:param] # → InputRef → "${workflow.input.param}" +wf.var(:counter) # → InputRef → "${workflow.variables.counter}" +``` + +--- + +## Testing + +### Test Structure + +``` +spec/ +├── conductor/ +│ ├── workflow/ +│ │ ├── dsl/ +│ │ │ └── workflow_builder_spec.rb # 52 DSL tests +│ │ └── llm_tasks_spec.rb # LLM helper tests +│ ├── client/ +│ ├── http/ +│ └── worker/ +└── integration/ # Requires live server +``` + +### Running Tests + +```bash +# Run all unit tests (REQUIRED after every change) +bundle exec rspec spec/conductor/ + +# Run DSL tests specifically +bundle exec rspec spec/conductor/workflow/dsl/ + +# Run worker tests +bundle exec rspec spec/conductor/worker/ + +# Run with documentation format +bundle exec rspec --format documentation + +# Integration tests (requires Conductor server) +CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ +``` + +--- + +## Making Changes + +### Adding a New Task Type + +1. Add constant to `lib/conductor/workflow/task_type.rb` +2. Add DSL method to `lib/conductor/workflow/dsl/workflow_builder.rb` +3. Handle conversion in `lib/conductor/workflow/dsl/task_ref.rb` +4. Add tests in `spec/conductor/workflow/dsl/workflow_builder_spec.rb` +5. Update examples in `examples/workflow_dsl.rb` +6. Run linter: `bundle exec rubocop -a` +7. Run tests: `bundle exec rspec spec/conductor/` + +### Adding a New API Endpoint + +1. Add method to appropriate Resource API in `lib/conductor/http/api/` +2. Add corresponding method to high-level client in `lib/conductor/client/` +3. Add tests in `spec/conductor/http/api/` and `spec/conductor/client/` +4. Run linter: `bundle exec rubocop -a` +5. Run tests: `bundle exec rspec spec/conductor/` + +### Modifying Worker Behavior + +1. Review `docs/design/WORKER_DESIGN.md` for detailed design +2. Modify `lib/conductor/worker/task_runner.rb` for polling behavior +3. Modify `lib/conductor/worker/worker.rb` for worker definition +4. Add tests in `spec/conductor/worker/` +5. Run linter: `bundle exec rubocop -a` +6. Run tests: `bundle exec rspec spec/conductor/` + +### Adding Event Listeners / Interceptors + +1. Create class implementing listener methods (`on_poll_started`, `on_task_execution_completed`, etc.) +2. Register with TaskHandler via `event_listeners:` option +3. Add tests in `spec/conductor/worker/events/` + +--- + +## Important Conventions + +- **Snake case** for methods and variables +- **Keyword arguments** for optional parameters +- **Blocks** for control flow (`parallel do`, `decide do`) +- **Symbol-to-string** task names are auto-converted +- **Output references** use `[]` operator (`task[:field]`) +- **Input references** use `wf[:param]` syntax +- **Thread safety** - Use Mutex for shared state in event system + +--- + +## Dependencies + +**Runtime:** +- `faraday ~> 2.0` - HTTP client +- `faraday-net_http_persistent ~> 2.0` - Connection pooling +- `faraday-retry ~> 2.0` - Automatic retries +- `concurrent-ruby ~> 1.2` - Thread pool executor + +**Development:** +- `rspec ~> 3.0` - Testing +- `webmock ~> 3.0` - HTTP mocking +- `rubocop ~> 1.0` - Linting diff --git a/CHANGELOG.md b/CHANGELOG.md index 8331c4e..4595cac 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,25 +5,11 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). -## [Unreleased] - -### Added - -- Port all 19 requested Python agent examples, with integration tests that execute the examples directly and CI using Conductor OSS playback plus the shared recording verification action. -- Add plan-and-compile agents with planner/fallback configuration, external worker declarations, explicit tool input schemas, SSE event callbacks, and structured approval responses. -- Compare every example's configuration to Python-generated fixtures and validate the agent schema. - -### Fixed - -- Execute tasks claimed by `update-v2` within the existing worker slot instead of leaving them in progress. -- Register individual task definitions without nesting the metadata request array. -- Inherit tool credentials through agent trees and register callable routers, tool guardrails, and hoisted conditional handoffs. - ## [0.1.0] ### Added -- Canonical (harmonized) metrics as the sole metric surface +- Canonical (harmonized) metrics as the sole metric surface -- [details](docs/METRICS_AND_INTERCEPTORS.md#detailed-technical-notes----unreleased) - Bounded `uri` label on `http_api_client_request_seconds`: uses path templates (e.g. `/workflow/{workflowId}`) instead of fully-resolved paths, preventing metric cardinality explosion - `WorkflowStatusProbe` in harness: opt-in probe (via `HARNESS_PROBE_RATE_PER_SEC`) that exercises UUID-bearing endpoints to validate template URI metrics @@ -50,11 +36,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - Task classes: `SimpleTask`, `SwitchTask`, `ForkTask`, `JoinTask`, `DoWhileTask`, `HttpTask`, `SubWorkflowTask`, `WaitTask`, `TerminateTask`, `SetVariableTask`, `DynamicForkTask`, `JavascriptTask`, `JsonJqTask`, `EventTask`, `HttpPollTask`, `DynamicTask`, `HumanTask`, `StartWorkflowTask`, `KafkaPublishTask`, `WaitForWebhookTask` - LLM task classes: `LlmChatCompleteTask`, `LlmTextCompleteTask`, `LlmGenerateEmbeddingsTask`, `LlmIndexTextTask`, `LlmIndexDocumentTask`, `LlmSearchIndexTask`, `LlmQueryEmbeddingsTask`, `LlmStoreEmbeddingsTask`, `LlmSearchEmbeddingsTask`, `GenerateImageTask`, `GenerateAudioTask`, `GetDocumentTask`, `ListMcpToolsTask`, `CallMcpToolTask` -### Fixed - -- `SchedulerResourceApi#pause_schedule` / `#resume_schedule` now work against both Conductor server families. The client sends `PUT` first and falls back to `GET` on a `405` -- and only on a `405`. OSS Conductor maps these two per-schedule routes `@PutMapping`-only, so the previous `GET`-only calls failed there outright; Orkes Conductor accepts both verbs as of the dual `@RequestMapping(method = {GET, PUT})` added in 2026-07, and is `GET`-only in deployments older than that. `pause_all_schedules` / `resume_all_schedules` remain `GET`, which is how both families map those admin endpoints. Matches the python-sdk, go-sdk, javascript-sdk, csharp-sdk and rust-sdk clients; `spec/conductor/http/api/scheduler_resource_api_spec.rb` pins the whole contract -- `Conductor::AuthenticationSettings` is no longer referenced as `Conductor::Configuration::AuthenticationSettings`, which raised `NameError: uninitialized constant`. The class has always been defined directly under `Conductor`. Fixed in `RactorTaskRunner`'s in-Ractor configuration rebuild (where it was a live failure) and in the `Conductor` / `OrkesClients` doc comments (where it told users to write the broken form) - ### Migration Guide **Before (old DSL):** @@ -78,26 +59,6 @@ end ### Added -- **Agents** (`require 'conductor/agents'`) - Ruby port of the Python SDK's agents package, same `agentConfig` on the wire - - `tool def` DSL: types from keyword defaults, secrets from `secret('...')` literals, `describe`, `requires_approval`, module scoping, RubyLLM::Tool adapter - - `Agent` with `add_tool`, `add_agent`, `hands_off_to`, `redact`, `stop_when`, `stop_after`, `on_approval`, `>>`; guardrails, termination conditions, handoffs, callbacks, memory, prompt templates - - `ConfigSerializer` verified against the Python SDK's 19 golden configs and `agent-schema.json` - - `AgentRuntime`: `call_sync`, `call_async` (SSE streaming with reconnect, polling fallback), `deploy`, `serve`; `Execution`, `ApprovalRequest` - - Tool workers registered with Python's task definition defaults; `_termination`, custom guardrail and callback workers - - `AgentResourceApi` / `AgentClient` for `/api/agent/*`, `OrkesClients#get_agent_client` - - `Task#runtime_metadata` (wire-only secret values), `TaskDef#runtime_metadata` (declared secret names), `TaskDef#enforce_schema` - - `TaskResourceApi#update_task_v2`; `Worker` option `lease_extend_enabled` - - Agent examples run against Conductor with shared LLM recordings and playback validation - -### Changed - -- `Configuration` caches the auth token per instance (two configurations no longer share a token); the class-level `Configuration.auth_token` accessors remain as a deprecated shim -- `TaskRunner` posts task results to `POST /tasks/update-v2` and falls back to `POST /tasks` once when the server does not serve it (Python SDK parity) - -### Removed - -- Unused `vcr` development dependency (`json_schemer` added for agent contract tests) - - **Core Infrastructure** - Configuration with environment variable support - Authentication (token management, TTL refresh, exponential backoff) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index e543c2c..5255a2a 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,48 +1,226 @@ -# Contributing +# Contributing to Conductor Ruby SDK -Use Ruby 3.0+ and Bundler. From the repository root: +Thank you for your interest in contributing to the Conductor Ruby SDK! This document provides guidelines and instructions for contributing. + +## Code of Conduct + +Please be respectful and constructive in all interactions. We welcome contributors of all experience levels. + +## Getting Started + +### Prerequisites + +- Ruby 2.6+ (Ruby 3.2+ recommended) +- Bundler +- Git + +### Setup + +1. Fork the repository on GitHub +2. Clone your fork: + ```bash + git clone https://github.com/YOUR_USERNAME/ruby-sdk.git + cd ruby-sdk + ``` + +3. Install dependencies: + ```bash + bundle install + ``` + +4. Run tests to verify setup: + ```bash + bundle exec rspec spec/conductor/ + ``` + +## Development Workflow + +### Branch Naming + +- `feature/description` - New features +- `fix/description` - Bug fixes +- `docs/description` - Documentation updates +- `refactor/description` - Code refactoring + +### Making Changes + +1. Create a new branch: + ```bash + git checkout -b feature/my-new-feature + ``` + +2. Make your changes + +3. Run tests: + ```bash + bundle exec rspec spec/conductor/ + ``` + +4. Run linting: + ```bash + bundle exec rubocop + ``` + +5. Commit your changes: + ```bash + git commit -m "Add feature: description of changes" + ``` + +6. Push to your fork: + ```bash + git push origin feature/my-new-feature + ``` + +7. Create a Pull Request + +## Testing + +### Running Tests ```bash -bundle install +# All unit tests bundle exec rspec spec/conductor/ -bundle exec rubocop -bundle exec ruby -Ilib -rconductor -e 'puts Conductor::VERSION' -gem build conductor_ruby.gemspec + +# Specific test file +bundle exec rspec spec/conductor/client/workflow_client_spec.rb + +# With documentation format +bundle exec rspec spec/conductor/ --format documentation + +# Integration tests (requires Conductor server) +CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ ``` -Add regression tests for fixes and tests for new behavior. Keep coverage at least -at its current level. Put unit tests in `spec/conductor/` and server tests in -`spec/integration/`. Update runnable examples when the public API changes. +### Writing Tests -For local OSS integration tests (requires Docker and Ruby): +- Place unit tests in `spec/conductor/` +- Place integration tests in `spec/integration/` +- Use descriptive test names +- Follow existing test patterns -```bash -scripts/run-integration-oss.sh +Example test structure: +```ruby +RSpec.describe Conductor::Client::WorkflowClient do + describe '#start' do + context 'when workflow exists' do + it 'returns a workflow ID' do + # test implementation + end + end + + context 'when workflow does not exist' do + it 'raises an ApiError' do + # test implementation + end + end + end +end ``` -Add agent examples in `examples/agents/` and run them against Conductor with -LLM recording enabled: +## Code Style + +We use RuboCop for code style enforcement. Key guidelines: +- Use 2 spaces for indentation +- Use snake_case for methods and variables +- Use CamelCase for classes and modules +- Add documentation comments for public methods +- Keep methods small and focused + +Run RuboCop to check your code: ```bash -bundle exec ruby -Ilib examples/agents/my_example.rb +bundle exec rubocop + +# Auto-fix safe issues +bundle exec rubocop -a ``` -Collect the recording JSON files and open a PR adding them to `llm-recordings/` -in [conductor-oss/conductor](https://github.com/conductor-oss/conductor). -Register the Ruby example in `examples/agents/catalog.rb` so CI replays it. +## Documentation -Without local Ruby, run unit tests and lint in Docker: +- Update README.md for user-facing changes +- Add YARD documentation for new public methods +- Update CHANGELOG.md for notable changes +- Include examples for new features -```bash -docker run --rm -v "$PWD":/app:z -w /app \ - -v ruby-sdk-bundle:/usr/local/bundle:z ruby:3.3 \ - bash -lc 'bundle install --quiet && bundle exec rspec spec/conductor/ && bundle exec rubocop' +### YARD Documentation Example + +```ruby +# Starts a workflow execution +# +# @param name [String] The workflow name +# @param input [Hash] The workflow input (default: {}) +# @param version [Integer, nil] The workflow version (optional) +# @return [String] The workflow ID +# @raise [ApiError] If the workflow doesn't exist +# +# @example Start a simple workflow +# workflow_id = client.start('my_workflow', input: { key: 'value' }) +# +def start(name, input: {}, version: nil) + # implementation +end ``` -The load harness runs with `bundle exec ruby harness/main.rb` using the same -server/auth environment variables as the SDK. It continuously starts workflows; -`HARNESS_WORKFLOWS_PER_SEC` defaults to 2. Kubernetes templates are in -[harness/manifests](harness/manifests/). +## Pull Request Guidelines + +### Before Submitting + +- [ ] Tests pass locally +- [ ] RuboCop passes (or issues are intentional) +- [ ] Documentation is updated +- [ ] CHANGELOG.md is updated (if applicable) +- [ ] Commits are clean and well-described + +### PR Description + +Include: +- Summary of changes +- Motivation/context +- How to test +- Screenshots (if UI-related) + +### Review Process + +1. Automated CI checks run +2. Maintainers review code +3. Address feedback +4. Merge when approved + +## Release Process + +Releases are managed by maintainers: + +1. Update version in `lib/conductor/version.rb` +2. Update CHANGELOG.md +3. Create GitHub release with tag +4. CI automatically publishes to RubyGems + +## Project Structure + +``` +lib/ +├── conductor.rb # Main entry point +├── conductor/ +│ ├── version.rb # Version constant +│ ├── configuration.rb # Configuration class +│ ├── exceptions.rb # Exception classes +│ ├── client/ # High-level clients +│ ├── http/ +│ │ ├── api/ # Resource API classes +│ │ ├── models/ # Model classes +│ │ ├── api_client.rb # HTTP client wrapper +│ │ └── rest_client.rb # Faraday client +│ ├── orkes/ # Orkes-specific code +│ ├── worker/ # Worker framework +│ └── workflow/ # Workflow DSL +``` + +## Getting Help + +- Open an issue for bugs or feature requests +- Join [Conductor Slack](https://join.slack.com/t/orkes-conductor/shared_invite/zt-2vdbx239s-Eacdyqya9giNLHfrCavfaA) +- Check existing issues and PRs + +## License -For releases, update `lib/conductor/version.rb` and `CHANGELOG.md`, then create a -GitHub release with a tag; CI publishes the gem. +By contributing, you agree that your contributions will be licensed under the Apache 2.0 License. diff --git a/DESIGN.md b/DESIGN.md index 86c4ca1..c17130b 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -242,27 +242,8 @@ lib/conductor/ ├── task_type.rb # Task type constants ├── timeout_policy.rb └── workflow_executor.rb -lib/conductor/agents.rb # require 'conductor/agents' -lib/conductor/agents/ # Agents (see docs/agents/README.md) -├── agent.rb, tool_def.rb, tools.rb # definition layer + `tool def` DSL -├── guardrail.rb, termination.rb, handoff.rb, callback_handler.rb, memory.rb, prompt_template.rb -├── config_serializer.rb # -> agentConfig, identical to the Python SDK -└── runtime/ # AgentRuntime, SseClient, Execution, ApprovalRequest, ToolRegistry, Dispatch, Secrets ``` -## Agents - -`Conductor::Agents` ports the Python SDK's `conductor.ai.agents` package. An `Agent` tree is -serialized by `ConfigSerializer` to the same `agentConfig` JSON Python sends; the server compiles -it into a workflow and runs the LLM loop. `AgentRuntime` starts the execution -(`POST /api/agent/start`), registers a Conductor worker for every task the server lists in -`requiredWorkers` (the user's `tool def` tools plus `_termination`, custom guardrails and -callbacks), and follows the run over SSE (`GET /api/agent/stream/{id}`, polling fallback). -Secrets travel as `TaskDef.runtimeMetadata` names and come back as `Task.runtimeMetadata` -values, read inside tools with `secret('NAME')`. The transport for `/api/agent/*` is -`AgentResourceApi` / `AgentClient` like every other resource. Decisions and the verified wire -contract are in `docs/design/AGENTS_IMPLEMENTATION_PLAN.md`. - ## Dependencies ```ruby diff --git a/README.md b/README.md index 16db78b..8f6c97d 100644 --- a/README.md +++ b/README.md @@ -1,45 +1,517 @@ # Conductor Ruby SDK -Ruby clients, workflow definitions, workers, and agents for [Conductor](https://github.com/conductor-oss/conductor). -Requires Ruby 3.0 or newer. +Official Ruby SDK for [Conductor OSS](https://github.com/conductor-oss/conductor) - a durable workflow orchestration engine. -## Install +[![Gem Version](https://badge.fury.io/rb/conductor_ruby.svg)](https://badge.fury.io/rb/conductor_ruby) +[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE) + +## Features + +- **Full Feature Parity** with Python SDK +- **Ruby-Idiomatic Workflow DSL** - Clean block-based syntax with 25+ task types +- **Worker Framework** - Multi-threaded task execution with class-based and block-based workers +- **LLM/AI Tasks** - Chat completion, embeddings, RAG, image/audio generation +- **Orkes Cloud Support** - Authentication, secrets, integrations, prompts +- **Comprehensive Testing** - 400+ unit tests, 110 integration tests + +## Installation + +Add to your Gemfile: ```ruby gem 'conductor_ruby' ``` -Set `CONDUCTOR_SERVER_URL` to your server's API URL (for example, -`http://localhost:8080/api`). For authenticated servers, also set -`CONDUCTOR_AUTH_KEY` and `CONDUCTOR_AUTH_SECRET`. +Or install directly: -## Start a workflow +```bash +gem install conductor_ruby +``` + +## Quick Start -This starts an existing workflow definition on your server: +### Hello World ```ruby require 'conductor' -client = Conductor::Client::WorkflowClient.new(Conductor::Configuration.new) -workflow_id = client.start('my_workflow', input: { 'name' => 'Ruby' }) -puts workflow_id +# Configuration (reads CONDUCTOR_SERVER_URL from environment) +config = Conductor::Configuration.new + +# Create clients +clients = Conductor::Orkes::OrkesClients.new(config) +executor = clients.get_workflow_executor + +# Define a worker +class GreetWorker + include Conductor::Worker::WorkerModule + worker_task 'greet' + + def execute(task) + name = get_input(task, 'name', 'World') + { 'result' => "Hello, #{name}!" } + end +end + +# Build workflow using new DSL +workflow = Conductor.workflow :greetings, version: 1, executor: executor do + greet = simple :greet, name: wf[:name] + output result: greet[:result] +end + +# Register and execute +workflow.register(overwrite: true) + +# Start workers +runner = Conductor::Worker::TaskRunner.new(config) +runner.register_worker(GreetWorker.new) +runner.start + +# Execute workflow +result = workflow.execute(input: { 'name' => 'Ruby' }, wait_for_seconds: 30) +puts "Result: #{result.output['result']}" # => "Hello, Ruby!" + +runner.stop +``` + +## Workflow DSL + +The SDK provides a clean, Ruby-idiomatic DSL for building workflows: + +```ruby +workflow = Conductor.workflow :order_processing, version: 1, executor: executor do + # Access workflow inputs with wf[:param] + user = simple :get_user, user_id: wf[:user_id] + + # Reference task outputs with task[:field] + order = simple :validate_order, email: user[:email] + + # HTTP calls + http :call_api, url: 'https://api.example.com', method: :post, body: { id: order[:id] } + + # Parallel execution + parallel do + simple :ship_order, order_id: order[:id] + simple :send_confirmation, email: user[:email] + end + + # Conditional branching + decide order[:region] do + on 'US' do + simple :us_shipping + end + on 'EU' do + simple :eu_shipping + end + otherwise do + terminate :failed, 'Unsupported region' + end + end + + # Set workflow output + output tracking: order[:tracking_number], status: 'completed' +end + +# Register and execute +workflow.register(overwrite: true) +result = workflow.execute(input: { user_id: 123 }, wait_for_seconds: 60) +``` + +### Task Methods Reference + +#### Basic Tasks + +```ruby +# Simple task (worker execution) +result = simple :task_name, input1: 'value', input2: wf[:param] + +# Inline code execution +jq :transform, query: '.items | map(.name)', input: previous[:data] +javascript :compute, script: 'return inputs.a + inputs.b', a: 1, b: 2 + +# Set workflow variables +set_variable :save_state, user_id: user[:id], status: 'active' + +# Human/manual task +human :approval, display_name: 'Manager Approval', form_template: 'approval_form' +``` + +#### HTTP Tasks + +```ruby +# HTTP request +http :call_api, + url: 'https://api.example.com/users', + method: :post, + headers: { 'Authorization' => 'Bearer ${workflow.secrets.api_token}' }, + body: { name: wf[:name], email: wf[:email] } + +# HTTP polling (wait for condition) +http_poll :wait_for_ready, + url: 'https://api.example.com/status/${workflow.input.job_id}', + method: :get, + termination_condition: '$.status == "ready"', + polling_interval: 5, + polling_strategy: :fixed +``` + +#### Control Flow + +```ruby +# Parallel execution (fork/join) +parallel do + simple :branch_a + simple :branch_b + simple :branch_c +end + +# Conditional branching +decide order[:status] do + on 'pending' do + simple :process_pending + end + on 'approved' do + simple :process_approved + end + otherwise do + simple :handle_unknown + end +end + +# Conditional shortcuts +when_true user[:is_premium] do + simple :apply_discount +end + +when_false order[:validated] do + terminate :failed, 'Order validation failed' +end + +# Loop over items +loop_over users[:list], as: :user do + simple :process_user, user_id: iteration[:user][:id] +end + +# Do-while loop +do_while :retry_loop, condition: '${retry_ref.output.success} == false' do + simple :retry_operation +end +``` + +#### Sub-workflows + +```ruby +# Call another workflow +sub_workflow :process_order, + workflow_name: 'order_processor', + version: 2, + input: { order_id: wf[:order_id] } + +# Start workflow (fire-and-forget) +start_workflow :trigger_notification, + workflow_name: 'send_notifications', + input: { user_id: user[:id] } + +# Inline sub-workflow definition +inline_workflow :nested_process do + simple :step1 + simple :step2 +end +``` + +#### Wait and Events + +```ruby +# Wait for duration +wait :pause, duration: '30s' # or '5m', '1h', '2d' + +# Wait until specific time +wait :scheduled, until: '2024-12-25T00:00:00Z' + +# Wait for external webhook +wait_for_webhook :external_callback, + matches: { 'type' => 'payment', 'order_id' => '${workflow.input.order_id}' } + +# Publish event +event :notify, sink: 'conductor:workflow_events', payload: { status: 'completed' } +``` + +#### Termination + +```ruby +# Complete workflow +terminate :success, 'Processing completed successfully' + +# Fail workflow +terminate :failed, 'Validation error: missing required field' +``` + +#### Dynamic Tasks + +```ruby +# Dynamic task name (resolved at runtime) +dynamic :run_handler, task_to_execute: wf[:handler_name] + +# Dynamic fork (parallel tasks determined at runtime) +dynamic_fork :process_all, + tasks_input: wf[:items], + task_name: 'process_item' +``` + +### LLM/AI Tasks + +```ruby +workflow = Conductor.workflow :ai_assistant, executor: executor do + # Chat completion (messages auto-converted from simple format) + response = llm_chat :chat, + provider: 'openai', + model: 'gpt-4', + messages: [ + { role: :system, message: 'You are a helpful assistant.' }, + { role: :user, message: wf[:question] } + ], + temperature: 0.7 + + # Text completion + llm_text :complete, + provider: 'anthropic', + model: 'claude-3-sonnet', + prompt: 'Summarize: ${workflow.input.text}' + + # Generate embeddings + embeddings = llm_embeddings :embed, + provider: 'openai', + model: 'text-embedding-3-small', + text: wf[:document] + + # Store embeddings in vector DB + llm_store_embeddings :store, + provider: 'pinecone', + index: 'documents', + embeddings: embeddings[:embeddings], + metadata: { doc_id: wf[:doc_id] } + + # Search embeddings + llm_search_embeddings :search, + provider: 'pinecone', + index: 'documents', + query: wf[:search_query], + max_results: 10 + + # Generate image + generate_image :create_image, + provider: 'openai', + model: 'dall-e-3', + prompt: 'A sunset over mountains', + size: '1024x1024' + + # Generate audio (text-to-speech) + generate_audio :speak, + provider: 'openai', + model: 'tts-1', + text: response[:content], + voice: 'nova' + + # MCP (Model Context Protocol) integration + tools = list_mcp_tools :get_tools, server_name: 'my_mcp_server' + + call_mcp_tool :use_tool, + server_name: 'my_mcp_server', + tool_name: 'search_documents', + arguments: { query: wf[:query] } + + output answer: response[:content] +end +``` + +### Output References + +The DSL uses a clean syntax for referencing outputs: + +```ruby +# Workflow input reference +wf[:user_id] # => '${workflow.input.user_id}' + +# Task output reference +task[:field] # => '${task_ref.output.field}' +task[:nested][:path] # => '${task_ref.output.nested.path}' + +# Loop iteration references (inside loop_over) +iteration[:current_item] # Current item being processed +iteration[:index] # Current index (0-based) +iteration[:user][:name] # If `as: :user` specified ``` ## Examples -- [Workflow DSL](examples/workflow_dsl.rb) -- [Worker configuration](examples/worker_configuration_example.rb) -- [Task context](examples/task_context_example.rb) and [event listeners](examples/task_listener_example.rb) -- [Agents](examples/agents/): tools, streaming, approvals, guardrails, and teams +The `examples/` directory contains comprehensive examples: + +| Example | Description | +|---------|-------------| +| [`helloworld/`](examples/helloworld/) | Simplest complete example - worker + workflow + execution | +| [`workflow_dsl.rb`](examples/workflow_dsl.rb) | Comprehensive new DSL showcase | +| [`simple_worker.rb`](examples/simple_worker.rb) | Worker patterns: class-based, block-based, error handling | +| [`kitchensink.rb`](examples/kitchensink.rb) | All major task types using new DSL | +| [`dynamic_workflow.rb`](examples/dynamic_workflow.rb) | Create and execute workflows at runtime | +| [`workflow_ops.rb`](examples/workflow_ops.rb) | Lifecycle operations: pause, resume, restart, retry | +| [`agentic_workflows/`](examples/agentic_workflows/) | LLM chat and AI workflow examples | + +Run examples: + +```bash +# Set environment variables +export CONDUCTOR_SERVER_URL=http://localhost:8080/api +# For Orkes Cloud: +# export CONDUCTOR_AUTH_KEY=your_key +# export CONDUCTOR_AUTH_SECRET=your_secret + +# Run hello world +cd examples/helloworld && bundle exec ruby helloworld.rb + +# Run DSL showcase +bundle exec ruby examples/workflow_dsl.rb + +# Run kitchen sink +bundle exec ruby examples/kitchensink.rb +``` + +## Worker Framework + +### Class-Based Workers + +```ruby +class ImageProcessor + include Conductor::Worker::WorkerModule + + worker_task 'process_image', poll_interval: 1, thread_count: 4 + + def execute(task) + url = get_input(task, 'image_url') + # Process image... + + result = Conductor::Http::Models::TaskResult.complete + result.add_output_data('processed_url', processed_url) + result.log('Image processed successfully') + result + end +end +``` + +### Block-Based Workers + +```ruby +worker = Conductor::Worker.define('simple_task') do |task| + input = task.input_data['value'] + { result: input * 2 } # Return hash for automatic TaskResult +end +``` + +### Running Workers + +```ruby +runner = Conductor::Worker::TaskRunner.new(config) +runner.register_worker(ImageProcessor.new) +runner.register_worker(worker) +runner.start(threads: 4) + +# Graceful shutdown +trap('INT') { runner.stop } +sleep while runner.running? +``` + +## Configuration + +### Environment Variables + +```bash +export CONDUCTOR_SERVER_URL=http://localhost:8080/api +export CONDUCTOR_AUTH_KEY=your_key # For Orkes Cloud +export CONDUCTOR_AUTH_SECRET=your_secret # For Orkes Cloud +``` + +### Programmatic + +```ruby +config = Conductor::Configuration.new( + server_api_url: 'https://play.orkes.io/api', + auth_key: 'your_key', + auth_secret: 'your_secret', + auth_token_ttl_min: 45, + verify_ssl: true +) +``` + +## API Coverage -To run an agent example from this checkout, configure an LLM integration on your -Conductor server, then run: +### Resource APIs (17 classes) + +| API | Description | +|-----|-------------| +| WorkflowResourceApi | Workflow execution and management | +| TaskResourceApi | Task polling and updates | +| MetadataResourceApi | Workflow/task definitions | +| SchedulerResourceApi | Scheduled workflows | +| EventResourceApi | Event handlers | +| WorkflowBulkResourceApi | Bulk operations | +| PromptResourceApi | AI prompt templates | +| SecretResourceApi | Secret management | +| IntegrationResourceApi | External integrations | +| + 8 more | Authorization, Users, Groups, Roles, etc. | + +### High-Level Clients (9 classes) + +```ruby +clients = Conductor::Orkes::OrkesClients.new(config) + +workflow_client = clients.get_workflow_client +task_client = clients.get_task_client +metadata_client = clients.get_metadata_client +scheduler_client = clients.get_scheduler_client +prompt_client = clients.get_prompt_client +secret_client = clients.get_secret_client +authorization_client = clients.get_authorization_client +workflow_executor = clients.get_workflow_executor +``` + +## Testing ```bash -bundle install -CONDUCTOR_AGENT_LLM_MODEL=openai/gpt-4o-mini \ - bundle exec ruby -Ilib examples/agents/01_basic_agent.rb +# Unit tests +bundle exec rspec spec/conductor/ + +# Integration tests (requires Conductor server) +CONDUCTOR_SERVER_URL=http://localhost:8080/api bundle exec rspec spec/integration/ ``` -See [CONTRIBUTING.md](CONTRIBUTING.md) for development commands and -[CHANGELOG.md](CHANGELOG.md) for release history. Licensed under [Apache 2.0](LICENSE). +## Requirements + +- Ruby 2.6+ (Ruby 3+ recommended) +- Conductor OSS 3.x or Orkes Cloud + +## Dependencies + +- `faraday ~> 2.0` - HTTP client +- `faraday-net_http_persistent ~> 2.0` - Connection pooling +- `faraday-retry ~> 2.0` - Automatic retries +- `concurrent-ruby ~> 1.2` - Thread-safe concurrency + +## Contributing + +1. Fork the repository +2. Create your feature branch (`git checkout -b feature/amazing-feature`) +3. Run tests (`bundle exec rspec`) +4. Commit your changes (`git commit -m 'Add amazing feature'`) +5. Push to the branch (`git push origin feature/amazing-feature`) +6. Open a Pull Request + +## License + +Apache 2.0 - see [LICENSE](LICENSE) for details. + +## Links + +- [Conductor OSS](https://github.com/conductor-oss/conductor) +- [Orkes Cloud](https://orkes.io) +- [Documentation](https://conductor-oss.org) +- [Python SDK](https://github.com/conductor-sdk/conductor-python) +- [Community Slack](https://join.slack.com/t/orkes-conductor/shared_invite/zt-2vdbx239s-Eacdyqya9giNLHfrCavfaA) diff --git a/docs/agents/README.md b/docs/agents/README.md new file mode 100644 index 0000000..f475a51 --- /dev/null +++ b/docs/agents/README.md @@ -0,0 +1,17 @@ +# Agents + +Durable AI agents on Conductor, written in Ruby. Tools run as Conductor worker +tasks, approvals and schedules wait server-side, and an execution survives a +process restart. + +Requires Ruby 3.0+ and a Conductor server with an LLM provider configured. + +- [Getting started](getting-started.md) +- Concepts: [agents](concepts/agents.md), [tools](concepts/tools.md), [multi-agent](concepts/multi-agent.md), + [guardrails](concepts/guardrails.md), [termination](concepts/termination.md), [callbacks](concepts/callbacks.md), + [stateful](concepts/stateful.md), [streaming and approval](concepts/streaming-hitl.md), + [structured output](concepts/structured-output.md), [runtime modes](concepts/deploy-serve-run.md), + [scheduling](concepts/scheduling.md) +- Reference: [API map](reference/api.md), [runtime](reference/runtime.md), [client](reference/client.md), + [agent fields](reference/agent-definition.md), [wire contract](reference/agent-schema.md) +- [Examples](../../examples/agents/) diff --git a/docs/agents/concepts/agents.md b/docs/agents/concepts/agents.md new file mode 100644 index 0000000..243c5a3 --- /dev/null +++ b/docs/agents/concepts/agents.md @@ -0,0 +1,26 @@ +# Agents + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def get_weather(city: String) + "Weather for #{city}" +end + +agent = Agent.new(name: 'weather', model: 'openai/gpt-4o-mini', + instructions: 'Answer concisely.', tools: [:get_weather]) +puts agent.call_sync('Weather in Seattle?') +``` + +- `name` must match `^[a-zA-Z_][a-zA-Z0-9_-]*$`. +- `model` is `provider/model`; the provider is an integration on the server. + Omit it only for sub-agents that inherit the parent's model. +- `instructions` is a String, a Proc evaluated at compile time, or a `PromptTemplate`. +- Per-call options: `session_id:`, `media:`, `context:`, `idempotency_key:`, + `timeout_seconds:`. Don't mutate a shared agent per request. + +`call_sync` compiles the agent, starts workers for its tools, and returns the +answer. A tool task stuck in `SCHEDULED` means no process is running its worker. + +See [tools](tools.md), [multi-agent](multi-agent.md), [runtime modes](deploy-serve-run.md). diff --git a/docs/agents/concepts/callbacks.md b/docs/agents/concepts/callbacks.md new file mode 100644 index 0000000..a519ba4 --- /dev/null +++ b/docs/agents/concepts/callbacks.md @@ -0,0 +1,18 @@ +# Callbacks + +```ruby +agent.callback(:before_tool) { |**event| logger.info(event[:tool_name]) } + +class Audit < CallbackHandler + def on_agent_end(**event) = logger.info(event) +end +agent.add_callback(Audit.new) +``` + +Positions: `before_agent`, `after_agent`, `before_model`, `after_model`, +`before_tool`, `after_tool`. Handler methods are `on_agent_start`, `on_agent_end`, +`on_model_start`, `on_model_end`, `on_tool_start`, `on_tool_end`. + +Callbacks run in your process and don't change the workflow. Keep them fast, and +don't make them the only record of anything: a restart drops them. Durable side +effects belong in a tool. diff --git a/docs/agents/concepts/deploy-serve-run.md b/docs/agents/concepts/deploy-serve-run.md new file mode 100644 index 0000000..50d6f3b --- /dev/null +++ b/docs/agents/concepts/deploy-serve-run.md @@ -0,0 +1,20 @@ +# Runtime modes + +| Call | Does | +|---|---| +| `runtime.compile(agent)` | Compile; returns `workflowDef` and `requiredWorkers`. | +| `runtime.deploy(agent)` | Compile and register on the server. | +| `runtime.serve(agent)` | Deploy, run tool workers, block until INT/TERM. | +| `runtime.call_sync(agent, prompt)` | Start, run workers, return the answer. | +| `runtime.call_async(agent, prompt)` | Same, but return an `Execution` immediately. | +| `Execution.find(id)` | Reattach to an existing execution. | + +`runtime` is `Conductor::Agents.runtime` (built from ENV) or your own +`AgentRuntime.new(configuration: ...)`. + +Local development: `call_sync`. CI/CD: `compile` to inspect, `deploy` to +release. Production: `serve` in a long-lived worker process, then start runs +from anywhere with the [client](../reference/client.md). + +Tools must be defined in the process that calls `serve` or `call_*`. Call +`shutdown` before a short-lived process exits. diff --git a/docs/agents/concepts/guardrails.md b/docs/agents/concepts/guardrails.md new file mode 100644 index 0000000..3f3e1be --- /dev/null +++ b/docs/agents/concepts/guardrails.md @@ -0,0 +1,24 @@ +# Guardrails + +Guardrails check input or output before the next step. + +```ruby +no_pii = Guardrail.new(name: 'no_pii', position: :output, on_fail: :retry) do |content| + GuardrailResult.new(passed: !content.match?(/\d{3}-\d{2}-\d{4}/), message: 'SSN found') +end + +agent = Agent.new(name: 'assistant', model: 'openai/gpt-4o-mini', guardrails: [ + RegexGuardrail.new([/\S+@\S+/], mode: :block, position: :output), + LlmGuardrail.new('openai/gpt-4o-mini', 'No medical advice.'), + no_pii +]) +agent.redact(%w[password token]) # shortcut for a blocking regex guardrail +``` + +`on_fail:` is `:raise`, `:retry` (up to `max_retries:`), `:fix`, or `:human`. +Retry only when a new model response could plausibly pass. + +Use `RegexGuardrail` for format checks, `LlmGuardrail` for policy, and a block +only when the rule needs application state. Put a guardrail on a `ToolDef` +(`guardrails:`) when it protects one side effect. Don't send secrets to an LLM +guardrail. Decisions show up in the execution history. diff --git a/docs/agents/concepts/multi-agent.md b/docs/agents/concepts/multi-agent.md new file mode 100644 index 0000000..4313e5d --- /dev/null +++ b/docs/agents/concepts/multi-agent.md @@ -0,0 +1,21 @@ +# Multi-agent + +```ruby +team = Agent.new(name: 'support', model: 'openai/gpt-4o-mini', + agents: [billing, shipping], strategy: :handoff) +pipeline = researcher >> writer >> editor # sequential +``` + +Strategies: `:sequential`, `:parallel`, `:handoff` (default; the model picks a +specialist), `:router` (a `router:` agent picks), `:swarm`, `:round_robin`, +`:random`, `:manual`, `:plan_execute`. + +- `a.hands_off_to(b, on: 'billing')` hands off when the text appears; `on:` also + takes a callable. +- `ToolDef.agent(child)` calls a child and returns its answer to the parent + instead of transferring control. +- `plan_execute(name:, tools:, model:, planner_instructions:, fallback_instructions:, fallback_max_turns:)` + builds a planner that writes a plan executed as a durable sub-workflow. + +Every open-ended design needs `max_turns:` and a [termination](termination.md) +condition. Each child appears in the execution history. diff --git a/docs/agents/concepts/scheduling.md b/docs/agents/concepts/scheduling.md new file mode 100644 index 0000000..948dcf0 --- /dev/null +++ b/docs/agents/concepts/scheduling.md @@ -0,0 +1,8 @@ +# Scheduling + +Deploy the agent, then create a workflow schedule for it with +`OrkesClients#get_scheduler_client`. Agents compile to workflows, so the normal +scheduler applies; there is no agent-specific schedule API. + +Use a stable schedule name, a timezone-aware cron, and idempotent input. Pause or +delete the schedule before deleting the agent or stopping its workers. diff --git a/docs/agents/concepts/stateful.md b/docs/agents/concepts/stateful.md new file mode 100644 index 0000000..7390a11 --- /dev/null +++ b/docs/agents/concepts/stateful.md @@ -0,0 +1,14 @@ +# Stateful agents + +```ruby +agent = Agent.new(name: 'chat', model: 'openai/gpt-4o-mini', stateful: true, + memory: ConversationMemory.new(max_messages: 50)) +agent.call_sync('Hi, I am Ana.', session_id: 'user-42') +agent.call_sync('What is my name?', session_id: 'user-42') +``` + +`stateful: true` keeps conversation state on the server per `session_id:` and +pins the run's tool workers to its own task domain. After a process restart, +`Execution.find(execution_id)` reattaches. + +Bound context with `max_messages:` rather than growing the prompt. diff --git a/docs/agents/concepts/streaming-hitl.md b/docs/agents/concepts/streaming-hitl.md new file mode 100644 index 0000000..69ebf82 --- /dev/null +++ b/docs/agents/concepts/streaming-hitl.md @@ -0,0 +1,27 @@ +# Streaming and approval + +```ruby +execution = agent.call_async('Transfer $500', on_event: ->(e) { puts e['event'] }) +execution.result(timeout: 120) +``` + +Events (`message`, `tool_call`, `tool_result`, `waiting`, `done`, `error`) +arrive over SSE; the runtime falls back to polling when SSE is unavailable. +`execution.partial_text` holds the text so far; nothing is final until +`execution.done?`. + +Tools with `approval_required: true` pause the execution: + +```ruby +agent.on_approval do |request| + puts request.tool_name, request.arguments + request.approve # or request.reject('no'), request.respond(hash) +end +``` + +`ToolDef.human` asks a person a question; `ToolDef.wait_for_message` waits for +`execution.signal(message)`. Without a handler, use `execution.approve`, +`reject`, `signal`, or the [client](../reference/client.md). Approval calls are +safe to repeat. + +Call `Conductor::Agents.shutdown` when a short-lived program is done. diff --git a/docs/agents/concepts/structured-output.md b/docs/agents/concepts/structured-output.md new file mode 100644 index 0000000..86c4260 --- /dev/null +++ b/docs/agents/concepts/structured-output.md @@ -0,0 +1,17 @@ +# Structured output + +```ruby +agent = Agent.new(name: 'extract', model: 'openai/gpt-4o-mini', + instructions: 'Extract the invoice.', + output_type: { 'type' => 'object', + 'properties' => { 'total' => { 'type' => 'number' } }, + 'required' => ['total'] }) +execution = agent.call_async(text) +execution.result +execution.output # => { 'total' => 42.0 } +``` + +`output_type:` is a JSON Schema hash, or any object with `to_json_schema`. The +server validates the final answer; on failure the run errors rather than +returning malformed data. Keep the schema small and ask for the same shape in +the instructions. diff --git a/docs/agents/concepts/termination.md b/docs/agents/concepts/termination.md new file mode 100644 index 0000000..de9ef85 --- /dev/null +++ b/docs/agents/concepts/termination.md @@ -0,0 +1,14 @@ +# Termination + +```ruby +agent.stop_when('DONE') +agent.stop_after(messages: 20) +agent.termination = Termination::TokenUsage.new(max_total_tokens: 50_000) | + Termination::TextMention.new('DONE') +``` + +Conditions: `Termination::MaxMessage`, `StopMessage`, `TextMention`, +`TokenUsage`. Combine with `&` and `|`. Always set `max_turns:` too. + +Stop a running execution with `execution.stop` or `AgentClient#stop`. Stopping +does not undo a tool call that already ran. diff --git a/docs/agents/concepts/tools.md b/docs/agents/concepts/tools.md new file mode 100644 index 0000000..be52078 --- /dev/null +++ b/docs/agents/concepts/tools.md @@ -0,0 +1,38 @@ +# Tools + +`tool def` turns a method into a tool. Keyword defaults give the schema +(`city: String`, `units: 'metric'`). Each call is a retryable Conductor task. + +```ruby +require 'conductor/agents' +include Conductor::Agents + +tool def create_issue(title: String) + Github.create_issue(title, token: secret('GITHUB_TOKEN')) + "created: #{title}" +end +describe :create_issue, 'Create a GitHub issue.' +requires_approval :create_issue +``` + +`secret('NAME')` inside the body both declares the credential and reads it at +run time. The server resolves it from its secret store into the task; nothing goes +through ENV or workflow input. Use `credentials:` for names built dynamically. + +Server-side tools need no worker: + +| Tool | Factory | +|---|---| +| HTTP endpoint | `ToolDef.http(name, url, ...)` | +| OpenAPI / Postman | `ToolDef.api(url)` | +| MCP server | `ToolDef.mcp(url)` | +| Human answer | `ToolDef.human(name, description:)` | +| Wait for a message | `ToolDef.wait_for_message(name, description:)` | +| Media, PDF, vectors | `ToolDef.image`, `.audio`, `.video`, `.pdf`, `.index`, `.search` | +| Another agent | `ToolDef.agent(child)` or `agent.add_tool(child)` | + +Options on `tool` and the factories: `retry_count:`, `retry_delay_seconds:`, +`timeout_seconds:`, `max_calls:`, `approval_required:`, `output_schema:`. + +A tool stuck in `SCHEDULED` has no worker polling. `CredentialNotFoundError` +means the secret is missing on the server. diff --git a/docs/agents/getting-started.md b/docs/agents/getting-started.md new file mode 100644 index 0000000..f95e6f9 --- /dev/null +++ b/docs/agents/getting-started.md @@ -0,0 +1,29 @@ +# Getting started + +```shell +gem install conductor_ruby +export CONDUCTOR_SERVER_URL=http://localhost:8080/api +export CONDUCTOR_AGENT_LLM_MODEL=openai/gpt-4o-mini +``` + +Set `CONDUCTOR_AUTH_KEY` and `CONDUCTOR_AUTH_SECRET` for authenticated servers. +The LLM provider is configured on the server, not in your code. + +```ruby +require 'conductor/agents' +include Conductor::Agents + +agent = Agent.new(name: 'greeter', model: 'openai/gpt-4o-mini', + instructions: 'You are a friendly assistant.') +puts agent.call_sync('Say hello.') +Conductor::Agents.shutdown +``` + +Or run the checked-in version: + +```shell +bundle exec ruby -Ilib examples/agents/01_basic_agent.rb +``` + +A model error means the provider is not set up on the server. A connection error +means `CONDUCTOR_SERVER_URL` is wrong. diff --git a/docs/agents/reference/agent-definition.md b/docs/agents/reference/agent-definition.md new file mode 100644 index 0000000..f91c988 --- /dev/null +++ b/docs/agents/reference/agent-definition.md @@ -0,0 +1,10 @@ +# Agent fields + +`Agent.new(name:, model:, instructions:, tools:, agents:, strategy:, router:, +output_type:, guardrails:, memory:, termination:, handoffs:, callbacks:, +credentials:, max_turns: 25, max_tokens:, timeout_seconds:, temperature:, +stateful:, metadata:, description:, planner:, fallback:, fallback_max_turns:)` + +`name` must match `^[a-zA-Z_][a-zA-Z0-9_-]*$`. A nil `model` inherits the +parent's. The serializer in `lib/conductor/agents/config_serializer.rb` is the +source of truth for what each field becomes on the wire. diff --git a/docs/agents/reference/agent-schema.md b/docs/agents/reference/agent-schema.md new file mode 100644 index 0000000..1d6b67e --- /dev/null +++ b/docs/agents/reference/agent-schema.md @@ -0,0 +1,13 @@ +# Wire contract + +`ConfigSerializer.serialize(agent)` produces the `agentConfig` hash sent to the +server. It is byte-for-byte the Python SDK's output: keys are camelCase, and +`agents`, `router`, `planner`, `fallback` nest recursively. + +The server schema is vendored at +[spec/fixtures/agents/agent-schema.json](../../../spec/fixtures/agents/agent-schema.json). +`spec/conductor/agents/contract_spec.rb` validates the golden configs against +it; `examples_spec.rb` compares every example to its Python-generated fixture. + +Adding a field means changing the serializer, the schema, and both specs +together. The Python SDK serializer is the parity source. diff --git a/docs/agents/reference/api.md b/docs/agents/reference/api.md new file mode 100644 index 0000000..799b16d --- /dev/null +++ b/docs/agents/reference/api.md @@ -0,0 +1,15 @@ +# API map + +Everything lives under `Conductor::Agents` after `require 'conductor/agents'`. + +| Need | Ruby | Guide | +|---|---|---| +| Define an agent | `Agent` | [agents](../concepts/agents.md) | +| Tools | `tool def`, `ToolDef.*` | [tools](../concepts/tools.md) | +| Run | `AgentRuntime`, `Conductor::Agents.runtime` | [runtime](runtime.md) | +| Control a run | `Execution`, `ApprovalRequest`, `Client::AgentClient` | [client](client.md) | +| Safety | `Guardrail`, `RegexGuardrail`, `LlmGuardrail`, `Termination::*` | [guardrails](../concepts/guardrails.md) | +| Compose | `strategy:`, `>>`, `hands_off_to`, `plan_execute` | [multi-agent](../concepts/multi-agent.md) | +| Wire format | `ConfigSerializer` | [contract](agent-schema.md) | + +Signatures are documented in YARD comments under `lib/conductor/agents/`. diff --git a/docs/agents/reference/client.md b/docs/agents/reference/client.md new file mode 100644 index 0000000..132e17f --- /dev/null +++ b/docs/agents/reference/client.md @@ -0,0 +1,16 @@ +# AgentClient + +`Conductor::Client::AgentClient` wraps `/api/agent/*` with the SDK's normal +transport and auth. Get it with `OrkesClients#get_agent_client`. Use it when the +tools are server-side or already served elsewhere and you only need control. + +| | Methods | +|---|---| +| Lifecycle | `compile_agent`, `deploy_agent`, `start_agent` | +| Inspect | `get_status`, `get_execution`, `list_executions`, `list_agents`, `get_agent`, `delete_agent` | +| Control | `approve`, `reject`, `respond`, `send_message`, `signal`, `pause`, `resume`, `stop`, `cancel` | +| Events | `stream_sse(execution_id, last_event_id:) { |event| }` | + +`stream_sse` raises `SseUnavailableError` when the server has no stream; poll +`get_status` instead. Control calls change a durable execution, so authorize +callers and make retries idempotent. diff --git a/docs/agents/reference/runtime.md b/docs/agents/reference/runtime.md new file mode 100644 index 0000000..80c3dbb --- /dev/null +++ b/docs/agents/reference/runtime.md @@ -0,0 +1,24 @@ +# AgentRuntime + +One per process. `Conductor::Agents.runtime` is the default, built from ENV; +`AgentRuntime.new(configuration:, agent_config:, logger:)` for anything else. + +| Method | Returns | +|---|---| +| `call_sync(agent, prompt, session_id:, timeout:, **opts)` | answer `String` | +| `call_async(agent, prompt, on_event:, **opts, &on_done)` | `Execution` | +| `compile(agent)` | `{ 'workflowDef', 'requiredWorkers' }` | +| `deploy(*agents)` | deployed names | +| `serve(*agents, blocking: true)` | blocks | +| `shutdown(timeout: 5)` | stops workers and streams | + +`opts`: `media:`, `context:`, `idempotency_key:`, `timeout_seconds:`. + +`AgentConfig.from_env` reads `CONDUCTOR_AGENT_WORKER_POLL_INTERVAL`, +`CONDUCTOR_AGENT_WORKER_THREADS`, `CONDUCTOR_AGENT_INTEGRATIONS_AUTO_REGISTER`, +`CONDUCTOR_AGENT_STREAMING_ENABLED`. Server URL and auth come from +`Configuration` (`CONDUCTOR_*`). + +`Execution`: `result(timeout:)`, `output`, `status`, `done?`, `waiting?`, +`partial_text`, `tool_calls`, `token_usage`, `approve`, `reject`, `signal`, +`pause`, `resume`, `stop`, `cancel`. `Execution.find(id)` reattaches. diff --git a/lib/conductor/agents/agent.rb b/lib/conductor/agents/agent.rb index 29b3050..d3c4867 100644 --- a/lib/conductor/agents/agent.rb +++ b/lib/conductor/agents/agent.rb @@ -123,8 +123,8 @@ def strategy_set? # Give the agent a tool. # @param tool [Symbol, String, ToolDef, Module, Class, Agent] a tool name defined with - # `tool def`, a ToolDef, a module that `extend Conductor::Agents::Tools`, a RubyLLM::Tool - # class, or another Agent (wrapped as an agent tool) + # `tool def`, a ToolDef, a module that `extend Conductor::Agents::Tools`, or another + # Agent (wrapped as an agent tool) # @param credentials [Array, nil] secret names when the scanner cannot see them # @return [self] def add_tool(tool, credentials: nil) @@ -312,13 +312,9 @@ def resolve_tool_defs(tool) [Tools.lookup(tool) || raise(ConfigurationError, "no tool named #{tool.inspect}; define it with `tool def #{tool}(...)` first")] when Agent then [ToolDef.agent(tool)] when Module - if tool.respond_to?(:tool_defs) - tool.tool_defs - elsif Tools::RubyLlmAdapter.ruby_llm_tool?(tool) - [Tools::RubyLlmAdapter.to_tool_def(tool)] - else - raise ConfigurationError, "#{tool} has no tools; use `extend Conductor::Agents::Tools` and `tool def ...`" - end + raise ConfigurationError, "#{tool} has no tools; use `extend Conductor::Agents::Tools` and `tool def ...`" unless tool.respond_to?(:tool_defs) + + tool.tool_defs else raise ConfigurationError, "cannot use #{tool.inspect} as a tool" end diff --git a/lib/conductor/agents/tools.rb b/lib/conductor/agents/tools.rb index 68e0257..f125826 100644 --- a/lib/conductor/agents/tools.rb +++ b/lib/conductor/agents/tools.rb @@ -5,7 +5,6 @@ require_relative 'tool_def' require_relative 'tools/schema_builder' require_relative 'tools/secret_scanner' -require_relative 'tools/ruby_llm_adapter' module Conductor module Agents diff --git a/lib/conductor/agents/tools/ruby_llm_adapter.rb b/lib/conductor/agents/tools/ruby_llm_adapter.rb deleted file mode 100644 index 243948b..0000000 --- a/lib/conductor/agents/tools/ruby_llm_adapter.rb +++ /dev/null @@ -1,73 +0,0 @@ -# frozen_string_literal: true - -module Conductor - module Agents - module Tools - # Adapts a RubyLLM::Tool class to a ToolDef so it can be given to an Agent as-is. - # - # RubyLLM exposes +name+, +description+ and +parameters+ (name => Parameter with - # +type+, +description+, +required+) and executes via +#execute(**args)+. RubyLLM is an - # optional dependency: this adapter only engages when RubyLLM::Tool is defined. - module RubyLlmAdapter - module_function - - # @return [Boolean] true when +klass+ is a RubyLLM::Tool subclass - def ruby_llm_tool?(klass) - return false unless defined?(::RubyLLM::Tool) - - klass.is_a?(Class) && klass < ::RubyLLM::Tool - end - - # @param klass [Class] a RubyLLM::Tool subclass - # @return [ToolDef] - def to_tool_def(klass) - instance = klass.new - name = read(klass, instance, :name) || snake_case(klass.name.to_s.split('::').last) - description = read(klass, instance, :description) || '' - params = read(klass, instance, :parameters) || {} - - properties = {} - required = [] - params.each do |param_name, param| - schema = { 'type' => (fetch(param, :type) || 'string').to_s } - desc = fetch(param, :description) - schema['description'] = desc if desc - properties[param_name.to_s] = schema - required << param_name.to_s if fetch(param, :required) != false - end - - input_schema = { 'type' => 'object', 'properties' => properties } - input_schema['required'] = required unless required.empty? - - ToolDef.new( - name: name.to_s, - description: description.to_s, - input_schema: input_schema, - output_schema: SchemaBuilder.default_output_schema, - func: ->(**args) { instance.execute(**args) } - ) - end - - def read(klass, instance, attr) - if klass.respond_to?(attr) - klass.public_send(attr) - elsif instance.respond_to?(attr) - instance.public_send(attr) - end - end - - def fetch(param, key) - if param.respond_to?(key) - param.public_send(key) - elsif param.respond_to?(:[]) - param[key] || param[key.to_s] - end - end - - def snake_case(name) - name.gsub(/([A-Z]+)([A-Z][a-z])/, '\1_\2').gsub(/([a-z\d])([A-Z])/, '\1_\2').downcase - end - end - end - end -end diff --git a/spec/conductor/agents/tools_spec.rb b/spec/conductor/agents/tools_spec.rb index 267860f..1f2054a 100644 --- a/spec/conductor/agents/tools_spec.rb +++ b/spec/conductor/agents/tools_spec.rb @@ -113,35 +113,4 @@ expect(weather.tool_defs.map(&:name)).to eq(%w[current forecast]) end end - - describe 'RubyLLM adapter' do - before do - stub_const('RubyLLM', SpecTools::FakeRubyLLM) - stub_const('WeatherLookup', Class.new(SpecTools::FakeRubyLLM::Tool) do - desc 'Gets current weather for a location' - param :latitude, type: :number, desc: 'Latitude' - param :longitude, type: :number, desc: 'Longitude' - param :units, type: :string, required: false - - def execute(latitude:, longitude:, units: 'metric') - { lat: latitude, lon: longitude, units: units } - end - end) - end - - it 'converts the class to a ToolDef and executes an instance' do - td = Conductor::Agents::Tools::RubyLlmAdapter.to_tool_def(WeatherLookup) - expect(td.name).to eq('weather_lookup') - expect(td.description).to eq('Gets current weather for a location') - expect(td.input_schema['required']).to eq(%w[latitude longitude]) - expect(td.input_schema['properties']['latitude']).to eq('type' => 'number', 'description' => 'Latitude') - expect(td.func.call(latitude: 1.0, longitude: 2.0)).to eq(lat: 1.0, lon: 2.0, units: 'metric') - end - - it 'is picked up by Agent#add_tool' do - agent = Conductor::Agents::Agent.new(name: 'a', model: 'openai/gpt-4o') - agent.add_tool WeatherLookup - expect(agent.tool('weather_lookup')).not_to be_nil - end - end end diff --git a/spec/fixtures/agents/README.md b/spec/fixtures/agents/README.md index 96b4ce6..a25ea17 100644 --- a/spec/fixtures/agents/README.md +++ b/spec/fixtures/agents/README.md @@ -1,24 +1,21 @@ # Agent contract fixtures -`agent-schema.json` and the original `configs/` files are the vendored server schema and -Python golden contracts. `configs/103_plan_and_compile.json` adds the named planner/fallback -configuration from Python's `plan_execute` helper. +Python-generated expectations for the Ruby `ConfigSerializer`. Nothing here is produced by Ruby. -`examples/*.json` contains the configurations of the actual 19 requested Python examples at -`conductor-oss/python-sdk@c99e2cf9871c21f7a64d823126ee1b77989b00ad`, checked against their -GitHub `main` source on 2026-09-17. Each file is an array in example execution order; repeated -runs of the same agent use one configuration. `21_regex_guardrails` has two configurations. -All `model` fields (including LLM guards) are normalized to `openai/gpt-4o-mini`. +- `agent-schema.json`: the server's agent config schema, vendored from Conductor. +- `configs/`: golden `agentConfig` outputs from the Python SDK's contract suite. + `103_plan_and_compile.json` adds the planner/fallback config from Python's `plan_execute`. +- `examples/`: one file per ported Python example, from + `conductor-oss/python-sdk@c99e2cf9871c21f7a64d823126ee1b77989b00ad` (checked against + `main` on 2026-09-17). Each file is an array of configs in execution order; an agent run + more than once appears once. All `model` fields are normalized to `openai/gpt-4o-mini`. -Generation used Python's `AgentConfigSerializer.serialize` on the imported example agents. -The second regex agent is constructed from its assignment in the example's main block; -`103` uses the example's factorial/write_summary/check_summary tools and planner instructions -with `plan_execute`, its fallback instructions, and fallback_max_turns=4. No Ruby serializer -output is used to produce these fixtures. To refresh, import the corresponding Python example, -serialize each agent passed to `runtime.run`/`start`, normalize models recursively, and write -sorted, indented JSON. Review the source revision and fixture changes together. +`spec/conductor/agents/examples_spec.rb` builds each Ruby example and compares its full +config with `examples/`. `contract_spec.rb` validates `configs/` against the schema. -`spec/conductor/agents/examples_spec.rb` builds the actual Ruby examples, compares full configs -with these fixtures, and validates them against the schema. These are contract expectations, -not separate implementations of the examples. Runtime recordings live exclusively in the -Conductor feature-branch checkout's `llm-recordings`; see the playback workflow. +To refresh `examples/`: import the Python example, serialize each agent it passes to +`runtime.run` or `start` with `AgentConfigSerializer.serialize`, normalize models +recursively, and write sorted, indented JSON. Review the source revision and the fixture +diff together. + +LLM recordings for playback live in `llm-recordings/` in conductor-oss/conductor, not here. diff --git a/spec/support/agent_tools.rb b/spec/support/agent_tools.rb index 9341ca0..4911a00 100644 --- a/spec/support/agent_tools.rb +++ b/spec/support/agent_tools.rb @@ -56,40 +56,4 @@ def self.positional(city, units: 'metric') [city, units] end end - - # A stand-in for RubyLLM::Tool so the adapter can be exercised without the gem - module FakeRubyLLM - class Tool - class Param - attr_reader :type, :description, :required - - def initialize(type:, description:, required:) - @type = type - @description = description - @required = required - end - end - - class << self - def desc(text = nil) - @description = text if text - @description - end - - attr_reader :description - - def param(name, type: :string, desc: nil, required: true) - (@parameters ||= {})[name] = Param.new(type: type, description: desc, required: required) - end - - def parameters - @parameters || {} - end - - def name - super.split('::').last.gsub(/([a-z\d])([A-Z])/, '\1_\2').downcase - end - end - end - end end From 21542169500d2ad04c0b375deb259ca5ab39b858 Mon Sep 17 00:00:00 2001 From: nicholascole Date: Thu, 24 Sep 2026 12:08:12 -0700 Subject: [PATCH 20/20] ToolDef -> Tool --- docs/agents/concepts/guardrails.md | 2 +- docs/agents/concepts/multi-agent.md | 2 +- docs/agents/concepts/streaming-hitl.md | 2 +- docs/agents/concepts/tools.md | 24 ++++++++------ docs/agents/reference/api.md | 2 +- examples/agents/04_http_and_mcp_tools.rb | 4 +-- examples/agents/16e_credentials_http_tool.rb | 2 +- examples/agents/golden_agents.rb | 2 +- lib/conductor/agents.rb | 2 +- lib/conductor/agents/agent.rb | 12 +++---- lib/conductor/agents/runtime/dispatch.rb | 2 +- lib/conductor/agents/{tool_def.rb => tool.rb} | 23 ++++++++++---- lib/conductor/agents/tools.rb | 31 ++++++++++--------- spec/conductor/agents/agent_spec.rb | 4 +-- .../agents/config_serializer_spec.rb | 4 +-- spec/conductor/agents/dispatch_spec.rb | 24 +++++++------- spec/conductor/agents/plans_spec.rb | 2 +- spec/conductor/agents/tool_registry_spec.rb | 4 +-- .../agents/{tool_def_spec.rb => tool_spec.rb} | 23 +++++++++++++- spec/conductor/agents/tools_spec.rb | 23 +++++++++++++- 20 files changed, 127 insertions(+), 67 deletions(-) rename lib/conductor/agents/{tool_def.rb => tool.rb} (93%) rename spec/conductor/agents/{tool_def_spec.rb => tool_spec.rb} (83%) diff --git a/docs/agents/concepts/guardrails.md b/docs/agents/concepts/guardrails.md index 3f3e1be..e30b0dd 100644 --- a/docs/agents/concepts/guardrails.md +++ b/docs/agents/concepts/guardrails.md @@ -19,6 +19,6 @@ agent.redact(%w[password token]) # shortcut for a blocking regex guardrail Retry only when a new model response could plausibly pass. Use `RegexGuardrail` for format checks, `LlmGuardrail` for policy, and a block -only when the rule needs application state. Put a guardrail on a `ToolDef` +only when the rule needs application state. Put a guardrail on a `Tool` (`guardrails:`) when it protects one side effect. Don't send secrets to an LLM guardrail. Decisions show up in the execution history. diff --git a/docs/agents/concepts/multi-agent.md b/docs/agents/concepts/multi-agent.md index 4313e5d..4164769 100644 --- a/docs/agents/concepts/multi-agent.md +++ b/docs/agents/concepts/multi-agent.md @@ -12,7 +12,7 @@ specialist), `:router` (a `router:` agent picks), `:swarm`, `:round_robin`, - `a.hands_off_to(b, on: 'billing')` hands off when the text appears; `on:` also takes a callable. -- `ToolDef.agent(child)` calls a child and returns its answer to the parent +- `Tool.agent(child)` calls a child and returns its answer to the parent instead of transferring control. - `plan_execute(name:, tools:, model:, planner_instructions:, fallback_instructions:, fallback_max_turns:)` builds a planner that writes a plan executed as a durable sub-workflow. diff --git a/docs/agents/concepts/streaming-hitl.md b/docs/agents/concepts/streaming-hitl.md index 69ebf82..0ad4bcb 100644 --- a/docs/agents/concepts/streaming-hitl.md +++ b/docs/agents/concepts/streaming-hitl.md @@ -19,7 +19,7 @@ agent.on_approval do |request| end ``` -`ToolDef.human` asks a person a question; `ToolDef.wait_for_message` waits for +`Tool.human` asks a person a question; `Tool.wait_for_message` waits for `execution.signal(message)`. Without a handler, use `execution.approve`, `reject`, `signal`, or the [client](../reference/client.md). Approval calls are safe to repeat. diff --git a/docs/agents/concepts/tools.md b/docs/agents/concepts/tools.md index be52078..087c6e8 100644 --- a/docs/agents/concepts/tools.md +++ b/docs/agents/concepts/tools.md @@ -23,16 +23,20 @@ Server-side tools need no worker: | Tool | Factory | |---|---| -| HTTP endpoint | `ToolDef.http(name, url, ...)` | -| OpenAPI / Postman | `ToolDef.api(url)` | -| MCP server | `ToolDef.mcp(url)` | -| Human answer | `ToolDef.human(name, description:)` | -| Wait for a message | `ToolDef.wait_for_message(name, description:)` | -| Media, PDF, vectors | `ToolDef.image`, `.audio`, `.video`, `.pdf`, `.index`, `.search` | -| Another agent | `ToolDef.agent(child)` or `agent.add_tool(child)` | - -Options on `tool` and the factories: `retry_count:`, `retry_delay_seconds:`, -`timeout_seconds:`, `max_calls:`, `approval_required:`, `output_schema:`. +| HTTP endpoint | `Tool.http(name, url, ...)` | +| OpenAPI / Postman | `Tool.api(url)` | +| MCP server | `Tool.mcp(url)` | +| Human answer | `Tool.human(name, description:)` | +| Wait for a message | `Tool.wait_for_message(name, description:)` | +| Media, PDF, vectors | `Tool.image`, `.audio`, `.video`, `.pdf`, `.index`, `.search` | +| Another agent | `Tool.agent(child)` or `agent.add_tool(child)` | + +Options on `tool`: `name:` (when the tool name differs from the method name), +`description:`, `guardrails:`, `credentials:`, `external:`, `stateful:`, +`retry_count:`, `retry_delay_seconds:`, `timeout_seconds:`, `max_calls:`, +`approval_required:`, `input_schema:`, `output_schema:`. The factories take the +subset that applies to them. `tool.with_guardrails(g)` returns a guarded copy of +any tool. A tool stuck in `SCHEDULED` has no worker polling. `CredentialNotFoundError` means the secret is missing on the server. diff --git a/docs/agents/reference/api.md b/docs/agents/reference/api.md index 799b16d..f8306fb 100644 --- a/docs/agents/reference/api.md +++ b/docs/agents/reference/api.md @@ -5,7 +5,7 @@ Everything lives under `Conductor::Agents` after `require 'conductor/agents'`. | Need | Ruby | Guide | |---|---|---| | Define an agent | `Agent` | [agents](../concepts/agents.md) | -| Tools | `tool def`, `ToolDef.*` | [tools](../concepts/tools.md) | +| Tools | `tool def`, `Tool.*` | [tools](../concepts/tools.md) | | Run | `AgentRuntime`, `Conductor::Agents.runtime` | [runtime](runtime.md) | | Control a run | `Execution`, `ApprovalRequest`, `Client::AgentClient` | [client](client.md) | | Safety | `Guardrail`, `RegexGuardrail`, `LlmGuardrail`, `Termination::*` | [guardrails](../concepts/guardrails.md) | diff --git a/examples/agents/04_http_and_mcp_tools.rb b/examples/agents/04_http_and_mcp_tools.rb index a52cb7b..5a98904 100644 --- a/examples/agents/04_http_and_mcp_tools.rb +++ b/examples/agents/04_http_and_mcp_tools.rb @@ -14,7 +14,7 @@ def format_report(title: String, body: String) tool :format_report, description: "Format a title and body into a structured report." def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) - reverse_api = ToolDef.http( + reverse_api = Tool.http( "reverse_string", "http://localhost:3001/api/string/reverse", description: "Reverse a string using the HTTP API", @@ -23,7 +23,7 @@ def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini credentials: ["HTTP_TEST_API_KEY"], input_schema: { "type" => "object", "properties" => { "text" => { "type" => "string", "description" => "Text to reverse" } }, "required" => ["text"] } ) - mcp_test_tools = ToolDef.mcp( + mcp_test_tools = Tool.mcp( "http://localhost:3001/mcp", name: "mcp_test_tools", description: "Deterministic test tools via MCP — math, string, collection, encoding, hash, datetime, validation, and conversion operations.", diff --git a/examples/agents/16e_credentials_http_tool.rb b/examples/agents/16e_credentials_http_tool.rb index f955c9d..9a7e222 100644 --- a/examples/agents/16e_credentials_http_tool.rb +++ b/examples/agents/16e_credentials_http_tool.rb @@ -9,7 +9,7 @@ module Example16eCredentialsHttpTool extend Conductor::Agents::Tools def self.build(model: ENV.fetch('CONDUCTOR_AGENT_LLM_MODEL', 'openai/gpt-4o-mini')) - list_repos = ToolDef.http( + list_repos = Tool.http( "list_github_repos", ENV.fetch('GITHUB_REPOS_URL', 'https://api.github.com/users/Conductor/repos?per_page=5&sort=updated'), description: "List public GitHub repositories for a user. Returns JSON array with name, url, and stars.", diff --git a/examples/agents/golden_agents.rb b/examples/agents/golden_agents.rb index 93213ab..5330e47 100644 --- a/examples/agents/golden_agents.rb +++ b/examples/agents/golden_agents.rb @@ -355,7 +355,7 @@ def on_model_end(**_kwargs) researcher = A::Agent.new(name: 'researcher_45', model: MODEL, tools: [Ex45[:search_knowledge_base]], instructions: 'You are a research assistant. Use search_knowledge_base to find ' \ 'information about topics. Provide concise summaries.') - A::Agent.new(name: 'manager_45', model: MODEL, tools: [A::ToolDef.agent(researcher), Ex45[:calculate]], + A::Agent.new(name: 'manager_45', model: MODEL, tools: [A::Tool.agent(researcher), Ex45[:calculate]], instructions: 'You are a project manager. Use the researcher tool to gather ' \ 'information and the calculate tool for math. Synthesize findings.') }, diff --git a/lib/conductor/agents.rb b/lib/conductor/agents.rb index 5bcd600..36031e9 100644 --- a/lib/conductor/agents.rb +++ b/lib/conductor/agents.rb @@ -15,7 +15,7 @@ require_relative '../conductor' require_relative 'agents/errors' require_relative 'agents/runtime/secrets' -require_relative 'agents/tool_def' +require_relative 'agents/tool' require_relative 'agents/tools' require_relative 'agents/guardrail' require_relative 'agents/termination' diff --git a/lib/conductor/agents/agent.rb b/lib/conductor/agents/agent.rb index d3c4867..3563294 100644 --- a/lib/conductor/agents/agent.rb +++ b/lib/conductor/agents/agent.rb @@ -1,7 +1,7 @@ # frozen_string_literal: true require_relative 'errors' -require_relative 'tool_def' +require_relative 'tool' require_relative 'tools' require_relative 'guardrail' require_relative 'termination' @@ -122,8 +122,8 @@ def strategy_set? # ── Tools ───────────────────────────────────────────────────────── # Give the agent a tool. - # @param tool [Symbol, String, ToolDef, Module, Class, Agent] a tool name defined with - # `tool def`, a ToolDef, a module that `extend Conductor::Agents::Tools`, or another + # @param tool [Symbol, String, Tool, Module, Class, Agent] a tool name defined with + # `tool def`, a Tool, a module that `extend Conductor::Agents::Tools`, or another # Agent (wrapped as an agent tool) # @param credentials [Array, nil] secret names when the scanner cannot see them # @return [self] @@ -143,7 +143,7 @@ def add_tools(*tools) self end - # @return [ToolDef, nil] + # @return [Tool, nil] def tool(name) @tools.find { |t| t.name == name.to_s } end @@ -307,10 +307,10 @@ def to_s def resolve_tool_defs(tool) case tool - when ToolDef then [tool] + when Tool then [tool] when Symbol, String [Tools.lookup(tool) || raise(ConfigurationError, "no tool named #{tool.inspect}; define it with `tool def #{tool}(...)` first")] - when Agent then [ToolDef.agent(tool)] + when Agent then [Tool.agent(tool)] when Module raise ConfigurationError, "#{tool} has no tools; use `extend Conductor::Agents::Tools` and `tool def ...`" unless tool.respond_to?(:tool_defs) diff --git a/lib/conductor/agents/runtime/dispatch.rb b/lib/conductor/agents/runtime/dispatch.rb index e62c254..488bc58 100644 --- a/lib/conductor/agents/runtime/dispatch.rb +++ b/lib/conductor/agents/runtime/dispatch.rb @@ -26,7 +26,7 @@ class MissingArgumentError < Error; end module_function # @param task [Http::Models::Task] - # @param tool_def [ToolDef] + # @param tool_def [Tool] # @return [Http::Models::TaskResult] def run_tool_task(task, tool_def, logger: nil) result = base_result(task) diff --git a/lib/conductor/agents/tool_def.rb b/lib/conductor/agents/tool.rb similarity index 93% rename from lib/conductor/agents/tool_def.rb rename to lib/conductor/agents/tool.rb index 183f29e..3255180 100644 --- a/lib/conductor/agents/tool_def.rb +++ b/lib/conductor/agents/tool.rb @@ -34,17 +34,20 @@ def self.valid?(type) end # A tool call with pre-filled arguments (Agent#prefill_tools) - PrefillToolCall = Struct.new(:tool_name, :arguments, :tool_def, keyword_init: true) do + PrefillToolCall = Struct.new(:tool_name, :arguments, :tool, keyword_init: true) do def to_h { 'toolName' => tool_name, 'arguments' => arguments || {} } end end - # Definition of one tool. Same fields and defaults as the Python SDK's ToolDef. + # A tool an agent can call. This is the developer-facing type: the counterpart of + # the Java SDK's @Tool / HttpTool / McpTool builders and the Python SDK's @tool / + # http_tool / mcp_tool functions. The wire-level tool config the server receives is + # produced by ConfigSerializer and never handed to developers. # # Worker tools are created by the Tools DSL (+tool def ...+); server-side tools by - # the factories below (ToolDef.http, .mcp, .human, .agent, ...). - class ToolDef + # the factories below (Tool.http, .mcp, .human, .agent, ...). + class Tool RETRY_POLICIES = %w[fixed linear_backoff exponential_backoff].freeze RETRY_LOGIC = { 'fixed' => 'FIXED', @@ -109,13 +112,21 @@ def add_credentials(*names) self end + # A copy of this tool guarded by +guardrails+ (replaces any existing ones) + # @return [Tool] + def with_guardrails(*guardrails) + copy = dup + copy.guardrails = guardrails.flatten + copy + end + # Build a pre-filled call for Agent#prefill_tools def call(**args) - PrefillToolCall.new(tool_name: @name, arguments: args.transform_keys(&:to_s), tool_def: self) + PrefillToolCall.new(tool_name: @name, arguments: args.transform_keys(&:to_s), tool: self) end def to_s - "#" + "#" end alias inspect to_s diff --git a/lib/conductor/agents/tools.rb b/lib/conductor/agents/tools.rb index f125826..91bc1bf 100644 --- a/lib/conductor/agents/tools.rb +++ b/lib/conductor/agents/tools.rb @@ -2,7 +2,7 @@ require_relative 'errors' require_relative 'runtime/secrets' -require_relative 'tool_def' +require_relative 'tool' require_relative 'tools/schema_builder' require_relative 'tools/secret_scanner' @@ -19,17 +19,17 @@ module Agents # describe :get_weather, 'Get the current weather for a city.' # requires_approval :get_weather # - # +tool+ receives the Symbol that +def+ returns, builds a ToolDef from the method + # +tool+ receives the Symbol that +def+ returns, builds a Tool from the method # (schema from keyword defaults, secrets from literal secret() calls) and registers it # both on the receiver (Weather[:get_weather], Weather.tool_defs) and in the global # registry that Agent#add_tool(:get_weather) consults. module Tools - TOOL_OPTIONS = %i[description input_schema output_schema approval_required timeout_seconds credentials - stateful max_calls retry_count retry_delay_seconds retry_policy external].freeze + TOOL_OPTIONS = %i[name description input_schema output_schema approval_required timeout_seconds credentials + guardrails stateful max_calls retry_count retry_delay_seconds retry_policy external].freeze # Weather[:current] on a module that `extend Conductor::Agents::Tools` module Lookup - # @return [ToolDef] + # @return [Tool] def [](name) fetch_tool(name) end @@ -42,7 +42,7 @@ def extended(base) base.extend(Secrets) end - # Global name => ToolDef registry shared by every scope that defines tools + # Global name => Tool registry shared by every scope that defines tools def registry @registry ||= {} end @@ -56,7 +56,7 @@ def register(tool_def) tool_def end - # @return [ToolDef, nil] + # @return [Tool, nil] def lookup(name) registry_mutex.synchronize { registry[name.to_s] } end @@ -66,7 +66,7 @@ def clear! registry_mutex.synchronize { registry.clear } end - # Build a ToolDef from a bound Method + # Build a Tool from a bound Method # @param method [Method] # @param name [String, nil] tool name override def build(method, name: nil, **options) @@ -77,7 +77,7 @@ def build(method, name: nil, **options) input_schema = options.fetch(:input_schema) { SchemaBuilder.input_schema(method) } credentials = SecretScanner.scan(method) - ToolDef.new( + Tool.new( name: tool_name, description: options.fetch(:description) { humanize(method.name) }, input_schema: input_schema, @@ -86,6 +86,7 @@ def build(method, name: nil, **options) approval_required: options.fetch(:approval_required, false), timeout_seconds: options[:timeout_seconds], credentials: credentials + Array(options[:credentials]), + guardrails: Array(options[:guardrails]), stateful: options.fetch(:stateful, false), max_calls: options[:max_calls], retry_count: options.fetch(:retry_count, 2), @@ -105,12 +106,14 @@ def humanize(name) # Mark a method as a tool # @param name [Symbol, String, Method] method name (what +def+ returns) or a Method - # @param options [Hash] ToolDef overrides: description:, output_schema:, approval_required:, - # timeout_seconds:, credentials:, stateful:, max_calls:, retry_count:, retry_delay_seconds:, retry_policy: - # @return [ToolDef] + # @param options [Hash] Tool overrides: name: (tool name when it differs from the method name), + # description:, output_schema:, approval_required:, timeout_seconds:, credentials:, guardrails:, + # stateful:, max_calls:, retry_count:, retry_delay_seconds:, retry_policy:, external: + # @return [Tool] def tool(name, **options) method = name.is_a?(Method) ? name : resolve_tool_method(name.to_sym) - tool_def = Tools.build(method, name: name.is_a?(Method) ? name.name : name, **options) + tool_name = options.delete(:name) || (name.is_a?(Method) ? name.name : name) + tool_def = Tools.build(method, name: tool_name, **options) tool_registry[tool_def.name] = tool_def Tools.register(tool_def) end @@ -131,7 +134,7 @@ def tool_credentials(name, *secret_names) end # Every tool defined in this scope, in definition order - # @return [Array] + # @return [Array] def tool_defs tool_registry.values end diff --git a/spec/conductor/agents/agent_spec.rb b/spec/conductor/agents/agent_spec.rb index 07e5ff1..3ecbf11 100644 --- a/spec/conductor/agents/agent_spec.rb +++ b/spec/conductor/agents/agent_spec.rb @@ -33,9 +33,9 @@ describe '#add_tool' do let(:agent) { described_class.new(name: 'a', model: model) } - it 'accepts a symbol from the global registry, a ToolDef, a Tools module and an Agent' do + it 'accepts a symbol from the global registry, a Tool, a Tools module and an Agent' do agent.add_tool :current - agent.add_tool Conductor::Agents::ToolDef.http('fetch', 'http://x') + agent.add_tool Conductor::Agents::Tool.http('fetch', 'http://x') agent.add_tool SpecTools::Github agent.add_tool described_class.new(name: 'helper', model: model) expect(agent.tools.map(&:name)).to eq(%w[current fetch create_issue gh_cli dynamic_secret helper]) diff --git a/spec/conductor/agents/config_serializer_spec.rb b/spec/conductor/agents/config_serializer_spec.rb index 2992e6c..f5923a9 100644 --- a/spec/conductor/agents/config_serializer_spec.rb +++ b/spec/conductor/agents/config_serializer_spec.rb @@ -66,7 +66,7 @@ def serialize(agent) it 'replaces the agent in an agent_tool config with agentConfig' do child = a::Agent.new(name: 'child', model: model) - parent = a::Agent.new(name: 'parent', model: model, tools: [a::ToolDef.agent(child, optional: false)]) + parent = a::Agent.new(name: 'parent', model: model, tools: [a::Tool.agent(child, optional: false)]) tool = serialize(parent)['tools'].first expect(tool['toolType']).to eq('agent_tool') expect(tool['config']['agentConfig']['name']).to eq('child') @@ -136,7 +136,7 @@ def serialize(agent) memory = a::ConversationMemory.new(max_messages: 5) memory.add_user_message('hi') agent = a::Agent.new(name: 'm', model: model, memory: memory, max_tokens: 100, temperature: 0.2, - metadata: { 'team' => 'x' }, prefill_tools: [a::ToolDef.new(name: 't').call(a: 1)]) + metadata: { 'team' => 'x' }, prefill_tools: [a::Tool.new(name: 't').call(a: 1)]) agent.callback(:after_model) { |**_| nil } config = serialize(agent) expect(config['memory']).to eq('messages' => [{ 'role' => 'user', 'message' => 'hi' }], 'maxMessages' => 5) diff --git a/spec/conductor/agents/dispatch_spec.rb b/spec/conductor/agents/dispatch_spec.rb index 550d5e3..b8925a0 100644 --- a/spec/conductor/agents/dispatch_spec.rb +++ b/spec/conductor/agents/dispatch_spec.rb @@ -49,26 +49,26 @@ def with_context(task) end it 'wraps scalar results and keeps _state_updates' do - scalar = Conductor::Agents::ToolDef.new(name: 's', func: ->(**) { 'plain' }, - input_schema: { 'type' => 'object', 'properties' => {} }) + scalar = Conductor::Agents::Tool.new(name: 's', func: ->(**) { 'plain' }, + input_schema: { 'type' => 'object', 'properties' => {} }) expect(described_class.run_tool_task(task_with({}), scalar).output_data).to eq('result' => 'plain') - stateful = Conductor::Agents::ToolDef.new(name: 's2', func: ->(**) { { ok: true, _state_updates: { 'n' => 1 } } }, - input_schema: { 'type' => 'object', 'properties' => {} }) + stateful = Conductor::Agents::Tool.new(name: 's2', func: ->(**) { { ok: true, _state_updates: { 'n' => 1 } } }, + input_schema: { 'type' => 'object', 'properties' => {} }) expect(described_class.run_tool_task(task_with({}), stateful).output_data).to eq('ok' => true, '_state_updates' => { 'n' => 1 }) end it 'reports tool exceptions as retryable failures with the reason' do - boom = Conductor::Agents::ToolDef.new(name: 'boom', func: ->(**) { raise 'kaput' }, - input_schema: { 'type' => 'object', 'properties' => {} }) + boom = Conductor::Agents::Tool.new(name: 'boom', func: ->(**) { raise 'kaput' }, + input_schema: { 'type' => 'object', 'properties' => {} }) result = described_class.run_tool_task(task_with({}), boom, logger: Logger.new(nil)) expect(result.status).to eq('FAILED') expect(result.reason_for_incompletion).to eq('RuntimeError: kaput') end it 'fails terminally on unserializable results' do - bad = Conductor::Agents::ToolDef.new(name: 'bad', func: ->(**) { { io: $stdout } }, - input_schema: { 'type' => 'object', 'properties' => {} }) + bad = Conductor::Agents::Tool.new(name: 'bad', func: ->(**) { { io: $stdout } }, + input_schema: { 'type' => 'object', 'properties' => {} }) allow(JSON).to receive(:generate).and_raise(JSON::GeneratorError, 'nope') result = described_class.run_tool_task(task_with({}), bad) expect(result.status).to eq('FAILED_WITH_TERMINAL_ERROR') @@ -99,11 +99,11 @@ def with_context(task) end it 'passes unknown keys only to tools that accept **kwargs' do - strict = Conductor::Agents::ToolDef.new(name: 'strict', func: ->(a:) { { a: a } }, - input_schema: { 'type' => 'object', 'properties' => { 'a' => {} } }) + strict = Conductor::Agents::Tool.new(name: 'strict', func: ->(a:) { { a: a } }, + input_schema: { 'type' => 'object', 'properties' => { 'a' => {} } }) expect(described_class.run_tool_task(task_with({ 'a' => 1, 'zzz' => 2 }), strict).output_data).to eq('a' => 1) - loose = Conductor::Agents::ToolDef.new(name: 'loose', func: ->(a:, **rest) { { a: a, rest: rest } }, - input_schema: { 'type' => 'object', 'properties' => { 'a' => {} } }) + loose = Conductor::Agents::Tool.new(name: 'loose', func: ->(a:, **rest) { { a: a, rest: rest } }, + input_schema: { 'type' => 'object', 'properties' => { 'a' => {} } }) expect(described_class.run_tool_task(task_with({ 'a' => 1, 'zzz' => 2 }), loose).output_data).to eq('a' => 1, 'rest' => { zzz: 2 }) end end diff --git a/spec/conductor/agents/plans_spec.rb b/spec/conductor/agents/plans_spec.rb index de996a7..9f0776b 100644 --- a/spec/conductor/agents/plans_spec.rb +++ b/spec/conductor/agents/plans_spec.rb @@ -4,7 +4,7 @@ require 'conductor/agents' RSpec.describe Conductor::Agents::Plans do - let(:tool) { Conductor::Agents::ToolDef.new(name: 'factorial', func: ->(n:) { (1..n).reduce(1, :*) }) } + let(:tool) { Conductor::Agents::Tool.new(name: 'factorial', func: ->(n:) { (1..n).reduce(1, :*) }) } it 'serializes named slots and makes recovery tools discoverable to the worker runtime' do agent = Conductor::Agents.plan_execute(name: 'math', tools: [tool], model: 'mock/mockLLM', diff --git a/spec/conductor/agents/tool_registry_spec.rb b/spec/conductor/agents/tool_registry_spec.rb index d56d265..2d76366 100644 --- a/spec/conductor/agents/tool_registry_spec.rb +++ b/spec/conductor/agents/tool_registry_spec.rb @@ -11,7 +11,7 @@ let(:weather) do agent = a::Agent.new(name: 'weather', model: 'openai/gpt-4o-mini', instructions: 'Answer weather questions.') agent.add_tool :current - agent.add_tool a::ToolDef.http('fetch', 'http://x') + agent.add_tool a::Tool.http('fetch', 'http://x') agent end @@ -57,7 +57,7 @@ it 'serves custom tool guardrails and function routers required by the server' do guard = a::Guardrail.new(name: 'tool_policy') { |content| content == 'safe' } - tool = a::ToolDef.new(name: 'action', func: -> { {} }, guardrails: [guard]) + tool = a::Tool.new(name: 'action', func: -> { {} }, guardrails: [guard]) child = a::Agent.new(name: 'child', model: 'm/x', tools: [tool]) team = a::Agent.new(name: 'team', agents: [child], strategy: :router, router: ->(prompt) { "#{prompt}_route" }) workers = registry.workers_for(team, required_workers: %w[tool_policy team_router_fn]) diff --git a/spec/conductor/agents/tool_def_spec.rb b/spec/conductor/agents/tool_spec.rb similarity index 83% rename from spec/conductor/agents/tool_def_spec.rb rename to spec/conductor/agents/tool_spec.rb index d314017..cce5e97 100644 --- a/spec/conductor/agents/tool_def_spec.rb +++ b/spec/conductor/agents/tool_spec.rb @@ -3,7 +3,7 @@ require 'spec_helper' require 'conductor/agents' -RSpec.describe Conductor::Agents::ToolDef do +RSpec.describe Conductor::Agents::Tool do describe '#initialize' do it 'applies the Python defaults' do td = described_class.new(name: 't') @@ -32,6 +32,27 @@ end end + describe '#with_guardrails' do + it 'returns a guarded copy and leaves the original untouched' do + guard = Conductor::Agents::Guardrail.new(name: 'g') { true } + original = described_class.new(name: 't', func: -> {}) + guarded = original.with_guardrails(guard) + expect(guarded).not_to equal(original) + expect(guarded.guardrails).to eq([guard]) + expect(guarded.name).to eq('t') + expect(guarded.local?).to be true + expect(original.guardrails).to eq([]) + end + + it 'replaces existing guardrails and flattens lists' do + old = Conductor::Agents::Guardrail.new(name: 'old') { true } + new1 = Conductor::Agents::Guardrail.new(name: 'new1') { true } + new2 = Conductor::Agents::Guardrail.new(name: 'new2') { true } + tool = described_class.new(name: 't', guardrails: [old]) + expect(tool.with_guardrails([new1, new2]).guardrails).to eq([new1, new2]) + end + end + describe '#call' do it 'builds a prefilled tool call' do call = described_class.new(name: 'get_weather').call(city: 'Lisbon') diff --git a/spec/conductor/agents/tools_spec.rb b/spec/conductor/agents/tools_spec.rb index 1f2054a..3941733 100644 --- a/spec/conductor/agents/tools_spec.rb +++ b/spec/conductor/agents/tools_spec.rb @@ -9,7 +9,7 @@ describe 'tool def' do it 'names the tool after the method and humanizes the description' do td = weather[:current] - expect(td).to be_a(Conductor::Agents::ToolDef) + expect(td).to be_a(Conductor::Agents::Tool) expect(td.name).to eq('current') expect(td.description).to eq('Current') expect(td.local?).to be true @@ -63,6 +63,27 @@ expect { weather.tool(:current, colour: 'red') }.to raise_error(Conductor::Agents::ConfigurationError, /unknown tool option/) end + it 'accepts a tool name that differs from the method name' do + scope = Module.new do + extend Conductor::Agents::Tools + def fetch_weather(city: String) = city + tool :fetch_weather, name: 'get_weather' + end + expect(scope[:get_weather].name).to eq('get_weather') + expect(scope.tool_defs.map(&:name)).to eq(['get_weather']) + expect(described_class.lookup('get_weather')).to equal(scope[:get_weather]) + end + + it 'attaches guardrails given at definition time' do + guard = Conductor::Agents::Guardrail.new(name: 'tool_policy') { |content| content == 'safe' } + scope = Module.new do + extend Conductor::Agents::Tools + def guarded(city: String) = city + end + scope.tool(:guarded, guardrails: [guard]) + expect(scope[:guarded].guardrails).to eq([guard]) + end + it 'treats keywords without defaults as required untyped properties' do expect(SpecTools::Plain[:lookup].input_schema).to eq( 'type' => 'object',