Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
bb3a209
docs: agents parity design specs and implementation plan
NicholasDCole Sep 8, 2026
708d6fd
Phase 0: agent transport, runtimeMetadata, instance token cache, upda…
NicholasDCole Sep 8, 2026
d58dc28
Phase 1: Conductor::Agents definition layer and agentConfig serializer
NicholasDCole Sep 8, 2026
aefd104
Phase 2: agent runtime, SSE streaming, tool dispatch, secrets, system…
NicholasDCole Sep 8, 2026
36994ff
Phase 3: replay the recorded tool_happy_path scenario against WireMock
NicholasDCole Sep 8, 2026
48ebeed
Phase 4: agents examples, user docs, README/AGENTS/DESIGN updates, ch…
NicholasDCole Sep 8, 2026
7a71432
Complete Ruby agents example parity and real-server playback CI
NicholasDCole Sep 17, 2026
40ad999
Use published agent recording revision for WireMock CI
NicholasDCole Sep 17, 2026
e72b8e4
Fix agent approval refresh and renew worker task leases
NicholasDCole Sep 21, 2026
b931174
Merge main into feature/conductor_agents
NicholasDCole Sep 21, 2026
15bfd59
Trim documentation to essential setup and contribution guidance
NicholasDCole Sep 21, 2026
6fd05d8
Restore documentation files inherited from main
NicholasDCole Sep 23, 2026
e1cc9cf
Make secret integration tests handle read-only OSS backends
NicholasDCole Sep 23, 2026
f62c6a2
Reuse shared Conductor startup for SDK agent playback
NicholasDCole Sep 23, 2026
1b7ccd4
Report validated guardrail rejections to the playback verifier
NicholasDCole Sep 23, 2026
1f49c71
Delegate agent playback outcome checks to Conductor
NicholasDCole Sep 23, 2026
6a4b505
Use shared Conductor HTTP and MCP playback services
NicholasDCole Sep 23, 2026
a953d80
Remove obsolete conductor-mocks replay job and harness
NicholasDCole Sep 23, 2026
06a0273
Use Conductor main for shared agent playback actions
NicholasDCole Sep 23, 2026
ccead57
Docs
NicholasDCole Sep 24, 2026
2154216
ToolDef -> Tool
NicholasDCole Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 68 additions & 0 deletions .github/workflows/agents-playback.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
name: Agents playback

on:
pull_request:
push:
branches: [main, develop]
workflow_dispatch:

jobs:
playback:
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
env:
CONDUCTOR_SERVER_URL: http://localhost:18080/api
CONDUCTOR_AGENT_LLM_MODEL: mock/mockLLM
CONDUCTOR_AGENTS_PLAYBACK: 'true'
CONDUCTOR_PLAYBACK_WORK_DIR: tmp/agent-playback
GITHUB_REPOS_URL: http://localhost:3002/users/Conductor/repos?per_page=5&sort=updated
steps:
- uses: actions/checkout@v4
- name: Checkout matching server and shared recordings
uses: actions/checkout@v4
with:
repository: conductor-oss/conductor
ref: main
path: tmp/conductor
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'
- uses: gradle/actions/setup-gradle@v4
- name: Build playback server
working-directory: tmp/conductor
run: ./gradlew --no-daemon :conductor-server:bootJar -x test
- uses: ruby/setup-ruby@v1
with:
ruby-version: '3.3'
bundler-cache: true
- name: Start HTTP and MCP services
uses: ./tmp/conductor/.github/actions/start-playback-services
- name: Start fresh playback server
id: server
uses: ./tmp/conductor/.github/actions/start-playback
- name: Run all agent examples
env:
CONDUCTOR_PLAYBACK_SERVICES_STARTED: 'true'
run: bash scripts/run-agents-playback.sh tmp/conductor
- name: Check playback outcomes
if: always() && steps.server.outcome == 'success'
uses: ./tmp/conductor/.github/actions/check-playback
with:
server-url: http://localhost:18080/api
- name: Upload playback diagnostics
if: always()
uses: actions/upload-artifact@v4
with:
name: agents-playback-logs
path: |
tmp/agent-playback/*.log
tmp/agent-playback/results.json
- name: Stop test services
if: always()
run: |
for file in tmp/agent-playback/*.pid; do
[ -f "$file" ] && kill "$(cat "$file")" 2>/dev/null || true
done
3 changes: 3 additions & 0 deletions .rspec_agents_status
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
example_id | status | run_time |
------------------------------------------------- | ------ | ------------ |
./spec/agents/replay/tool_happy_path_spec.rb[1:1] | passed | 3.11 seconds |
7 changes: 7 additions & 0 deletions .rubocop.yml
Original file line number Diff line number Diff line change
Expand Up @@ -199,3 +199,10 @@ RSpec/FilePath:

RSpec/VerifiedDoubles:
Enabled: false

# The agents Tools DSL has its own `describe :tool_name, 'text'`; spec support files use it
RSpec/DescribeSymbol:
Exclude:
- 'spec/support/**/*'
- 'spec/agents/support/**/*'
- 'examples/**/*'
5 changes: 0 additions & 5 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,11 +36,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- Task classes: `SimpleTask`, `SwitchTask`, `ForkTask`, `JoinTask`, `DoWhileTask`, `HttpTask`, `SubWorkflowTask`, `WaitTask`, `TerminateTask`, `SetVariableTask`, `DynamicForkTask`, `JavascriptTask`, `JsonJqTask`, `EventTask`, `HttpPollTask`, `DynamicTask`, `HumanTask`, `StartWorkflowTask`, `KafkaPublishTask`, `WaitForWebhookTask`
- LLM task classes: `LlmChatCompleteTask`, `LlmTextCompleteTask`, `LlmGenerateEmbeddingsTask`, `LlmIndexTextTask`, `LlmIndexDocumentTask`, `LlmSearchIndexTask`, `LlmQueryEmbeddingsTask`, `LlmStoreEmbeddingsTask`, `LlmSearchEmbeddingsTask`, `GenerateImageTask`, `GenerateAudioTask`, `GetDocumentTask`, `ListMcpToolsTask`, `CallMcpToolTask`

### Fixed

- `SchedulerResourceApi#pause_schedule` / `#resume_schedule` now work against both Conductor server families. The client sends `PUT` first and falls back to `GET` on a `405` -- and only on a `405`. OSS Conductor maps these two per-schedule routes `@PutMapping`-only, so the previous `GET`-only calls failed there outright; Orkes Conductor accepts both verbs as of the dual `@RequestMapping(method = {GET, PUT})` added in 2026-07, and is `GET`-only in deployments older than that. `pause_all_schedules` / `resume_all_schedules` remain `GET`, which is how both families map those admin endpoints. Matches the python-sdk, go-sdk, javascript-sdk, csharp-sdk and rust-sdk clients; `spec/conductor/http/api/scheduler_resource_api_spec.rb` pins the whole contract
- `Conductor::AuthenticationSettings` is no longer referenced as `Conductor::Configuration::AuthenticationSettings`, which raised `NameError: uninitialized constant`. The class has always been defined directly under `Conductor`. Fixed in `RactorTaskRunner`'s in-Ractor configuration rebuild (where it was a live failure) and in the `Conductor` / `OrkesClients` doc comments (where it told users to write the broken form)

### Migration Guide

**Before (old DSL):**
Expand Down
10 changes: 8 additions & 2 deletions Gemfile.lock
Original file line number Diff line number Diff line change
Expand Up @@ -47,10 +47,16 @@ GEM
fiber-local (1.1.0)
fiber-storage
fiber-storage (1.0.1)
hana (1.3.7)
hashdiff (1.2.1)
io-console (0.8.2)
io-event (1.16.0)
json (2.7.6)
json_schemer (2.5.0)
bigdecimal
hana (~> 1.3)
regexp_parser (~> 2.0)
simpleidn (~> 0.2)
method_source (1.1.0)
metrics (0.15.0)
net-http-persistent (4.0.8)
Expand Down Expand Up @@ -105,9 +111,9 @@ GEM
rubocop-capybara (~> 2.17)
ruby-progressbar (1.13.0)
ruby2_keywords (0.0.5)
simpleidn (0.3.0)
traces (0.18.2)
unicode-display_width (2.6.0)
vcr (6.1.0)
webmock (3.26.1)
addressable (>= 2.8.0)
crack (>= 0.3.2)
Expand All @@ -120,13 +126,13 @@ PLATFORMS
DEPENDENCIES
async (~> 2.0)
conductor_ruby!
json_schemer (~> 2.0)
prometheus-client (~> 4.0)
pry (~> 0.14)
rake (~> 13.0)
rspec (~> 3.0)
rubocop (~> 1.0)
rubocop-rspec (~> 2.0)
vcr (~> 6.0)
webmock (~> 3.0)
webrick (~> 1.8)

Expand Down
2 changes: 1 addition & 1 deletion conductor_ruby.gemspec
Original file line number Diff line number Diff line change
Expand Up @@ -41,11 +41,11 @@ Gem::Specification.new do |spec|
spec.add_dependency 'json', '>= 2.0'

# Development dependencies (alphabetically sorted)
spec.add_development_dependency 'json_schemer', '~> 2.0'
spec.add_development_dependency 'pry', '~> 0.14'
spec.add_development_dependency 'rake', '~> 13.0'
spec.add_development_dependency 'rspec', '~> 3.0'
spec.add_development_dependency 'rubocop', '~> 1.0'
spec.add_development_dependency 'rubocop-rspec', '~> 2.0'
spec.add_development_dependency 'vcr', '~> 6.0'
spec.add_development_dependency 'webmock', '~> 3.0'
end
17 changes: 17 additions & 0 deletions docs/agents/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Agents

Durable AI agents on Conductor, written in Ruby. Tools run as Conductor worker
tasks, approvals and schedules wait server-side, and an execution survives a
process restart.

Requires Ruby 3.0+ and a Conductor server with an LLM provider configured.

- [Getting started](getting-started.md)
- Concepts: [agents](concepts/agents.md), [tools](concepts/tools.md), [multi-agent](concepts/multi-agent.md),
[guardrails](concepts/guardrails.md), [termination](concepts/termination.md), [callbacks](concepts/callbacks.md),
[stateful](concepts/stateful.md), [streaming and approval](concepts/streaming-hitl.md),
[structured output](concepts/structured-output.md), [runtime modes](concepts/deploy-serve-run.md),
[scheduling](concepts/scheduling.md)
- Reference: [API map](reference/api.md), [runtime](reference/runtime.md), [client](reference/client.md),
[agent fields](reference/agent-definition.md), [wire contract](reference/agent-schema.md)
- [Examples](../../examples/agents/)
26 changes: 26 additions & 0 deletions docs/agents/concepts/agents.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
# Agents

```ruby
require 'conductor/agents'
include Conductor::Agents

tool def get_weather(city: String)
"Weather for #{city}"
end

agent = Agent.new(name: 'weather', model: 'openai/gpt-4o-mini',
instructions: 'Answer concisely.', tools: [:get_weather])
puts agent.call_sync('Weather in Seattle?')
```

- `name` must match `^[a-zA-Z_][a-zA-Z0-9_-]*$`.
- `model` is `provider/model`; the provider is an integration on the server.
Omit it only for sub-agents that inherit the parent's model.
- `instructions` is a String, a Proc evaluated at compile time, or a `PromptTemplate`.
- Per-call options: `session_id:`, `media:`, `context:`, `idempotency_key:`,
`timeout_seconds:`. Don't mutate a shared agent per request.

`call_sync` compiles the agent, starts workers for its tools, and returns the
answer. A tool task stuck in `SCHEDULED` means no process is running its worker.

See [tools](tools.md), [multi-agent](multi-agent.md), [runtime modes](deploy-serve-run.md).
18 changes: 18 additions & 0 deletions docs/agents/concepts/callbacks.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Callbacks

```ruby
agent.callback(:before_tool) { |**event| logger.info(event[:tool_name]) }

class Audit < CallbackHandler
def on_agent_end(**event) = logger.info(event)
end
agent.add_callback(Audit.new)
```

Positions: `before_agent`, `after_agent`, `before_model`, `after_model`,
`before_tool`, `after_tool`. Handler methods are `on_agent_start`, `on_agent_end`,
`on_model_start`, `on_model_end`, `on_tool_start`, `on_tool_end`.

Callbacks run in your process and don't change the workflow. Keep them fast, and
don't make them the only record of anything: a restart drops them. Durable side
effects belong in a tool.
20 changes: 20 additions & 0 deletions docs/agents/concepts/deploy-serve-run.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Runtime modes

| Call | Does |
|---|---|
| `runtime.compile(agent)` | Compile; returns `workflowDef` and `requiredWorkers`. |
| `runtime.deploy(agent)` | Compile and register on the server. |
| `runtime.serve(agent)` | Deploy, run tool workers, block until INT/TERM. |
| `runtime.call_sync(agent, prompt)` | Start, run workers, return the answer. |
| `runtime.call_async(agent, prompt)` | Same, but return an `Execution` immediately. |
| `Execution.find(id)` | Reattach to an existing execution. |

`runtime` is `Conductor::Agents.runtime` (built from ENV) or your own
`AgentRuntime.new(configuration: ...)`.

Local development: `call_sync`. CI/CD: `compile` to inspect, `deploy` to
release. Production: `serve` in a long-lived worker process, then start runs
from anywhere with the [client](../reference/client.md).

Tools must be defined in the process that calls `serve` or `call_*`. Call
`shutdown` before a short-lived process exits.
24 changes: 24 additions & 0 deletions docs/agents/concepts/guardrails.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# Guardrails

Guardrails check input or output before the next step.

```ruby
no_pii = Guardrail.new(name: 'no_pii', position: :output, on_fail: :retry) do |content|
GuardrailResult.new(passed: !content.match?(/\d{3}-\d{2}-\d{4}/), message: 'SSN found')
end

agent = Agent.new(name: 'assistant', model: 'openai/gpt-4o-mini', guardrails: [
RegexGuardrail.new([/\S+@\S+/], mode: :block, position: :output),
LlmGuardrail.new('openai/gpt-4o-mini', 'No medical advice.'),
no_pii
])
agent.redact(%w[password token]) # shortcut for a blocking regex guardrail
```

`on_fail:` is `:raise`, `:retry` (up to `max_retries:`), `:fix`, or `:human`.
Retry only when a new model response could plausibly pass.

Use `RegexGuardrail` for format checks, `LlmGuardrail` for policy, and a block
only when the rule needs application state. Put a guardrail on a `Tool`
(`guardrails:`) when it protects one side effect. Don't send secrets to an LLM
guardrail. Decisions show up in the execution history.
21 changes: 21 additions & 0 deletions docs/agents/concepts/multi-agent.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Multi-agent

```ruby
team = Agent.new(name: 'support', model: 'openai/gpt-4o-mini',
agents: [billing, shipping], strategy: :handoff)
pipeline = researcher >> writer >> editor # sequential
```

Strategies: `:sequential`, `:parallel`, `:handoff` (default; the model picks a
specialist), `:router` (a `router:` agent picks), `:swarm`, `:round_robin`,
`:random`, `:manual`, `:plan_execute`.

- `a.hands_off_to(b, on: 'billing')` hands off when the text appears; `on:` also
takes a callable.
- `Tool.agent(child)` calls a child and returns its answer to the parent
instead of transferring control.
- `plan_execute(name:, tools:, model:, planner_instructions:, fallback_instructions:, fallback_max_turns:)`
builds a planner that writes a plan executed as a durable sub-workflow.

Every open-ended design needs `max_turns:` and a [termination](termination.md)
condition. Each child appears in the execution history.
8 changes: 8 additions & 0 deletions docs/agents/concepts/scheduling.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Scheduling

Deploy the agent, then create a workflow schedule for it with
`OrkesClients#get_scheduler_client`. Agents compile to workflows, so the normal
scheduler applies; there is no agent-specific schedule API.

Use a stable schedule name, a timezone-aware cron, and idempotent input. Pause or
delete the schedule before deleting the agent or stopping its workers.
14 changes: 14 additions & 0 deletions docs/agents/concepts/stateful.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Stateful agents

```ruby
agent = Agent.new(name: 'chat', model: 'openai/gpt-4o-mini', stateful: true,
memory: ConversationMemory.new(max_messages: 50))
agent.call_sync('Hi, I am Ana.', session_id: 'user-42')
agent.call_sync('What is my name?', session_id: 'user-42')
```

`stateful: true` keeps conversation state on the server per `session_id:` and
pins the run's tool workers to its own task domain. After a process restart,
`Execution.find(execution_id)` reattaches.

Bound context with `max_messages:` rather than growing the prompt.
27 changes: 27 additions & 0 deletions docs/agents/concepts/streaming-hitl.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Streaming and approval

```ruby
execution = agent.call_async('Transfer $500', on_event: ->(e) { puts e['event'] })
execution.result(timeout: 120)
```

Events (`message`, `tool_call`, `tool_result`, `waiting`, `done`, `error`)
arrive over SSE; the runtime falls back to polling when SSE is unavailable.
`execution.partial_text` holds the text so far; nothing is final until
`execution.done?`.

Tools with `approval_required: true` pause the execution:

```ruby
agent.on_approval do |request|
puts request.tool_name, request.arguments
request.approve # or request.reject('no'), request.respond(hash)
end
```

`Tool.human` asks a person a question; `Tool.wait_for_message` waits for
`execution.signal(message)`. Without a handler, use `execution.approve`,
`reject`, `signal`, or the [client](../reference/client.md). Approval calls are
safe to repeat.

Call `Conductor::Agents.shutdown` when a short-lived program is done.
17 changes: 17 additions & 0 deletions docs/agents/concepts/structured-output.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Structured output

```ruby
agent = Agent.new(name: 'extract', model: 'openai/gpt-4o-mini',
instructions: 'Extract the invoice.',
output_type: { 'type' => 'object',
'properties' => { 'total' => { 'type' => 'number' } },
'required' => ['total'] })
execution = agent.call_async(text)
execution.result
execution.output # => { 'total' => 42.0 }
```

`output_type:` is a JSON Schema hash, or any object with `to_json_schema`. The
server validates the final answer; on failure the run errors rather than
returning malformed data. Keep the schema small and ask for the same shape in
the instructions.
14 changes: 14 additions & 0 deletions docs/agents/concepts/termination.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Termination

```ruby
agent.stop_when('DONE')
agent.stop_after(messages: 20)
agent.termination = Termination::TokenUsage.new(max_total_tokens: 50_000) |
Termination::TextMention.new('DONE')
```

Conditions: `Termination::MaxMessage`, `StopMessage`, `TextMention`,
`TokenUsage`. Combine with `&` and `|`. Always set `max_turns:` too.

Stop a running execution with `execution.stop` or `AgentClient#stop`. Stopping
does not undo a tool call that already ran.
Loading
Loading