An observable, bounded AI agent control loop designed for incident triage and root cause investigation across distributed services.
This repository implements Problem 4: Observable Agent Loop for the Caygnus Product Engineering technical assessment. It demonstrates:
- Autonomous Tool Selection & Execution: Dynamic multi-step tool calling across telemetry metrics, application logs, service health status, and remediation runbooks.
- Strict Evidence Grounding (AC6): Decouples empirical tool evidence (facts) from agent inferences and recommendations.
- Bounded Execution & Safety (AC5): Enforces configurable step limits (
max_steps) to prevent infinite execution traps and runaway resource consumption. - Resilient Failure Handling (AC4): Catches tool runtime exceptions and enables adaptive fallback strategies without crashing the orchestrator.
- Trace Observability & Sanitization (AC3): Structured audit logging with automated redaction of sensitive credentials, API keys, and authorization tokens.
- Hermetic Testing: 100% offline unit and integration test suite requiring zero paid API keys or network dependencies.
observable-agent-loop-solution/
├── main.py # CLI entry point with pre-configured scenario runners
├── README.md # Project documentation and quick-start guide
├── SUBMISSION.md # Complete Caygnus submission report
├── src/
│ ├── __init__.py
│ ├── agent.py # Agent control loop orchestrator and bounded step engine
│ ├── cli.py # Terminal formatting, colored timeline, and report visualizer
│ ├── model.py # Pluggable model adapters (Dynamic & Scripted Deterministic)
│ ├── tools.py # Tool base class, schema validation, and ToolRegistry
│ ├── types.py # Core dataclasses, state enums, step records, and traces
│ └── data/
│ ├── __init__.py
│ └── synthetic_incident_data.py # Realistic production outage incident fixtures
└── tests/
├── __init__.py
└── test_agent.py # 12 unit & integration tests covering AC1 through AC6
- Python 3.10 or higher.
- No external packages or third-party dependencies required.
# Multi-Step Root Cause Investigation (AC2, AC3, AC6)
python main.py --scenario multi-step
# Single-Step Tool Selection (AC1)
python main.py --scenario single-step
# Tool Failure & Adaptive Fallback (AC4)
python main.py --scenario tool-failure
# Bounded Step Limit Enforcement (AC5)
python main.py --scenario step-limit
# Raw JSON Summary Output
python main.py --scenario multi-step --json
# Interactive Custom Prompt Mode
python main.py --scenario interactiveRun the complete test suite using Python's built-in unittest runner:
python -m unittest discover -s tests -p "test_*.py" -vExpected output:
test_ac1_appropriate_tool_selection ... ok
test_ac2_multi_step_investigation ... ok
test_ac3_observable_ordered_trace ... ok
test_ac4_tool_failure_and_recovery ... ok
test_ac5_execution_limit_enforcement ... ok
test_ac6_evidence_vs_conclusions_separation ... ok
test_search_logs_tool_validation_success ... ok
test_tool_validation_invalid_enum ... ok
test_tool_validation_invalid_type ... ok
test_tool_validation_missing_required_parameter ... ok
test_unregistered_tool_execution ... ok
test_redact_secrets ... ok
----------------------------------------------------------------------
Ran 12 tests in 0.003s
OK