Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

plugpull

Pull the plug on your AI agent in the middle of a run, and see what breaks.

An agent run often lives longer than one request. Processes crash. Deploys restart containers. Users close the tab. Agent frameworks save state so a run can continue after that. plugpull breaks a run on purpose, lets the framework recover, and prints what went wrong.

  • No API key. No LLM is called. In our checks, uv run plugpull gave the same verdicts in 15 runs, and the LangGraph scenario with durability="async" gave the same verdicts in 130 repeated runs. One line can change from run to run: LangGraph durability="async" before-effect prints FAIL run could not be resumed after the kill when the kill lands before LangGraph has saved any checkpoint. With "async", LangGraph saves in the background while the tool runs, and its docs name the risk. When saves are slow this happens more often: with the disk under load on our Mac, 1 run in 40 on Python 3.10 and 3 in 40 on 3.13. CI on Python 3.10 printed a FAIL on that line once (its log cut the text short), and a local run once printed this text; neither kept resume's error output, so for those two this is the most likely cause, not a proven one. DECISIONS.md §9 has the evidence.
  • It does not measure tokens or cost. It shows where the behavior breaks.

Full results table, findings, and reproduction commands: docs/results.md, docs/finding-langgraph-kill-mid-tool.md, docs/finding-openai-agents-4775.md, docs/finding-agent-framework-8292.md.

How is this different?

plugpull runs the framework's own recovery code, not a mock of it: kill-mid-tool compiles a real LangGraph graph with a real SqliteSaver checkpointer, and continues the run from that checkpoint (app.invoke(None, ...)); the lost-session-write scenario runs a real OpenAI Agents SDK Runner and restores it with the SDK's own RunState.from_json (see src/plugpull/targets/langgraph_target.py and src/plugpull/targets/openai_agents_target.py). For kill-mid-tool, that also means an actual SIGKILL mid-tool-call, not a simulated one. lost-session-write does not kill the process; it fails one Session write instead and lets start exit with the error, then runs resume for real — see "Before it gives a verdict" under How kill-mid-tool works and How lost-session-write works for exactly what each scenario does and checks before trusting a result.

As of 2026-09-14, on the OpenAI Agents SDK 0.22.2 (the version on PyPI on that date), it prints FAIL on ack-lost for a case the SDK's own docs say should pass ("the SDK preserves one durable InputItem occurrence" — docs). That gap is tracked upstream in openai/openai-agents-python#4775, open as of 2026-09-14, with a fix proposed in PR #4906, not yet merged as of 2026-09-14. Run uv run plugpull yourself to see the same result.

As of 2026-09-15, on Microsoft Agent Framework 1.18.0 (the version on PyPI on that date), --agent on a functional workflow example prints FAIL: two orders check out concurrently, the process is killed in one order's payment step, and after the checkpoint restore that order is never charged. A separate script without plugpull, repro_without_plugpull.py, where both orders owe money, also shows the other order charged twice. It happens only when the two concurrent calls to the same @step start in a different order on resume than on the first run; with the gather order swapped, it passes. The cause is tracked upstream in microsoft/agent-framework#8292, open with no fix PR as of 2026-09-15. Commands and real output: docs/finding-agent-framework-8292.md.

Try it (about 5 minutes)

You need Python 3.10+ and uv.

git clone https://github.com/chaudl113/plugpull
cd plugpull
uv sync --extra langgraph --extra openai-agents
uv run plugpull

Example output

Real output on LangGraph 1.2.11 and OpenAI Agents SDK 0.22.2, macOS, Python 3.13:

$ uv run plugpull
PASS   kill-mid-tool       before-effect  langgraph 1.2.11 durability=async (default)  the provider saw exactly 1 charge_card effect
FAIL   kill-mid-tool       after-effect   langgraph 1.2.11 durability=async (default)  effect duplicated: the provider saw 2 charge_card effects after resume (expected 1) (documented: make the tool idempotent, see https://docs.langchain.com/oss/python/langgraph/functional-api#idempotency)
PASS   kill-mid-tool       before-effect  langgraph 1.2.11 durability=sync             the provider saw exactly 1 charge_card effect
FAIL   kill-mid-tool       after-effect   langgraph 1.2.11 durability=sync             effect duplicated: the provider saw 2 charge_card effects after resume (expected 1) (documented: make the tool idempotent, see https://docs.langchain.com/oss/python/langgraph/functional-api#idempotency)
PASS   lost-session-write  write-lost     openai-agents 0.22.2                         write 4: all 5 Session items stored exactly once
FAIL   lost-session-write  ack-lost       openai-agents 0.22.2                         write 4: the staged input duplicated: stored 2 times after resume (expected 1) (reported as openai/openai-agents-python#4775; the docs promise the opposite: "the SDK preserves one durable InputItem occurrence", https://openai.github.io/openai-agents-python/results/)

How kill-mid-tool works

The tool in the agent does not call a real service. It calls record_effect(name, key=...), and that call goes to a fake provider inside plugpull. The provider works like a payment API with idempotency keys:

  • two calls with the same effect name and the same key are one effect;
  • the same key on two different effect names is two effects;
  • every call without a key is a new effect.

plugpull kills the agent with SIGKILL at one of two kill points, then resumes the run:

Kill point The process dies when
before-effect the tool has started, and the provider has not recorded the effect yet
after-effect the provider has recorded the effect, and the tool has not returned yet

After resume, for every effect name the run used, the provider must have seen exactly 1 effect. 0 means the effect was lost. 2 or more means it was duplicated. Each kill point gets its own result line.

Before it gives a verdict, plugpull checks that the kill really landed at the kill point. The provider is still holding the tool's call. The operating system says which process made that call, and that process is in the process group plugpull kills. The agent died from SIGKILL. The call closed right after. If any of this is not true, or the operating system does not say which process made the call, the result is ERROR.

What each kind of tool gets:

Tool body before-effect after-effect
no idempotency key PASS FAIL (duplicated)
a key that is the same after resume PASS PASS
a new key every time the tool runs PASS FAIL (duplicated)
writes "done" to its own ledger, then calls the provider FAIL (lost) PASS

tests/ checks this table on a fake framework that runs an unfinished tool again on resume. On LangGraph it checks the first, second (with the tool call id as the key) and fourth rows with durability="async" and "sync". With "async", the test first makes the tool wait until LangGraph has saved the checkpoint the tool runs from, because otherwise the verdict depends on timing: the tool call id can change after resume (see below), and a kill before any checkpoint is saved gives FAIL at both kill points, because resume crashes before any tool runs. A separate test checks that FAIL.

How lost-session-write works

The agent keeps its conversation history in a Session, and the Session's store is a fake store inside plugpull, like a database on another machine. One Session call is one store call. The store keeps exactly what it is given: no keys, no deduplication.

plugpull gives the agent a staged input (a text with a random token) to add to its paused run. First it runs start once with every write answered, to learn which Session writes carry the staged input. Then, for each of those writes, it runs start again, and the store takes that one write and closes the call without answering, at one of two fault points:

Fault point The store
write-lost did not keep the items, and the call fails
ack-lost kept the items, and the call fails

The process is not killed. start saves the run the way the framework's docs say, and exits with the error. resume retries with the store answering normally. After resume, every item ever written must be stored exactly once. 0 means lost, 2 or more means duplicated. Each fault point and each write gets its own result line; the text starts with the write's number.

plugpull picks the write, not the agent. Before it gives a verdict, it checks that start reached the lost write, that the lost write carries the staged input, and that up to it the staged input is in the same write numbers as in the first run; that start failed (non-zero exit); and that it wrote nothing to the Session after the lost write. It does not compare the other items in the writes. If a check fails, the result is ERROR.

On openai-agents, the app follows the SDK docs, sections "Add input before resuming" and "Recover a failed resumed Session write": RunState.add_input, approve, resume with the Session; on error save RunState.to_json, and resume restores it with from_json and runs again.

Why ack-lost fails on OpenAI Agents SDK 0.22.2

When the store kept the write but the call failed, resume writes the staged input again, so the Session holds it twice. This is reported as openai/openai-agents-python#4775. It is not documented behavior: the docs promise the opposite.

"Across serialization, resume, and replay-safe retries, the SDK preserves one durable InputItem occurrence." — OpenAI Agents SDK docs, Results, Add input before resuming

With a proposed fix (the v0.22.2 tag merged with PR #4906's branch, head 58de902e, not merged upstream; its RunState schema is 1.18), both fault points pass. Reproduction command and real output: docs/finding-openai-agents-4775.md.

What the result means

Verdict Meaning
PASS The scenario ran as designed, and the expectation held: the provider saw exactly 1 effect (kill-mid-tool), or every Session item is stored exactly once (lost-session-write).
FAIL The scenario ran as designed, and the expectation broke: an effect or a Session item was lost or duplicated, or the run crashed or hung on resume. This is not always a framework bug: sometimes the framework documents it, and the fix belongs in your code; sometimes the docs promise the opposite. The result text says which, when we know.
ERROR plugpull could not run the scenario as designed. For example, the agent exited before the tool called the provider, the kill did not land at the kill point, start did not reach the lost Session write or did not fail on it, or no Session write carried the staged input. There is no PASS or FAIL. In text mode, the target's stderr is printed under the result line. Please open an issue.

In text mode, plugpull also prints the target's stderr under a FAIL line, but only when resume crashed or hung. It does not print stderr for PASS, or for a FAIL where resume exited 0 (for example, an effect duplicated but the process still exited cleanly). See DECISIONS.md §7.

Options: --json prints machine-readable results, {"schema": 1, "results": [...]} (fields in DECISIONS.md §7). --strict exits with code 1 on any FAIL, for use in CI. --target langgraph or --target openai-agents runs one framework (repeatable). --agent COMMAND runs kill-mid-tool on your own agent (see below).

Exit code Meaning
0 Every scenario ran. Without --strict, FAILs are printed but the code is still 0.
1 --strict was given, and at least one result is FAIL.
2 At least one result is ERROR, or a framework you asked for (or any supported one) is not installed.

Scenarios

Scenario What plugpull does What should be true after Frameworks
kill-mid-tool Kills the process (SIGKILL) inside a tool, before and after the effect reaches the provider, then resumes the run The provider saw exactly 1 effect LangGraph
lost-session-write Fails one Session write that carries a staged input, with the store keeping it (ack-lost) or not (write-lost), then resumes the run Every Session item is stored exactly once OpenAI Agents SDK

Planned next. Each one comes from public bug reports:

Scenario Public reports
cancel-mid-stream: the client cancels while the answer streams langgraph#5672
resume-after-approval: a human approves a tool call, then the run continues langgraph#3275, pydantic-ai#3274
resume-after-upgrade: a run is saved, the framework version changes, then it resumes wanted: send us a report

Why after-effect fails on LangGraph

The tool in plugpull's LangGraph agent has no idempotency key. LangGraph saves a checkpoint before a step runs, not in the middle of it. When the process dies inside a tool, resume starts that step again, so the tool calls the provider a second time. This is documented:

"A task that started but did not finish may run again on that resume, so design side effects to be idempotent. Use idempotency keys or verify existing results to avoid unintended duplication." — LangGraph docs, Functional API, Idempotency

It happens with durability="sync" too. If your tool charges a card, sends an email, or writes to a database, give it an idempotency key that stays the same after resume. Do not mark the work "done" before the call: at before-effect that loses the effect.

A tool call id is not always a stable key on LangGraph

plugpull's LangGraph agent makes a new tool call id each time its model runs, like a real LLM. We gave charge_card the tool call id (InjectedToolCallId) as its idempotency key, and ran kill-mid-tool many times:

durability runs before-effect after-effect
"async" (default) 20 PASS 20 FAIL 9, PASS 11
"sync" 20 PASS 20 PASS 20

Measured on LangGraph 1.2.11, macOS, Python 3.13. Run it yourself (the numbers change from run to run, because they depend on timing):

uv run python tests/measure_langgraph_rig.py tool-call-id async 20
uv run python tests/measure_langgraph_rig.py tool-call-id sync 20

With durability="async" (LangGraph's default), the id is not always saved when the process dies. Resume then runs the model again, the model makes a new id, and the provider sees a second charge with a different key. LangGraph's docs name this risk:

""async": LangGraph persists changes asynchronously while the next step executes. This provides good performance and durability, but there is a small risk that LangGraph does not write checkpoints if the process crashes during execution." — LangGraph docs, Checkpointers, Durability modes

With durability="sync", LangGraph writes the checkpoint before the next step starts, and the key was the same after resume in every run above. With durability="async", the key stayed the same in the tests only when the test made the tool wait until the checkpoint was saved; a real app does not wait, so we have not yet measured a key that stays the same with durability="async" in an ordinary run.

Run it on your own LangGraph agent

plugpull --agent COMMAND runs kill-mid-tool on your agent instead of plugpull's. plugpull runs COMMAND start, kills it inside the tool, then runs COMMAND resume. The verdicts, exit codes and --json format are the ones above. Compared with the same agent without plugpull, you add 7 lines, in 4 steps (examples/langgraph_agent/agent.py marks them # plugpull; it has two versions of the tool, so the record_effect line is there twice):

  1. Report the effect. from plugpull import record_effect, and where the tool calls the outside service, record_effect("charge_card", key=...) with the key it sends (2 lines). Do it in the copy of the tool you test, or behind a switch: outside plugpull the call raises.
  2. Save checkpoints in PLUGPULL_WORKDIR. plugpull sets it to a new directory for each kill point, and puts none of its own files there (1 line). Use a checkpointer that outlives the process, like SqliteSaver, and a thread_id that start and resume both know.
  3. Take start and resume as the last argument. start runs app.invoke(input, config); resume runs app.invoke(None, config) (4 lines).
  4. Run it, from an environment where your agent and plugpull are installed:
uv run plugpull --agent "python examples/langgraph_agent/agent.py no-key"
uv run plugpull --agent "python examples/langgraph_agent/agent.py tool-call-id"

Real output on LangGraph 1.2.11, macOS, Python 3.13 (and the same verdicts on 3.10):

$ uv run plugpull --agent "python examples/langgraph_agent/agent.py no-key"
PASS   kill-mid-tool  before-effect  python examples/langgraph_agent/agent.py no-key  the provider saw exactly 1 charge_card effect
FAIL   kill-mid-tool  after-effect   python examples/langgraph_agent/agent.py no-key  effect duplicated: the provider saw 2 charge_card effects after resume (expected 1)
$ uv run plugpull --agent "python examples/langgraph_agent/agent.py tool-call-id"
PASS   kill-mid-tool  before-effect  python examples/langgraph_agent/agent.py tool-call-id  the provider saw exactly 1 charge_card effect
PASS   kill-mid-tool  after-effect   python examples/langgraph_agent/agent.py tool-call-id  the provider saw exactly 1 charge_card effect

The sample runs with durability="sync", so these verdicts do not depend on timing. With the default "async", a before-effect line can change from run to run, and a tool call id is not always the same key after resume (see the two sections above); both results are true for an "async" agent. A FAIL on your agent has no docs link, because plugpull does not know your framework: for LangGraph, the duplicate is the documented idempotency rule.

ERROR means plugpull could not check your agent: the tool never called record_effect, it called it from a process outside the process group plugpull kills (for example a worker started with start_new_session=True), or COMMAND could not be started. --agent is repeatable, runs no built-in target unless you also pass --target, and is not stable before 0.1. Details and what is not supported: DECISIONS.md §10.

Supported frameworks

Framework Scenario Status
LangGraph kill-mid-tool works
Your own LangGraph agent (--agent) kill-mid-tool works, see above
OpenAI Agents SDK lost-session-write works
Microsoft Agent Framework functional workflow (--agent) kill-mid-tool example in examples/maf_functional_agent, see the finding
Pydantic AI next

Project status

Version 0.0.1. The target interface (start, resume, record_effect(name, *, key=None), PLUGPULL_WORKDIR, --agent, and the Session store calls in plugpull.session_write) and the --json format may change before 0.1. Design decisions are in DECISIONS.md. Contributions: CONTRIBUTING.md.

License

MIT

About

Pull the plug on your AI agent in the middle of a run, and see what breaks.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages