Pull the plug on your AI agent in the middle of a run, and see what breaks.
An agent run often lives longer than one request. Processes crash. Deploys restart containers. Users close the tab. Agent frameworks save state so a run can continue after that. plugpull breaks a run on purpose, lets the framework recover, and prints what went wrong.
- No API key. No LLM is called. In our checks,
uv run plugpullgave the same verdicts in 15 runs, and the LangGraph scenario withdurability="async"gave the same verdicts in 130 repeated runs. One line can change from run to run: LangGraphdurability="async"before-effectprintsFAIL run could not be resumed after the killwhen the kill lands before LangGraph has saved any checkpoint. With"async", LangGraph saves in the background while the tool runs, and its docs name the risk. When saves are slow this happens more often: with the disk under load on our Mac, 1 run in 40 on Python 3.10 and 3 in 40 on 3.13. CI on Python 3.10 printed a FAIL on that line once (its log cut the text short), and a local run once printed this text; neither kept resume's error output, so for those two this is the most likely cause, not a proven one. DECISIONS.md §9 has the evidence. - It does not measure tokens or cost. It shows where the behavior breaks.
Full results table, findings, and reproduction commands: docs/results.md, docs/finding-langgraph-kill-mid-tool.md, docs/finding-openai-agents-4775.md, docs/finding-agent-framework-8292.md.
plugpull runs the framework's own recovery code, not a mock of it: kill-mid-tool compiles a
real LangGraph graph with a real SqliteSaver checkpointer, and continues the run from that
checkpoint (app.invoke(None, ...)); the lost-session-write scenario runs a real OpenAI
Agents SDK Runner and restores it with the
SDK's own RunState.from_json (see src/plugpull/targets/langgraph_target.py and
src/plugpull/targets/openai_agents_target.py). For kill-mid-tool, that also means an actual
SIGKILL mid-tool-call, not a simulated one. lost-session-write does not kill the process; it
fails one Session write instead and lets start exit with the error, then runs resume for
real — see "Before it gives a verdict" under How kill-mid-tool
works and How lost-session-write
works for exactly what each scenario does and checks before
trusting a result.
As of 2026-09-14, on the OpenAI Agents SDK 0.22.2 (the version on PyPI on that date), it prints
FAIL on ack-lost for a case the SDK's own docs say should pass ("the SDK preserves one durable
InputItem occurrence" — docs). That
gap is tracked upstream in openai/openai-agents-python#4775,
open as of 2026-09-14, with a fix proposed in PR #4906,
not yet merged as of 2026-09-14. Run uv run plugpull yourself to see the same result.
As of 2026-09-15, on Microsoft Agent Framework 1.18.0 (the version on PyPI on that date),
--agent on a functional workflow example prints FAIL: two orders check out concurrently, the
process is killed in one order's payment step, and after the checkpoint restore that order is
never charged. A separate script without plugpull,
repro_without_plugpull.py, where both
orders owe money, also shows the other order charged twice. It happens only when the two
concurrent calls to the same @step start in a different order on resume than on the first run;
with the gather order swapped, it passes. The cause is
tracked upstream in microsoft/agent-framework#8292,
open with no fix PR as of 2026-09-15. Commands and real output:
docs/finding-agent-framework-8292.md.
You need Python 3.10+ and uv.
git clone https://github.com/chaudl113/plugpull
cd plugpull
uv sync --extra langgraph --extra openai-agents
uv run plugpullReal output on LangGraph 1.2.11 and OpenAI Agents SDK 0.22.2, macOS, Python 3.13:
$ uv run plugpull
PASS kill-mid-tool before-effect langgraph 1.2.11 durability=async (default) the provider saw exactly 1 charge_card effect
FAIL kill-mid-tool after-effect langgraph 1.2.11 durability=async (default) effect duplicated: the provider saw 2 charge_card effects after resume (expected 1) (documented: make the tool idempotent, see https://docs.langchain.com/oss/python/langgraph/functional-api#idempotency)
PASS kill-mid-tool before-effect langgraph 1.2.11 durability=sync the provider saw exactly 1 charge_card effect
FAIL kill-mid-tool after-effect langgraph 1.2.11 durability=sync effect duplicated: the provider saw 2 charge_card effects after resume (expected 1) (documented: make the tool idempotent, see https://docs.langchain.com/oss/python/langgraph/functional-api#idempotency)
PASS lost-session-write write-lost openai-agents 0.22.2 write 4: all 5 Session items stored exactly once
FAIL lost-session-write ack-lost openai-agents 0.22.2 write 4: the staged input duplicated: stored 2 times after resume (expected 1) (reported as openai/openai-agents-python#4775; the docs promise the opposite: "the SDK preserves one durable InputItem occurrence", https://openai.github.io/openai-agents-python/results/)
The tool in the agent does not call a real service. It calls record_effect(name, key=...),
and that call goes to a fake provider inside plugpull. The provider works like a payment
API with idempotency keys:
- two calls with the same effect name and the same key are one effect;
- the same key on two different effect names is two effects;
- every call without a key is a new effect.
plugpull kills the agent with SIGKILL at one of two kill points, then resumes the run:
| Kill point | The process dies when |
|---|---|
before-effect |
the tool has started, and the provider has not recorded the effect yet |
after-effect |
the provider has recorded the effect, and the tool has not returned yet |
After resume, for every effect name the run used, the provider must have seen exactly 1 effect. 0 means the effect was lost. 2 or more means it was duplicated. Each kill point gets its own result line.
Before it gives a verdict, plugpull checks that the kill really landed at the kill point. The provider is still holding the tool's call. The operating system says which process made that call, and that process is in the process group plugpull kills. The agent died from SIGKILL. The call closed right after. If any of this is not true, or the operating system does not say which process made the call, the result is ERROR.
What each kind of tool gets:
| Tool body | before-effect |
after-effect |
|---|---|---|
| no idempotency key | PASS | FAIL (duplicated) |
| a key that is the same after resume | PASS | PASS |
| a new key every time the tool runs | PASS | FAIL (duplicated) |
| writes "done" to its own ledger, then calls the provider | FAIL (lost) | PASS |
tests/ checks this table on a fake framework that runs an unfinished tool again on resume.
On LangGraph it checks the first, second (with the tool call id as the key) and fourth rows with
durability="async" and "sync". With "async", the test first makes the tool wait until
LangGraph has saved the checkpoint the tool runs from, because otherwise the verdict depends on
timing: the tool call id can change after resume
(see below), and a kill before any
checkpoint is saved gives FAIL at both kill points, because resume crashes before any tool runs.
A separate test checks that FAIL.
The agent keeps its conversation history in a Session, and the Session's store is a fake store inside plugpull, like a database on another machine. One Session call is one store call. The store keeps exactly what it is given: no keys, no deduplication.
plugpull gives the agent a staged input (a text with a random token) to add to its paused run.
First it runs start once with every write answered, to learn which Session writes carry the
staged input. Then, for each of those writes, it runs start again, and the store takes that
one write and closes the call without answering, at one of two fault points:
| Fault point | The store |
|---|---|
write-lost |
did not keep the items, and the call fails |
ack-lost |
kept the items, and the call fails |
The process is not killed. start saves the run the way the framework's docs say, and exits
with the error. resume retries with the store answering normally. After resume, every item
ever written must be stored exactly once. 0 means lost, 2 or more means duplicated. Each
fault point and each write gets its own result line; the text starts with the write's number.
plugpull picks the write, not the agent. Before it gives a verdict, it checks that start
reached the lost write, that the lost write carries the staged input, and that up to it the
staged input is in the same write numbers as in the first run; that start failed (non-zero
exit); and that it wrote nothing to the Session after the lost write. It does not compare the
other items in the writes. If a check fails, the result is ERROR.
On openai-agents, the app follows the SDK docs, sections "Add input before resuming" and
"Recover a failed resumed Session write": RunState.add_input, approve, resume with the Session;
on error save RunState.to_json, and resume restores it with from_json and runs again.
When the store kept the write but the call failed, resume writes the staged input again, so the Session holds it twice. This is reported as openai/openai-agents-python#4775. It is not documented behavior: the docs promise the opposite.
"Across serialization, resume, and replay-safe retries, the SDK preserves one durable
InputItemoccurrence." — OpenAI Agents SDK docs, Results, Add input before resuming
With a proposed fix (the v0.22.2 tag merged with PR
#4906's branch, head 58de902e, not
merged upstream; its RunState schema is 1.18), both fault points pass. Reproduction command and
real output: docs/finding-openai-agents-4775.md.
| Verdict | Meaning |
|---|---|
| PASS | The scenario ran as designed, and the expectation held: the provider saw exactly 1 effect (kill-mid-tool), or every Session item is stored exactly once (lost-session-write). |
| FAIL | The scenario ran as designed, and the expectation broke: an effect or a Session item was lost or duplicated, or the run crashed or hung on resume. This is not always a framework bug: sometimes the framework documents it, and the fix belongs in your code; sometimes the docs promise the opposite. The result text says which, when we know. |
| ERROR | plugpull could not run the scenario as designed. For example, the agent exited before the tool called the provider, the kill did not land at the kill point, start did not reach the lost Session write or did not fail on it, or no Session write carried the staged input. There is no PASS or FAIL. In text mode, the target's stderr is printed under the result line. Please open an issue. |
In text mode, plugpull also prints the target's stderr under a FAIL line, but only when resume crashed or hung. It does not print stderr for PASS, or for a FAIL where resume exited 0 (for example, an effect duplicated but the process still exited cleanly). See DECISIONS.md §7.
Options: --json prints machine-readable results, {"schema": 1, "results": [...]} (fields in
DECISIONS.md §7). --strict
exits with code 1 on any FAIL, for use in CI. --target langgraph or --target openai-agents
runs one framework (repeatable). --agent COMMAND runs kill-mid-tool on your own agent
(see below).
| Exit code | Meaning |
|---|---|
| 0 | Every scenario ran. Without --strict, FAILs are printed but the code is still 0. |
| 1 | --strict was given, and at least one result is FAIL. |
| 2 | At least one result is ERROR, or a framework you asked for (or any supported one) is not installed. |
| Scenario | What plugpull does | What should be true after | Frameworks |
|---|---|---|---|
kill-mid-tool |
Kills the process (SIGKILL) inside a tool, before and after the effect reaches the provider, then resumes the run | The provider saw exactly 1 effect | LangGraph |
lost-session-write |
Fails one Session write that carries a staged input, with the store keeping it (ack-lost) or not (write-lost), then resumes the run |
Every Session item is stored exactly once | OpenAI Agents SDK |
Planned next. Each one comes from public bug reports:
| Scenario | Public reports |
|---|---|
cancel-mid-stream: the client cancels while the answer streams |
langgraph#5672 |
resume-after-approval: a human approves a tool call, then the run continues |
langgraph#3275, pydantic-ai#3274 |
resume-after-upgrade: a run is saved, the framework version changes, then it resumes |
wanted: send us a report |
The tool in plugpull's LangGraph agent has no idempotency key. LangGraph saves a checkpoint before a step runs, not in the middle of it. When the process dies inside a tool, resume starts that step again, so the tool calls the provider a second time. This is documented:
"A task that started but did not finish may run again on that resume, so design side effects to be idempotent. Use idempotency keys or verify existing results to avoid unintended duplication." — LangGraph docs, Functional API, Idempotency
It happens with durability="sync" too. If your tool charges a card, sends an email, or writes
to a database, give it an idempotency key that stays the same after resume. Do not mark the
work "done" before the call: at before-effect that loses the effect.
plugpull's LangGraph agent makes a new tool call id each time its model runs, like a real LLM.
We gave charge_card the tool call id (InjectedToolCallId) as its idempotency key, and ran
kill-mid-tool many times:
durability |
runs | before-effect |
after-effect |
|---|---|---|---|
"async" (default) |
20 | PASS 20 | FAIL 9, PASS 11 |
"sync" |
20 | PASS 20 | PASS 20 |
Measured on LangGraph 1.2.11, macOS, Python 3.13. Run it yourself (the numbers change from run to run, because they depend on timing):
uv run python tests/measure_langgraph_rig.py tool-call-id async 20
uv run python tests/measure_langgraph_rig.py tool-call-id sync 20With durability="async" (LangGraph's default), the id is not always saved when the process
dies. Resume then runs the model again, the model makes a new id, and the provider sees a
second charge with a different key. LangGraph's docs name this risk:
"
"async": LangGraph persists changes asynchronously while the next step executes. This provides good performance and durability, but there is a small risk that LangGraph does not write checkpoints if the process crashes during execution." — LangGraph docs, Checkpointers, Durability modes
With durability="sync", LangGraph writes the checkpoint before the next step starts, and the
key was the same after resume in every run above. With durability="async", the key stayed the
same in the tests only when the test made the tool wait until the checkpoint was saved; a real
app does not wait, so we have not yet measured a key that stays the same with
durability="async" in an ordinary run.
plugpull --agent COMMAND runs kill-mid-tool on your agent instead of plugpull's. plugpull
runs COMMAND start, kills it inside the tool, then runs COMMAND resume. The verdicts, exit
codes and --json format are the ones above. Compared with the same agent without plugpull, you
add 7 lines, in 4 steps
(examples/langgraph_agent/agent.py marks them # plugpull;
it has two versions of the tool, so the record_effect line is there twice):
- Report the effect.
from plugpull import record_effect, and where the tool calls the outside service,record_effect("charge_card", key=...)with the key it sends (2 lines). Do it in the copy of the tool you test, or behind a switch: outside plugpull the call raises. - Save checkpoints in
PLUGPULL_WORKDIR. plugpull sets it to a new directory for each kill point, and puts none of its own files there (1 line). Use a checkpointer that outlives the process, likeSqliteSaver, and athread_idthatstartandresumeboth know. - Take
startandresumeas the last argument.startrunsapp.invoke(input, config);resumerunsapp.invoke(None, config)(4 lines). - Run it, from an environment where your agent and plugpull are installed:
uv run plugpull --agent "python examples/langgraph_agent/agent.py no-key"
uv run plugpull --agent "python examples/langgraph_agent/agent.py tool-call-id"Real output on LangGraph 1.2.11, macOS, Python 3.13 (and the same verdicts on 3.10):
$ uv run plugpull --agent "python examples/langgraph_agent/agent.py no-key"
PASS kill-mid-tool before-effect python examples/langgraph_agent/agent.py no-key the provider saw exactly 1 charge_card effect
FAIL kill-mid-tool after-effect python examples/langgraph_agent/agent.py no-key effect duplicated: the provider saw 2 charge_card effects after resume (expected 1)
$ uv run plugpull --agent "python examples/langgraph_agent/agent.py tool-call-id"
PASS kill-mid-tool before-effect python examples/langgraph_agent/agent.py tool-call-id the provider saw exactly 1 charge_card effect
PASS kill-mid-tool after-effect python examples/langgraph_agent/agent.py tool-call-id the provider saw exactly 1 charge_card effect
The sample runs with durability="sync", so these verdicts do not depend on timing. With the
default "async", a before-effect line can change from run to run, and a tool call id is not
always the same key after resume (see the two sections above); both results are true for an
"async" agent. A FAIL on your agent has no docs link, because plugpull does not know your
framework: for LangGraph, the duplicate is the
documented idempotency rule.
ERROR means plugpull could not check your agent: the tool never called record_effect, it called
it from a process outside the process group plugpull kills (for example a worker started with
start_new_session=True), or COMMAND could not be started. --agent is repeatable, runs no
built-in target unless you also pass --target, and is not stable before 0.1. Details and what is
not supported: DECISIONS.md §10.
| Framework | Scenario | Status |
|---|---|---|
| LangGraph | kill-mid-tool |
works |
Your own LangGraph agent (--agent) |
kill-mid-tool |
works, see above |
| OpenAI Agents SDK | lost-session-write |
works |
Microsoft Agent Framework functional workflow (--agent) |
kill-mid-tool |
example in examples/maf_functional_agent, see the finding |
| Pydantic AI | next |
Version 0.0.1. The target interface (start, resume, record_effect(name, *, key=None),
PLUGPULL_WORKDIR, --agent, and the Session store calls in plugpull.session_write) and the
--json format may change before 0.1.
Design decisions are in DECISIONS.md. Contributions: CONTRIBUTING.md.
MIT