Durable Execution: Resuming Long-Running Agents

The Kitebase support agent picks up ticket KITE-142, “Customer locked out after SSO change”. It looks the ticket up, finds the fix in the help center and drafts a reply to Northwind Studio. The reply has to wait for priya, the support lead, to approve it, the approval gate from the last article. She’s in meetings until the afternoon. At lunch a deploy restarts the worker, and everything the agent knew was in a Python list.

Starting over costs you every model call again. It can also be worse than that. If the crash lands a moment after the reply went out, the restarted agent doesn’t know it sent anything, and Northwind Studio gets the same reply twice.

What you’ll build: a Kitebase agent that saves its progress to SQLite after every step, pauses for priya’s approval without keeping a process alive, gets killed right after sending the reply, and resumes in a fresh process that finishes the job. The customer gets exactly one reply. It runs offline; an API key swaps in Claude.

Why an agent dies with its process

The loop from the agent loop in code keeps its whole state in memory: the messages list with the goal, every model response and every tool result (state and memory covers why that list is the state). Kill the process and it’s gone.

A five-second run can just be retried. A run that takes minutes or hours will get interrupted: a deploy replaces its container, it runs out of memory, the server crashes. A process sitting in input() for three hours waiting for an approval will almost certainly meet one of those.

Durable execution means a run keeps going across restarts because its progress lives in storage, not in memory. Any process can pick it up where the last one stopped. Here’s the run you’ll build, as the rows it leaves in the database:

PROCESSES AGENT.DB: ONE ROW PER CHECKPOINT PROCESS 1 Saves the goal, then 3 model calls and 2 tool calls. Stops at the reply: it needs a human. Exits. Nothing waits. 1 running goal: "Ticket KITE-142 just came in..." 2 running model asks: get_ticket("KITE-142") 3 running result: Northwind Studio, priya 4 running model asks: search_help("locked out...") 5 running result: [account-recovery#2] ... 6 waiting_for_approval model asks: reply_to_customer(...) HOURS LATER priya approves it approvals table: toolu_03 approved by priya PROCESS 2 Loads 6, sends reply #1, then gets killed. 7 never written: the crash came after the reply went out, before the save PROCESS 3 Loads 6 again. The help desk knows the key and sends nothing. 7 running result: {"reply_id": 1, "status": "sent"} 8 done "Replied to Northwind Studio on KITE-142..." Each process starts with a run id and a database file. Everything else comes from the latest row.
One run, three processes. The companion code prints exactly these rows.

The rest of this article builds it in four pieces: save, resume, wait, and make the risky steps safe to repeat.

Step 1: Save a checkpoint after every step

A checkpoint is a saved copy of everything the run needs to continue. For an agent that’s short: the message list, plus a status so you can find runs that aren’t finished. Each run gets a run id, a name you choose that’s unique per run, like run-kite-142, so any process can find its checkpoints.

The example stores checkpoints in SQLite, a database that lives in one file and ships with Python as the sqlite3 module. One row per save:

CREATE TABLE checkpoints (
    run_id     TEXT    NOT NULL,
    seq        INTEGER NOT NULL,   -- 1, 2, 3, ... per run
    status     TEXT    NOT NULL,   -- running | waiting_for_approval | done
    messages   TEXT    NOT NULL,   -- the whole message list, as JSON
    created_at TEXT    NOT NULL,
    PRIMARY KEY (run_id, seq)
);

Saving adds a row with the next seq:

def save(self, run_id: str, status: str, messages: list[dict]) -> int:
    # `with self.db` commits when the block ends. Once save() returns, the row
    # is on disk and a crash can't take it back.
    with self.db:
        (last,) = self.db.execute(
            "SELECT COALESCE(MAX(seq), 0) FROM checkpoints WHERE run_id = ?", (run_id,)).fetchone()
        self.db.execute("INSERT INTO checkpoints VALUES (?, ?, ?, ?, ?)",
                        (run_id, last + 1, status, json.dumps(messages), now()))
    return last + 1

A commit is the moment the database makes a write permanent. Until then, a crash throws the write away. The loop only moves on after save() returns, so every step it has taken is on disk.

When do you save? Every time the message list grows. That’s twice per step: once when the model answers, and once when the tools return. Here’s the first process of the worked example. Every line is one row written:

== Process 1: start the run ==
checkpoint 1  running               user: Ticket KITE-142 just came in. Look it up, find the ...
loaded checkpoint 1 (running)
checkpoint 2  running               model asks for get_ticket(ticket_id="KITE-142")
checkpoint 3  running               tool result: {"id": "KITE-142", "title": "Customer locked...
checkpoint 4  running               model asks for search_help(query="locked out after SSO ch...
checkpoint 5  running               tool result: [account-recovery#2] Locked out of your acco...
checkpoint 6  waiting_for_approval  model asks for reply_to_customer(ticket_id="KITE-142", me...
paused: reply_to_customer needs a human, so process 1 exits

Saving the model’s answer before running its tools looks redundant. It isn’t. Checkpoint 6 fixes the decision: which tool, which arguments, and the tool call’s id, toolu_03. Step 4 depends on that id never changing.

Checkpoints are small. Checkpoint 6 is 1.9 KB of JSON, and all eight rows for the finished run come to about 9 KB.

One gotcha: the SDK’s response blocks are Python objects, so json.dumps(response.content) fails with Object of type ToolUseBlock is not JSON serializable. Convert each block with block.model_dump(mode="json", exclude_none=True) first, which keeps every field the API wants back. The example’s to_dict does this.

Why save the whole message list every time instead of just the new message?

Because resume becomes one read with nothing to rebuild. The latest row is the complete state, so there’s no way to end up with half a history.

It does repeat data: row 8 contains everything in rows 1 to 7. At 9 KB a run that doesn’t matter. When runs get long, keep only the latest row per run, or store one row per message and rebuild the list with an ORDER BY. Keeping every row also gives you a history of how the run got there, handy for debugging.

Step 2: Resume by looking at the last message

Resuming means: load the latest checkpoint and keep going. What keeps this simple is that the loop carries nothing between iterations except the message list. Each time round, it looks at the last message and decides what to do from that alone:

AGENT.DB latest row LOOK AT THE LAST MESSAGE status_of() Goal or tool results e.g. checkpoint 5 Tool call, approved checkpoint 6, after priya Gated call, no decision checkpoint 6, before priya Model text, no tools checkpoint 8 CALL THE MODEL append its reply RUN THE TOOLS append the results STOP AND WAIT the process can exit DONE nothing left to do save a checkpoint, then look again A fresh run and a resumed run go through the same code.
The same loop starts a run and resumes one. Only the last message decides the next step.

In code, that decision is a small function:

def status_of(store: Store, run_id: str, messages: list[dict]) -> str:
    calls = pending_tool_calls(messages)  # tool calls in the last message, if it's the model's
    if not calls:
        return "done" if messages[-1]["role"] == "assistant" else "running"
    if any(c["name"] in NEEDS_APPROVAL and store.decision(run_id, c["id"]) is None for c in calls):
        return "waiting_for_approval"
    return "running"

And the loop, trimmed of logging, the step limit and the hook that simulates the crash:

def run_agent(store, model, kb, run_id) -> str:
    messages = store.latest(run_id).messages
    while (status := status_of(store, run_id, messages)) == "running":
        calls = pending_tool_calls(messages)
        if calls:
            results = [run_tool(kb, store, run_id, call) for call in calls]
            messages.append({"role": "user", "content": results})
        else:
            response = model(messages)
            messages.append({"role": "assistant", "content": [to_dict(b) for b in response.content]})
        checkpoint(store, run_id, status_of(store, run_id, messages), messages)
    return status  # "done", or "waiting_for_approval": nothing waits in memory

There’s no separate resume path: starting a run is saving checkpoint 1 and calling run_agent; resuming is calling it again with the same run id. The step limit counts the model’s messages in the checkpoint, so a crash-looping run can’t reset its 10-call limit.

Something still has to call run_agent after a crash. The example’s third process asks the store for every run whose latest checkpoint isn’t done:

SELECT run_id FROM checkpoints c
WHERE seq = (SELECT MAX(seq) FROM checkpoints WHERE run_id = c.run_id)
  AND status != 'done'

In production, a worker runs this when it starts up, and a sweeper (a small job on a timer) runs it every few minutes to catch runs whose worker died. Hold that thought; it’s where durable-workflow engines come in.

Step 3: Wait for a human without holding a process

Guardrails and human in the loop put a gate in front of reply_to_customer: the customer sees it, so a person approves it first. The naive version blocks on input("Approve? ") until priya answers, holding a worker for hours and losing the wait on the next restart.

With checkpoints, waiting is just a status. When the loop meets a gated call with no decision, status_of returns waiting_for_approval, the loop saves checkpoint 6 and returns, and the process exits. Nothing is running. Priya’s decision arrives through the support tool, which writes one row:

store.decide("run-kite-142", "toolu_03", approved=True, decided_by="priya")

Any later process finds the decision and carries on; a three-day wait costs one row. A rejection becomes an error tool_result (“A support lead rejected this reply: Loop in their IT admin first.”), so the model hears why instead of retrying the same call.

The gotcha is the human who never answers. A run in waiting_for_approval sits there forever, cheaply, while the customer waits. Give every wait a deadline: the sweeper finds runs that have waited more than, say, four working hours and pings someone else or rejects the reply with a note.

Step 4: Make side effects safe to replay

A side effect is anything a tool does to the world outside your program: sending a reply, charging a card, assigning a ticket. Reads like get_ticket and search_help have none, so running them twice is harmless. Writes are where resume gets dangerous.

Here’s the worst moment to crash. Priya approved the reply. Process 2 resumes, sends the reply to the help desk, and gets killed before it saves checkpoint 7. Process 3 loads the latest checkpoint, which is still 6. It says: reply pending, approved. So process 3 sends it again.

WITHOUT AN IDEMPOTENCY KEY agent help desk process 2: send_reply stored as reply #1 CRASH before checkpoint 7 process 3 loads checkpoint 6 send_reply again, no key stored as reply #2 Northwind Studio gets 2 replies WITH KEY run-kite-142:toolu_03 agent help desk process 2: send_reply + key stored as reply #1, key saved CRASH before checkpoint 7 process 3 loads checkpoint 6 send_reply again, same key key seen: returns reply #1, sends nothing Northwind Studio gets 1 reply The key is built from the run id and the tool call's id, both saved in checkpoint 6, so a retry sends the same key.
Saving more often can't close this gap. The help desk and your database are two different systems.

You can’t fix this by saving more often. There will always be a moment after the help desk accepted the reply and before your database heard about it. They’re two separate systems, with no way to commit to both at once.

The fix is the one from retries, timeouts and failure handling: make the write idempotent, meaning doing it twice has the same effect as doing it once. You send an idempotency key, a unique string for this one logical action, and the receiver remembers the keys it has seen. A repeat gets the original result back and changes nothing. The example’s help desk works the way Stripe’s API and many others do:

def send_reply(self, ticket_id: str, message: str, idempotency_key: str | None) -> dict:
    if idempotency_key:
        row = self.db.execute("SELECT id FROM replies WHERE idempotency_key = ?",
                              (idempotency_key,)).fetchone()
        if row:
            return {"reply_id": row[0], "duplicate": True}
    ...  # insert the reply and return its new id

For a retry inside one process, any key made before the first attempt works. For a resume, the key must also survive the crash, so build it from things that are in the checkpoint:

# Stable across resumes: the run id and the tool call's id are both in the checkpoint.
args["idempotency_key"] = f"{run_id}:{call['id']}"

This is why Step 1 saved the model’s answer before running the tools. Process 3 reads toolu_03 from checkpoint 6 and sends the same key as process 2. If the loop had asked the model again instead, the new response would carry a new tool call id, maybe with different wording, and the help desk would see a brand-new reply. Here are processes 2 and 3:

== Process 2: resume, then get killed right after the reply goes out ==
loaded checkpoint 6 (waiting_for_approval)
  help desk: sent reply #1 to Northwind Studio (key run-kite-142:toolu_03)
CRASH: killed after reply_to_customer ran, before its result was saved

== Process 3: a fresh worker picks up unfinished runs ==
loaded checkpoint 6 (waiting_for_approval)
  help desk: key run-kite-142:toolu_03 seen before, returned reply #1, sent nothing
checkpoint 7  running               tool result: {"reply_id": 1, "status": "sent"}
checkpoint 8  done                  model: Replied to Northwind Studio on KITE-142 with the S...

== Result ==
Replies Northwind Studio got on KITE-142: 1
Model calls across all three processes: 4

Four model calls for a four-step run: the resumes repeated none of them.

What if the system doesn’t take a key? Check before you act: look for a reply on KITE-142 carrying a marker from this run, and skip the send if it’s there. If you can’t check either, record in your own database that the action is starting before you call it. After a crash, a “started, never finished” row tells a human which action to verify, instead of a silent duplicate.

Why can't the agent just check its checkpoint to see if the reply went out?

Because the checkpoint is the thing that didn’t get written. The crash happened after the help desk accepted the reply and before checkpoint 7, so checkpoint 6 honestly says “not sent yet”.

Only the system that did the work knows it happened. That’s why the protection has to live there, as an idempotency key or a lookup, and not in your own records.

When to reach for a durable-workflow engine

Everything above is a couple of hundred lines of plain Python and two tables. For one worker and a modest number of long runs, that’s enough. It’s also a small version of what a durable-workflow engine does: a service that runs your long-lived code and makes it survive crashes for you. Temporal is the best-known; Restate, Inngest and DBOS are others in the same space.

The model, as Temporal’s documentation describes it: you write the workflow as ordinary code, and every call to the outside world goes through the engine as an activity. The engine records each activity’s result in the workflow’s history. After a crash it runs your workflow code again from the top, handing back recorded results for activities that already finished instead of calling them again. In this article’s terms, activities are the model and tool calls, and the history is the checkpoints table.

Two consequences carry over from what you’ve built. Workflow code must be deterministic, meaning it takes the same path every time it’s replayed; don’t read the clock or random numbers directly, because the engine gives you replay-safe versions. And an activity can still run twice if the crash lands at the wrong moment, the same gap as Step 4, so the idempotency keys stay.

What you’d be buying is everything the example leaves out:

  • Automatic resume, instead of your own sweeper calling unfinished_runs().
  • One worker per run. Two workers resuming the same run would both run its tools. Preventing that needs a lease: a lock with an expiry, so a dead worker’s claim runs out.
  • Timers. “Escalate if nobody approves within four hours” becomes a line of workflow code, not a cron job.
  • Per-step retries and timeouts, configured instead of hand-written.
  • A UI showing where every run is and what it’s waiting for.

The default: start with the checkpoint table. Move to an engine when you find yourself writing leases, timers and a retry scheduler around it. At that point you’re building one, and theirs has had years of production use.

Try it yourself

The companion example is the whole worked example in three simulated processes. Each one opens the database files fresh and carries nothing over in memory.

Download the runnable example (zip)

cd 07-durable-execution
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. Set USE_IDEMPOTENCY_KEY = False in main.py. Process 3 sends the reply again, and the result line says Northwind Studio got 2.
  2. In main(), approve with approved=False, note="Loop in their IT admin first.". Nothing is sent, and the model’s last message says why.
  3. After a run, look inside: sqlite3 agent.db "select seq, status, length(messages) from checkpoints". Each row is the previous one plus one message.

pip install pytest && pytest -q runs the offline tests. They check what each checkpoint holds, that a resume lands on the right step, that no model call is repeated, and that the customer gets exactly one reply.

Common beginner mistakes

  • Saving only at the end. A crash then loses the whole run. Save every time the message list grows.
  • Holding a process open for a human. An input() or a polling loop ties up a worker and dies with it. Save, exit, and let the approval resume the run.
  • Asking the model again on resume. You pay for the call, you may get a different decision, and the new tool call id makes a new idempotency key. Reuse the saved decision.
  • A fresh key per attempt. f"{run_id}:{uuid4()}" looks safe and dedupes nothing. Build the key only from values in the checkpoint.
  • Checkpoints on a disk that dies with the container. A SQLite file inside a container without a mounted volume vanishes on the very deploy you’re guarding against. On containers, use Postgres or a mounted volume.
  • Keeping the step counter in memory. Count from the checkpoint, or a crash-looping run resets its limit every restart.

Questions you will face in production

“SQLite or Postgres for the checkpoints?” SQLite on one machine with a disk that survives restarts. Postgres once workers run on several machines or containers come and go: every worker can reach it, and the code barely changes.

“What happens to runs in flight when I deploy a change to the tools?” They resume under the new code with checkpoints written by the old code. Renaming reply_to_customer breaks every run waiting on a call to it. Keep old tool names and argument shapes working until those runs finish, or store a version in each checkpoint and let old runs finish on code that still understands them.

“Does resuming cost tokens?” Only for work that wasn’t saved. In the worked example the resumes repeat no model calls: four calls for four steps. A crash between the model answering and the save costs that one call again. The history itself isn’t free, though: each call sends the whole message list, resumed or not.

Check your understanding

The agent crashes after search_help returns but before checkpoint 5 is saved. What happens on resume, and is it a problem?

The latest checkpoint is 4, whose last message asks for search_help, so the resumed loop runs the search again and saves checkpoint 5. That’s fine: a search has no side effects. You lose a few milliseconds, not a reply.

A teammate changes the key to f"{run_id}:{uuid.uuid4()}" "to make it more unique". What breaks?

Every resume makes a new key, so the help desk never sees a repeat and the crash from Step 4 sends a second reply. A key only protects you if every attempt at the same action produces the same key. Build it from the run id and the saved tool call id.

Why does the loop save the model's response before running the tools, instead of once per step after the tools finish?

To freeze the decision. After a crash, the resumed run reads the saved tool call and runs exactly that: same arguments, same id, same idempotency key. Without it, the loop has to ask the model again, pays for that call, may get different wording, and gets a new tool call id, so the help desk sees a new reply.

A deploy restarts three workers at once, and each one resumes every unfinished run on startup. What goes wrong?

All three can pick up run-kite-142 and run its tools. The idempotency key saves the customer from three replies, but all three workers append their own checkpoints to the same run, so its history stops describing one run. You need a lease, so only one worker owns a run at a time. It’s one of the main things a durable-workflow engine does for you.

What to remember

  • An agent’s state is its message list. Save it to a database every time it grows, and commit before moving on.
  • Resuming is the same loop: load the latest checkpoint and let the last message decide the next step.
  • Waiting for a human is a status in a row, not a process sitting in memory. Give every wait a deadline.
  • Saving can’t close the gap between an external write and your checkpoint. Side effects need an idempotency key built from the saved run id and tool call id.
  • Start with a checkpoint table. Move to a durable-workflow engine when you’re about to write leases, timers and retries yourself.

What to study next

A run that survives crashes can still loop quietly, burn tokens or stall in waiting_for_approval without anyone noticing. Observability and cost control for agents covers tracing what a long run actually did, step by step, and capping what it spends.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.