Retries, Timeouts, and Failure Handling
Your Kitebase support agent is handling ticket KITE-142: a Northwind Studio user is locked out after their company changed its single sign-on setup. It reads the ticket, finds the right help article, and calls reply_to_customer. Kitebase is slow that afternoon. The email goes out, but the response takes too long, your code decides the call failed, and it tries again. The customer gets the same email twice.
Nothing crashed, and every line did what it was told. What makes failures inside an agent run different is the second decision-maker in the loop, the model. It can route around a failure, or make it worse by calling the same broken thing again and again.
What you’ll build: the Kitebase agent loop, run nine times against a fake Kitebase that drops connections, hangs, and loses responses, with a scripted fake model playing Claude. Each run breaks one thing, and a test checks what the loop does about it.
This builds on the loop from Plan, Act, Observe: The Agent Loop in Code and the tools from Giving an Agent Tools. Retries and backoff for the Claude call itself are in Your First LLM Integration, and the SDK already does them, so this article is about everything else.
The run you’ll break
The agent has the five Kitebase tools every article in this topic uses: search_help(query), get_ticket(ticket_id), search_tickets(query, status), assign_ticket(ticket_id, assignee) and reply_to_customer(ticket_id, message). When nothing goes wrong, handling KITE-142 takes four model turns:
[1] get_ticket {"ticket_id": "KITE-142"}
-> {"id": "KITE-142", "title": "Customer locked out after SSO change", ...
[2] search_help {"query": "locked out SSO"}
-> [{"id": "account-recovery#2", "text": "Locked out of your account: If...
[3] reply_to_customer {"ticket_id": "KITE-142", "message": "Your workspace uses single sign...
-> {"reply_id": "rpl_001", "sent_to": "Northwind Studio"}
DONE (end_turn) after 4 step(s), 1 email(s) sent
A turn (or step) is one model call plus the tools it asks for. Three of these tools only read. Two of them change something: assign_ticket and reply_to_customer, the write tools. Keep that split in mind, because it decides what’s safe to retry.
In the companion code you switch on faults per tool: "drop" fails before doing anything, and "hang" does the work and then answers too late. A scripted fake model plays Claude, so every run is repeatable and the tests need no API key.
Retry in code, or hand it to the model?
When a tool fails in a normal service, your code has two choices: retry or give up. In an agent there’s a third. You can put the failure in the tool result, mark it with is_error: true, and let the model decide what to do. That’s the rule from the agent loop article: never let a tool exception crash the loop.
But handing every blip to the model is wasteful. A dropped connection that would work on the second try doesn’t need a model turn, which costs seconds and money, to find that out. So each failure goes one of three ways:
- Retry in code when the failure is transient (it might work next time: a dropped connection, a timeout, a 503) and the tool is safe to call twice. Do it once or twice, quickly, and the model never knows.
- Return it to the model as
is_errorwhen the failure is permanent (the same input fails the same way: a ticket that doesn’t exist, a bad argument), or when your retries ran out. The model can change the input, use another tool, or give up and say why. - Stop the run when neither can fix it. That’s the second half of this article.
Here’s the function that runs every tool call. It returns the content and an is_error flag, and it never raises:
def run_tool(kitebase, name: str, args: dict, log=print) -> tuple[str, bool]:
attempts = TOOL_ATTEMPTS if retry_safe(name, kitebase) else 1
for attempt in range(1, attempts + 1):
try:
return json.dumps(call_with_timeout(kitebase, name, args, TOOL_TIMEOUT)), False
except ToolError as e: # permanent: the model has to change something
return f"Error: {e}", True
except (TimeoutError, ConnectionError) as e: # transient: might work next time
problem = f"{e} ({type(e).__name__})"
if attempt < attempts:
time.sleep(BACKOFF)
except Exception as e: # a bug in the tool, or arguments it can't take
return f"Error: {name} crashed: {type(e).__name__}: {e}", True
tries = "1 try" if attempts == 1 else f"{attempts} tries"
error = f"Error: {name} failed after {tries}: {problem}. Kitebase may be slow or down."
if name not in READ_ONLY:
error += " The change may have gone through anyway, so check before you repeat it."
return error, True
TOOL_ATTEMPTS is 2, so one retry. That’s deliberately few: after a failed tool, the model can go another way, and every extra retry is time the user waits before that can happen.
Run python main.py flaky and search_help drops its connection once:
[2] search_help {"query": "locked out SSO"}
attempt 1: connection reset by Kitebase (KitebaseUnavailable). Retrying in 0.1s.
-> [{"id": "account-recovery#2", "text": "Locked out of your account: If...
The model’s transcript shows a normal result. The blip cost 0.1 seconds, not a model turn.
A timeout the model can act on
A tool that fails fast is easy. A tool that never answers is worse, because the loop runs one tool at a time and everything waits for it. The model can’t recover from a failure it hasn’t been told about yet.
So every tool that talks to a network gets a timeout: a cap on how long you wait for an answer. The example wraps each call in a thread and waits a fixed time for the result:
def call_with_timeout(kitebase, name: str, args: dict, timeout: float):
future = _pool.submit(kitebase.call, name, args)
try:
return future.result(timeout=timeout)
except futures.TimeoutError:
# This ends the wait. The thread keeps going, and so does the work it started.
raise TimeoutError(f"no response after {timeout}s") from None
The example uses 0.2 seconds so the demo runs fast. Real tools want 5 to 10. In your own code, set the timeout on the HTTP client the tool uses (httpx.Client(timeout=5.0), requests.get(url, timeout=5)), which is simpler than threads.
In python main.py slow, get_ticket hangs on both tries. The model gets this tool result, straight from the run:
{"type": "tool_result", "tool_use_id": "toolu_01",
"content": "Error: get_ticket failed after 2 tries: no response after 0.2s (TimeoutError). "
"Kitebase may be slow or down.",
"is_error": True}
And the scripted model does what a model can and a retry loop can’t: it goes another way.
[2] search_tickets {"query": "locked out SSO", "status": "in_progress"}
-> [{"id": "KITE-142", "title": "Customer locked out after SSO change", ...
The error text is a prompt now, so write it for the model. “failed after 2 tries: no response after 0.2s, Kitebase may be slow” tells it the input was fine and the service wasn’t, so trying something else makes sense. Error: 500 tells it nothing, and it may well send the same call again.
The gotcha is in the comment inside call_with_timeout: a timeout stops you waiting, not the work. The request already went to Kitebase. For a read, that’s just a wasted request. For a write, it’s the email in the opening.
Retrying a write: idempotency keys
A timed-out write leaves you not knowing if it happened. Your code can’t tell “sent, answer lost” from “never sent”.
The fix is to make the write idempotent: calling it twice with the same input has the same effect as calling it once. assign_ticket already is, because it sets the assignee: setting it to priya twice leaves it priya. Sending an email isn’t, so it needs an idempotency key: a string you send with the request so the API can spot a repeat and return the first result instead of acting again. Stripe’s API works this way, and the fake Kitebase does too.
Derive the key from the arguments, not from a random ID per attempt, so every retry of the same reply carries the same key:
def reply_key(ticket_id: str, message: str) -> str:
"""Same ticket and same text give the same key, so a retry can't send twice."""
return hashlib.sha256(f"{ticket_id}\n{message.strip()}".encode()).hexdigest()[:16]
For KITE-142 and the SSO reply, that’s 8d9c503c3702ab68. MCP Server Best Practices builds the same key into an MCP server. What’s new in an agent is that there are two retriers: your loop, and the model.
The top run is python main.py duplicate. The reply goes out, the response is lost, the loop retries with the same key, and Kitebase answers rpl_001 without sending anything. One email.
The bottom run is python main.py no-keys, and it’s the one to remember. The loop does the careful thing: a write with no key isn’t safe to retry, so retry_safe says no, and the model gets an error saying the reply may have gone through. The model calls reply_to_customer again anyway, which a model will sometimes do after a timeout. Second email.
[3] reply_to_customer {"ticket_id": "KITE-142", "message": "Your workspace uses single sign...
-> is_error: Error: reply_to_customer failed after 1 try: no response after 0.2s ... The change
may have gone through anyway, so check before you repeat it.
[4] reply_to_customer {"ticket_id": "KITE-142", "message": "Your workspace uses single sign...
-> {"reply_id": "rpl_002", "sent_to": "Northwind Studio"}
DONE (end_turn) after 5 step(s), 2 email(s) sent
Telling the model to check first helps, but you can’t count on it. The key is what protects the customer, and it works whichever retrier sends the repeat, as long as the text is the same. If the model rewrites the message before resending, the key changes and a second email goes out. That’s why the error text also tells the model the first one may have landed.
What if the API I'm calling has no idempotency keys?
Look before you write. Before sending, fetch the ticket’s recent replies and skip the send if an identical one is already there. It’s slower, since it’s one extra read per write, and there’s a small window where two sends can both look and both miss. But it turns “definitely twice” into “almost never twice”.
When the model loops on a failing call
Handing errors back to the model has a failure mode of its own. Sometimes the model reads the error and calls the exact same thing again. In python main.py loop, the model has transposed two digits and asks for KITE-412:
[1] get_ticket {"ticket_id": "KITE-412"}
-> is_error: Error: No ticket KITE-412. Ticket ids look like KITE-142.
[2] get_ticket {"ticket_id": "KITE-412"}
-> is_error: Error: No ticket KITE-412. Ticket ids look like KITE-142.
That’s a permanent failure, so the loop never retries it in code. But the model does, and each repeat is a paid turn that can’t succeed. The fix is a repeat guard: count failures per exact call (tool name plus input), and once the same call has failed twice, stop running it.
signature = (block.name, json.dumps(block.input, sort_keys=True))
if failed_calls[signature] >= MAX_SAME_FAILURES:
blocked += 1
if blocked > 1: # it ignored the first warning
return give_up("stuck", f"Claude kept repeating {block.name} ...", step, writes)
content, is_error = ("Error: not run. This exact call already failed 2 times. Don't repeat "
"it: change the input, use another tool, or stop and explain."), True
else:
content, is_error = run_tool(kitebase, block.name, block.input, log)
if is_error:
failed_calls[signature] += 1
The first blocked call gets a nudge instead of the real tool, which is often enough for the model to change course. The second one ends the run. In the example, calls 3 and 4 never reach Kitebase, and the run stops with FAILED (stuck) after 4 step(s). The guard only counts failed calls: repeating a search that worked is fine.
Step and time limits
Some runs fail without a single error. In python main.py runaway, every tool call works, and the model just never finishes: it keeps searching the help center with slightly different queries. The repeat guard never fires, because nothing failed. What stops it is the step limit from the agent loop article: at most MAX_STEPS model turns per run.
A step limit caps cost. It doesn’t cap how long the user waits, because one turn can be fast or slow. That’s what the deadline is for: a wall-clock limit for the whole run, checked before every model call.
In the code, both are the first thing each turn checks, and the last thing the loop does:
for step in range(1, max_steps + 1):
if clock() - start > deadline:
return give_up("deadline", f"the run hit its {deadline:g}s time limit", step, writes)
response = client.messages.create(...)
...
return give_up("max_steps", f"the run hit its {max_steps}-step limit", max_steps, writes)
The defaults are 8 steps and 30 seconds. KITE-142 needs 4 or 5 steps, so 8 leaves room for a detour. Set yours from real runs: log the steps and seconds successful runs take, and set each limit at about twice the slowest normal one.
The deadline is checked between steps, so a turn with hung tools can overshoot it by that turn’s timeouts and retries, as the 1.0s in the diagram shows. That’s why the per-tool timeout matters even with a deadline: the deadline decides when to stop, and the tool timeout decides how late you notice.
Stop reasons mid-run: max_tokens and refusal
The model’s own turn can fail too, and it arrives as a successful response. Only stop_reason tells you, as the LLM integration article showed. Inside a loop, two values need handling:
"max_tokens": the turn was cut off at yourmax_tokenscap. In an agent, that can happen halfway through writing a tool call, so the last block may be atool_usewith half its arguments. Never run it. Anthropic’s docs say to retry the request with a highermax_tokens."refusal": the model declined to continue, and the content may be empty. Sending the same request again usually gets the same answer. Stop, and hand the ticket to a person. (For some refusals, the API docs suggest retrying on a different Claude model; that’s a production fallback, not a default.)
if response.stop_reason == "max_tokens" and max_tokens == MAX_TOKENS:
# A cut-off turn may hold a half-written tool call. Don't run it; redo the turn with room.
max_tokens *= 2
continue
if response.stop_reason == "max_tokens":
return give_up("cut_off", f"Claude's reply was cut off twice, even at max_tokens={max_tokens}", ...)
if response.stop_reason == "refusal":
return give_up("refusal", "Claude declined to continue", step, writes)
The continue skips appending the cut-off turn, so the model redoes it from the same point with twice the room. One retry is enough: if 2,048 tokens still can’t hold one turn, something else is wrong, and more tokens won’t fix it.
[2] cut off at max_tokens=1024. Redoing the turn with 2048.
[3] search_help {"query": "locked out SSO"}
Failing gracefully
What a run says when it stops matters as much as when. “Error: max_steps” in a log helps nobody. The person who picks the ticket up needs three things: why it stopped, what already changed, and that it’s now theirs.
def give_up(reason: str, why: str, step: int, writes: list[str]) -> RunResult:
done = "; ".join(writes) if writes else "none, nothing was sent to the customer or changed"
message = f"I couldn't finish this: {why}.\nChanges made: {done}.\nA person needs to pick this up."
return RunResult("failed", reason, message, step, writes)
The loop records every successful write as it goes. So from python main.py deadline:
FAILED (deadline) after 3 step(s), 0 email(s) sent
I couldn't finish this: the run hit its 0.8s time limit.
Changes made: none, nothing was sent to the customer or changed.
A person needs to pick this up.
“Changes made” is the line that matters most. A support lead who reads “a reply was sent: rpl_001” won’t send another. One who reads “nothing was sent” knows the customer is still waiting. The short reason (deadline, stuck, …) is for dashboards; message is for people.
The same goes for the model call itself. If Claude’s API is still failing after the SDK’s own retries, the loop catches anthropic.APIError and returns the same kind of note, with reason model_unavailable, instead of a stack trace.
Try it yourself
The companion example is the loop from this article, a fake Kitebase with switchable faults, and nine scenarios. It runs offline in about three seconds.
Download the runnable example (zip)
cd 05-retries-and-failures
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py # all nine scenarios
python main.py no-keys # just one
pip install pytest && pytest -q runs 13 offline tests, one or more per failure path. With ANTHROPIC_API_KEY set, python main.py slow --live runs a tool-fault scenario against real Claude.
Then try these:
- Set
TOOL_ATTEMPTS = 1inmain.pyand runpython main.py flaky. The dropped connection now reaches the model asis_error. A real model would spend a turn on something one retry would have fixed. - In
retry_safe, returnTrueforreply_to_customereven without a key, and runpython main.py no-keys. Still two emails, but now your code sends the second one instead of the model. - Set
RUN_DEADLINE = 0.3and runpython main.py slow. One hung tool, tried twice, uses the whole budget before the model gets a second turn.
Common beginner mistakes
- Retrying everything in code. A permanent error fails the same way every time, and each retry delays the model’s chance to fix the input. Retry transient failures only.
- No timeout on a tool. The loop waits for each tool, so one hung call freezes the whole run, and the step limit never fires because no step finishes.
- Retrying a write with no idempotency key. Or with a new random key per attempt, which is the same thing. Derive the key from the arguments.
- Running a cut-off tool call. A
max_tokensturn can end mid-argument. Redo the turn; never execute what came back. - A failure message with no “what changed”. Whoever takes over needs to know whether the customer already got a reply.
Questions you will face in production
“The run hit a limit halfway through. Can it pick up where it left off?” Not with an in-memory loop like this one. The message list lives in the process, so a crash or a deploy loses it. Saving each step so a run can resume is durable execution, in Durable Execution: Resuming Long-Running Agents. Resuming is also the strongest reason for idempotency keys: a resumed run replays the last step, and the key makes that replay safe.
“What if a tool is slow for everyone, not just this run?” Then every run pays the timeout, again and again. Track tool failure rates, and when one tool keeps failing, take it out of the tool list for a while, or tell the model up front that it’s down. The pattern is a circuit breaker: after N failures in a row, stop calling for a minute. Observability and Cost Control for Agents covers the tracing you need to see it happening.
Check your understanding
search_tickets returns a 503 on the first try. Should your code retry it, or return it to the model?
Retry it in code, once, after a short wait. A 503 is transient and search_tickets only reads, so repeating it is harmless. If the retry fails too, return an is_error result saying the service seems down, so the model can take another route or stop.
assign_ticket times out. Is it safe for your code to retry?
Yes. It sets the assignee rather than toggling or adding, so running it twice with priya leaves the ticket assigned to priya. That’s why retry_safe allows it without a key. A tool like add_comment wouldn’t be safe: two retries, two comments.
Your loop never retries reply_to_customer, but customers still sometimes get two emails. How?
The model retries it. After a timeout, it gets is_error and may call the tool again with the same message. Not retrying in your code only removes one of the two retriers. Add an idempotency key derived from the ticket and the message, so a repeat from either one is recognized.
A turn comes back with stop_reason "max_tokens" and a tool_use block at the end. What do you do with that block?
Nothing. Its input may be cut off mid-argument. Don’t run it and don’t append the turn to the messages. Redo the same turn with a higher max_tokens, once, and give up with a clear message if it’s cut off again.
What to remember
- A failed tool goes one of three ways: a quick retry in code (transient and safe to repeat), an
is_errorresult the model can act on, or the end of the run. - Every tool gets a timeout. A timeout stops the waiting, not the work, so a timed-out write may have happened.
- Retry writes only with an idempotency key derived from the arguments. The model retries too, so the key has to protect you from both.
- A repeat guard stops identical failing calls; a step limit caps cost; a deadline caps the wait. You need all three.
- Check
stop_reasonevery turn: redo amax_tokensturn once with more room, and stop on arefusal. - When a run gives up, say why, say what already changed, and hand it to a person.
What to study next
Every guard here assumes the model is trying to do the right thing and failing. Next is what to do when it’s about to do the wrong thing successfully: Guardrails and Human in the Loop adds approval gates in front of write tools like reply_to_customer. After that, Durable Execution makes a run survive a crash and resume, which is where idempotency keys pay off most.
Further reading
- Anthropic: Handling stop reasons. Every
stop_reasonvalue, including what to do with amax_tokenscut-off during tool use. - AWS Builders’ Library: Timeouts, retries, and backoff with jitter. The standard writeup on why retries need limits and jitter.
- Stripe: Idempotent requests. How idempotency keys work in a real API you can copy the pattern from.
- Anthropic: Building effective agents. Context on the agent loop these failures live inside, including why you want stopping conditions.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.