Retries, Timeouts, and Failure Handling

Your Kitebase support agent is handling ticket KITE-142: a Northwind Studio user is locked out after their company changed its single sign-on setup. It reads the ticket, finds the right help article, and calls reply_to_customer. Kitebase is slow that afternoon. The email goes out, but the response takes too long, your code decides the call failed, and it tries again. The customer gets the same email twice.

Nothing crashed, and every line did what it was told. What makes failures inside an agent run different is the second decision-maker in the loop, the model. It can route around a failure, or make it worse by calling the same broken thing again and again.

What you’ll build: the Kitebase agent loop, run nine times against a fake Kitebase that drops connections, hangs, and loses responses, with a scripted fake model playing Claude. Each run breaks one thing, and a test checks what the loop does about it.

This builds on the loop from Plan, Act, Observe: The Agent Loop in Code and the tools from Giving an Agent Tools. Retries and backoff for the Claude call itself are in Your First LLM Integration, and the SDK already does them, so this article is about everything else.

The run you’ll break

The agent has the five Kitebase tools every article in this topic uses: search_help(query), get_ticket(ticket_id), search_tickets(query, status), assign_ticket(ticket_id, assignee) and reply_to_customer(ticket_id, message). When nothing goes wrong, handling KITE-142 takes four model turns:

  [1] get_ticket {"ticket_id": "KITE-142"}
      -> {"id": "KITE-142", "title": "Customer locked out after SSO change", ...
  [2] search_help {"query": "locked out SSO"}
      -> [{"id": "account-recovery#2", "text": "Locked out of your account: If...
  [3] reply_to_customer {"ticket_id": "KITE-142", "message": "Your workspace uses single sign...
      -> {"reply_id": "rpl_001", "sent_to": "Northwind Studio"}
  DONE (end_turn) after 4 step(s), 1 email(s) sent

A turn (or step) is one model call plus the tools it asks for. Three of these tools only read. Two of them change something: assign_ticket and reply_to_customer, the write tools. Keep that split in mind, because it decides what’s safe to retry.

In the companion code you switch on faults per tool: "drop" fails before doing anything, and "hang" does the work and then answers too late. A scripted fake model plays Claude, so every run is repeatable and the tests need no API key.

Retry in code, or hand it to the model?

When a tool fails in a normal service, your code has two choices: retry or give up. In an agent there’s a third. You can put the failure in the tool result, mark it with is_error: true, and let the model decide what to do. That’s the rule from the agent loop article: never let a tool exception crash the loop.

But handing every blip to the model is wasteful. A dropped connection that would work on the second try doesn’t need a model turn, which costs seconds and money, to find that out. So each failure goes one of three ways:

What failed Your code What Claude gets What happens next DROPPED search_help connection reset RETRY IN CODE transient, and a read: try again after 0.1s A NORMAL RESULT account-recovery#2 from attempt 2 CARRIES ON Claude never saw the failure HUNG get_ticket no reply in 0.2s RETRY, THEN STOP attempt 2 hangs too: hand it to Claude IS_ERROR: TRUE "get_ticket failed after 2 tries ..." CHANGES COURSE search_tickets finds KITE-142 PERMANENT get_ticket KITE-412 no such ticket DON'T RETRY the same input fails the same way IS_ERROR: TRUE "No ticket KITE-412." and it calls it again, twice RUN STOPS guard blocks call 3, handoff note at call 4 Retry in code only when the failure is transient and the tool is safe to repeat. Everything else goes to Claude as is_error. Stop the run when neither can fix it.
Where each failure ends up. Every value comes from running the example.
  • Retry in code when the failure is transient (it might work next time: a dropped connection, a timeout, a 503) and the tool is safe to call twice. Do it once or twice, quickly, and the model never knows.
  • Return it to the model as is_error when the failure is permanent (the same input fails the same way: a ticket that doesn’t exist, a bad argument), or when your retries ran out. The model can change the input, use another tool, or give up and say why.
  • Stop the run when neither can fix it. That’s the second half of this article.

Here’s the function that runs every tool call. It returns the content and an is_error flag, and it never raises:

def run_tool(kitebase, name: str, args: dict, log=print) -> tuple[str, bool]:
    attempts = TOOL_ATTEMPTS if retry_safe(name, kitebase) else 1
    for attempt in range(1, attempts + 1):
        try:
            return json.dumps(call_with_timeout(kitebase, name, args, TOOL_TIMEOUT)), False
        except ToolError as e:  # permanent: the model has to change something
            return f"Error: {e}", True
        except (TimeoutError, ConnectionError) as e:  # transient: might work next time
            problem = f"{e} ({type(e).__name__})"
            if attempt < attempts:
                time.sleep(BACKOFF)
        except Exception as e:  # a bug in the tool, or arguments it can't take
            return f"Error: {name} crashed: {type(e).__name__}: {e}", True
    tries = "1 try" if attempts == 1 else f"{attempts} tries"
    error = f"Error: {name} failed after {tries}: {problem}. Kitebase may be slow or down."
    if name not in READ_ONLY:
        error += " The change may have gone through anyway, so check before you repeat it."
    return error, True

TOOL_ATTEMPTS is 2, so one retry. That’s deliberately few: after a failed tool, the model can go another way, and every extra retry is time the user waits before that can happen.

Run python main.py flaky and search_help drops its connection once:

  [2] search_help {"query": "locked out SSO"}
      attempt 1: connection reset by Kitebase (KitebaseUnavailable). Retrying in 0.1s.
      -> [{"id": "account-recovery#2", "text": "Locked out of your account: If...

The model’s transcript shows a normal result. The blip cost 0.1 seconds, not a model turn.

A timeout the model can act on

A tool that fails fast is easy. A tool that never answers is worse, because the loop runs one tool at a time and everything waits for it. The model can’t recover from a failure it hasn’t been told about yet.

So every tool that talks to a network gets a timeout: a cap on how long you wait for an answer. The example wraps each call in a thread and waits a fixed time for the result:

def call_with_timeout(kitebase, name: str, args: dict, timeout: float):
    future = _pool.submit(kitebase.call, name, args)
    try:
        return future.result(timeout=timeout)
    except futures.TimeoutError:
        # This ends the wait. The thread keeps going, and so does the work it started.
        raise TimeoutError(f"no response after {timeout}s") from None

The example uses 0.2 seconds so the demo runs fast. Real tools want 5 to 10. In your own code, set the timeout on the HTTP client the tool uses (httpx.Client(timeout=5.0), requests.get(url, timeout=5)), which is simpler than threads.

In python main.py slow, get_ticket hangs on both tries. The model gets this tool result, straight from the run:

{"type": "tool_result", "tool_use_id": "toolu_01",
 "content": "Error: get_ticket failed after 2 tries: no response after 0.2s (TimeoutError). "
            "Kitebase may be slow or down.",
 "is_error": True}

And the scripted model does what a model can and a retry loop can’t: it goes another way.

  [2] search_tickets {"query": "locked out SSO", "status": "in_progress"}
      -> [{"id": "KITE-142", "title": "Customer locked out after SSO change", ...

The error text is a prompt now, so write it for the model. “failed after 2 tries: no response after 0.2s, Kitebase may be slow” tells it the input was fine and the service wasn’t, so trying something else makes sense. Error: 500 tells it nothing, and it may well send the same call again.

The gotcha is in the comment inside call_with_timeout: a timeout stops you waiting, not the work. The request already went to Kitebase. For a read, that’s just a wasted request. For a write, it’s the email in the opening.

Retrying a write: idempotency keys

A timed-out write leaves you not knowing if it happened. Your code can’t tell “sent, answer lost” from “never sent”.

The fix is to make the write idempotent: calling it twice with the same input has the same effect as calling it once. assign_ticket already is, because it sets the assignee: setting it to priya twice leaves it priya. Sending an email isn’t, so it needs an idempotency key: a string you send with the request so the API can spot a repeat and return the first result instead of acting again. Stripe’s API works this way, and the fake Kitebase does too.

Derive the key from the arguments, not from a random ID per attempt, so every retry of the same reply carries the same key:

def reply_key(ticket_id: str, message: str) -> str:
    """Same ticket and same text give the same key, so a retry can't send twice."""
    return hashlib.sha256(f"{ticket_id}\n{message.strip()}".encode()).hexdigest()[:16]

For KITE-142 and the SSO reply, that’s 8d9c503c3702ab68. MCP Server Best Practices builds the same key into an MCP server. What’s new in an agent is that there are two retriers: your loop, and the model.

WITH AN IDEMPOTENCY KEY: python main.py duplicate Claude Agent loop Kitebase reply_to_customer(KITE-142, ...) send, key 8d9c503c... email 1 sent no reply in 0.2s retry after 0.1s, same key key seen: rpl_001, no email rpl_001 1 email to Northwind Studio NO KEY: python main.py no-keys Claude Agent loop Kitebase reply_to_customer(KITE-142, ...) send, no key email 1 sent no reply in 0.2s a write with no key: don't retry is_error: "may have gone through" reply_to_customer again send, no key email 2 sent 2 emails to Northwind Studio
Not retrying in your code isn't enough. The model retries too.

The top run is python main.py duplicate. The reply goes out, the response is lost, the loop retries with the same key, and Kitebase answers rpl_001 without sending anything. One email.

The bottom run is python main.py no-keys, and it’s the one to remember. The loop does the careful thing: a write with no key isn’t safe to retry, so retry_safe says no, and the model gets an error saying the reply may have gone through. The model calls reply_to_customer again anyway, which a model will sometimes do after a timeout. Second email.

  [3] reply_to_customer {"ticket_id": "KITE-142", "message": "Your workspace uses single sign...
      -> is_error: Error: reply_to_customer failed after 1 try: no response after 0.2s ... The change
         may have gone through anyway, so check before you repeat it.
  [4] reply_to_customer {"ticket_id": "KITE-142", "message": "Your workspace uses single sign...
      -> {"reply_id": "rpl_002", "sent_to": "Northwind Studio"}
  DONE (end_turn) after 5 step(s), 2 email(s) sent

Telling the model to check first helps, but you can’t count on it. The key is what protects the customer, and it works whichever retrier sends the repeat, as long as the text is the same. If the model rewrites the message before resending, the key changes and a second email goes out. That’s why the error text also tells the model the first one may have landed.

What if the API I'm calling has no idempotency keys?

Look before you write. Before sending, fetch the ticket’s recent replies and skip the send if an identical one is already there. It’s slower, since it’s one extra read per write, and there’s a small window where two sends can both look and both miss. But it turns “definitely twice” into “almost never twice”.

When the model loops on a failing call

Handing errors back to the model has a failure mode of its own. Sometimes the model reads the error and calls the exact same thing again. In python main.py loop, the model has transposed two digits and asks for KITE-412:

  [1] get_ticket {"ticket_id": "KITE-412"}
      -> is_error: Error: No ticket KITE-412. Ticket ids look like KITE-142.
  [2] get_ticket {"ticket_id": "KITE-412"}
      -> is_error: Error: No ticket KITE-412. Ticket ids look like KITE-142.

That’s a permanent failure, so the loop never retries it in code. But the model does, and each repeat is a paid turn that can’t succeed. The fix is a repeat guard: count failures per exact call (tool name plus input), and once the same call has failed twice, stop running it.

signature = (block.name, json.dumps(block.input, sort_keys=True))
if failed_calls[signature] >= MAX_SAME_FAILURES:
    blocked += 1
    if blocked > 1:  # it ignored the first warning
        return give_up("stuck", f"Claude kept repeating {block.name} ...", step, writes)
    content, is_error = ("Error: not run. This exact call already failed 2 times. Don't repeat "
                         "it: change the input, use another tool, or stop and explain."), True
else:
    content, is_error = run_tool(kitebase, block.name, block.input, log)
    if is_error:
        failed_calls[signature] += 1

The first blocked call gets a nudge instead of the real tool, which is often enough for the model to change course. The second one ends the run. In the example, calls 3 and 4 never reach Kitebase, and the run stops with FAILED (stuck) after 4 step(s). The guard only counts failed calls: repeating a search that worked is fine.

Step and time limits

Some runs fail without a single error. In python main.py runaway, every tool call works, and the model just never finishes: it keeps searching the help center with slightly different queries. The repeat guard never fires, because nothing failed. What stops it is the step limit from the agent loop article: at most MAX_STEPS model turns per run.

A step limit caps cost. It doesn’t cap how long the user waits, because one turn can be fast or slow. That’s what the deadline is for: a wall-clock limit for the whole run, checked before every model call.

deadline: every Kitebase call hangs, and the run has 0.8s Claude turn 1 turn 2 Kitebase get_ticket retry search_tickets retry 0.1s wait 0.1s wait 0s 0.2 0.4 0.6 0.8 1.0 0.8s limit Before turn 3: 1.0s is past 0.8s. Stop and write the handoff note. Each hung call costs 0.2s + 0.1s + 0.2s. Two of them and the budget is gone. runaway: every call works, Claude never answers, 8-step limit 1 get_ticket 2 search 3 search 4 search 5 search 6 search 7 search 8 search search_help with four queries, cycled. Every call works, so the repeat guard never fires. After turn 8, the step limit ends the run.
Two runs that no single tool error explains. The deadline catches slow, the step limit catches endless.

In the code, both are the first thing each turn checks, and the last thing the loop does:

for step in range(1, max_steps + 1):
    if clock() - start > deadline:
        return give_up("deadline", f"the run hit its {deadline:g}s time limit", step, writes)
    response = client.messages.create(...)
    ...
return give_up("max_steps", f"the run hit its {max_steps}-step limit", max_steps, writes)

The defaults are 8 steps and 30 seconds. KITE-142 needs 4 or 5 steps, so 8 leaves room for a detour. Set yours from real runs: log the steps and seconds successful runs take, and set each limit at about twice the slowest normal one.

The deadline is checked between steps, so a turn with hung tools can overshoot it by that turn’s timeouts and retries, as the 1.0s in the diagram shows. That’s why the per-tool timeout matters even with a deadline: the deadline decides when to stop, and the tool timeout decides how late you notice.

Stop reasons mid-run: max_tokens and refusal

The model’s own turn can fail too, and it arrives as a successful response. Only stop_reason tells you, as the LLM integration article showed. Inside a loop, two values need handling:

  • "max_tokens": the turn was cut off at your max_tokens cap. In an agent, that can happen halfway through writing a tool call, so the last block may be a tool_use with half its arguments. Never run it. Anthropic’s docs say to retry the request with a higher max_tokens.
  • "refusal": the model declined to continue, and the content may be empty. Sending the same request again usually gets the same answer. Stop, and hand the ticket to a person. (For some refusals, the API docs suggest retrying on a different Claude model; that’s a production fallback, not a default.)
if response.stop_reason == "max_tokens" and max_tokens == MAX_TOKENS:
    # A cut-off turn may hold a half-written tool call. Don't run it; redo the turn with room.
    max_tokens *= 2
    continue
if response.stop_reason == "max_tokens":
    return give_up("cut_off", f"Claude's reply was cut off twice, even at max_tokens={max_tokens}", ...)
if response.stop_reason == "refusal":
    return give_up("refusal", "Claude declined to continue", step, writes)

The continue skips appending the cut-off turn, so the model redoes it from the same point with twice the room. One retry is enough: if 2,048 tokens still can’t hold one turn, something else is wrong, and more tokens won’t fix it.

  [2] cut off at max_tokens=1024. Redoing the turn with 2048.
  [3] search_help {"query": "locked out SSO"}

Failing gracefully

What a run says when it stops matters as much as when. “Error: max_steps” in a log helps nobody. The person who picks the ticket up needs three things: why it stopped, what already changed, and that it’s now theirs.

def give_up(reason: str, why: str, step: int, writes: list[str]) -> RunResult:
    done = "; ".join(writes) if writes else "none, nothing was sent to the customer or changed"
    message = f"I couldn't finish this: {why}.\nChanges made: {done}.\nA person needs to pick this up."
    return RunResult("failed", reason, message, step, writes)

The loop records every successful write as it goes. So from python main.py deadline:

  FAILED (deadline) after 3 step(s), 0 email(s) sent
  I couldn't finish this: the run hit its 0.8s time limit.
  Changes made: none, nothing was sent to the customer or changed.
  A person needs to pick this up.

“Changes made” is the line that matters most. A support lead who reads “a reply was sent: rpl_001” won’t send another. One who reads “nothing was sent” knows the customer is still waiting. The short reason (deadline, stuck, …) is for dashboards; message is for people.

The same goes for the model call itself. If Claude’s API is still failing after the SDK’s own retries, the loop catches anthropic.APIError and returns the same kind of note, with reason model_unavailable, instead of a stack trace.

Try it yourself

The companion example is the loop from this article, a fake Kitebase with switchable faults, and nine scenarios. It runs offline in about three seconds.

Download the runnable example (zip)

cd 05-retries-and-failures
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py              # all nine scenarios
python main.py no-keys      # just one

pip install pytest && pytest -q runs 13 offline tests, one or more per failure path. With ANTHROPIC_API_KEY set, python main.py slow --live runs a tool-fault scenario against real Claude.

Then try these:

  1. Set TOOL_ATTEMPTS = 1 in main.py and run python main.py flaky. The dropped connection now reaches the model as is_error. A real model would spend a turn on something one retry would have fixed.
  2. In retry_safe, return True for reply_to_customer even without a key, and run python main.py no-keys. Still two emails, but now your code sends the second one instead of the model.
  3. Set RUN_DEADLINE = 0.3 and run python main.py slow. One hung tool, tried twice, uses the whole budget before the model gets a second turn.

Common beginner mistakes

  • Retrying everything in code. A permanent error fails the same way every time, and each retry delays the model’s chance to fix the input. Retry transient failures only.
  • No timeout on a tool. The loop waits for each tool, so one hung call freezes the whole run, and the step limit never fires because no step finishes.
  • Retrying a write with no idempotency key. Or with a new random key per attempt, which is the same thing. Derive the key from the arguments.
  • Running a cut-off tool call. A max_tokens turn can end mid-argument. Redo the turn; never execute what came back.
  • A failure message with no “what changed”. Whoever takes over needs to know whether the customer already got a reply.

Questions you will face in production

“The run hit a limit halfway through. Can it pick up where it left off?” Not with an in-memory loop like this one. The message list lives in the process, so a crash or a deploy loses it. Saving each step so a run can resume is durable execution, in Durable Execution: Resuming Long-Running Agents. Resuming is also the strongest reason for idempotency keys: a resumed run replays the last step, and the key makes that replay safe.

“What if a tool is slow for everyone, not just this run?” Then every run pays the timeout, again and again. Track tool failure rates, and when one tool keeps failing, take it out of the tool list for a while, or tell the model up front that it’s down. The pattern is a circuit breaker: after N failures in a row, stop calling for a minute. Observability and Cost Control for Agents covers the tracing you need to see it happening.

Check your understanding

search_tickets returns a 503 on the first try. Should your code retry it, or return it to the model?

Retry it in code, once, after a short wait. A 503 is transient and search_tickets only reads, so repeating it is harmless. If the retry fails too, return an is_error result saying the service seems down, so the model can take another route or stop.

assign_ticket times out. Is it safe for your code to retry?

Yes. It sets the assignee rather than toggling or adding, so running it twice with priya leaves the ticket assigned to priya. That’s why retry_safe allows it without a key. A tool like add_comment wouldn’t be safe: two retries, two comments.

Your loop never retries reply_to_customer, but customers still sometimes get two emails. How?

The model retries it. After a timeout, it gets is_error and may call the tool again with the same message. Not retrying in your code only removes one of the two retriers. Add an idempotency key derived from the ticket and the message, so a repeat from either one is recognized.

A turn comes back with stop_reason "max_tokens" and a tool_use block at the end. What do you do with that block?

Nothing. Its input may be cut off mid-argument. Don’t run it and don’t append the turn to the messages. Redo the same turn with a higher max_tokens, once, and give up with a clear message if it’s cut off again.

What to remember

  • A failed tool goes one of three ways: a quick retry in code (transient and safe to repeat), an is_error result the model can act on, or the end of the run.
  • Every tool gets a timeout. A timeout stops the waiting, not the work, so a timed-out write may have happened.
  • Retry writes only with an idempotency key derived from the arguments. The model retries too, so the key has to protect you from both.
  • A repeat guard stops identical failing calls; a step limit caps cost; a deadline caps the wait. You need all three.
  • Check stop_reason every turn: redo a max_tokens turn once with more room, and stop on a refusal.
  • When a run gives up, say why, say what already changed, and hand it to a person.

What to study next

Every guard here assumes the model is trying to do the right thing and failing. Next is what to do when it’s about to do the wrong thing successfully: Guardrails and Human in the Loop adds approval gates in front of write tools like reply_to_customer. After that, Durable Execution makes a run survive a crash and resume, which is where idempotency keys pay off most.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.