Your First LLM Integration: APIs, Errors, Retries

Kitebase, a small project-tracking app, wants its support inbox sorted by AI. The first version takes ten minutes: send each email to a model, ask for a category, print the answer. Then it meets real traffic. A Monday-morning burst gets rejected with 429 Too Many Requests. One reply takes long enough that the customer gives up on the spinner. And the “JSON” your code parses sometimes arrives wrapped in a Markdown code fence, so json.loads crashes.

None of that is the model being bad. It’s what happens when you call any slow, busy, paid HTTP API. Article 01: What LLMs Are gave you the mental model; this article is the code around the call.

What you’ll build: a triage helper for Kitebase’s inbox. It sends one customer email to Claude, gets back a checked JSON ticket, streams a first reply, and survives rate limits on the way. It runs offline against a scripted fake API that fails twice before it answers, so you can watch the retries without a key.

The email that runs through the whole article:

Hi, I've been locked out of Kitebase since this morning. The reset link never
arrives and my whole team is waiting on me. We have a client demo at 3pm. Please help!

The smallest call

Install the SDK (pip install anthropic) and get an API key from the Anthropic Console: a secret string that identifies you and your bill. The SDK reads it from the ANTHROPIC_API_KEY environment variable, so it never goes in your code.

import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    system="You are the support assistant for Kitebase... Three sentences at most.",
    messages=[{"role": "user", "content": EMAIL}],
)

Three parameters are required. model picks the model. max_tokens caps the length of the reply in tokens, the chunks of text models read and bill in (about 4 characters of English each, as article 01 showed). messages is the conversation so far; here it’s one user message holding the email. system is optional, and the next section covers it.

What comes back isn’t a string:

1. YOUR REQUEST model: "claude-opus-5" max_tokens: 1024 system: "You are the support assistant for Kitebase..." messages: [{role: "user", content: "Hi, I've been locked out of Kitebase..."}] 2. CLAUDE writes the reply token by token 3. THE RESPONSE content: [{ type: "text", text: "Sorry you're locked out, especially..."}] stop_reason: "end_turn" usage: {input_tokens: 93, output_tokens: 60} system holds the rules. messages carry this request's input. No temperature on claude-opus-5. A list of blocks. Join the ones whose type is "text". Check stop_reason first: max_tokens means cut off, refusal means declined. usage is the bill: $0.0020.
One call, and the three parts of the response you actually read.

Read three things, in this order (this is draft_reply in the example):

if response.stop_reason == "refusal":
    ...  # no reply to show: pass the email to a person
# content is a list of blocks, and the first one isn't always text.
text = "".join(block.text for block in response.content if block.type == "text")
if response.stop_reason in CUT_OFF:  # "max_tokens" and friends
    text += " [cut off]"
  • stop_reason says why the model stopped. "end_turn" means it finished. "max_tokens" means it hit your cap mid-sentence. "refusal" means it declined, and the content may be empty. All three come back as a successful response, so no exception warns you.
  • content is a list of blocks, each with a type. You’ll see response.content[0].text in a lot of tutorials. It breaks the day the first block is something else, like a tool call. Join the text blocks instead.
  • usage is what you pay for. The reply above used 93 input tokens and 60 output tokens. At claude-opus-5’s $5 per million input tokens and $25 per million output tokens, that’s $0.000465 + $0.0015 = $0.0020. Output tokens cost five times as much as input, so a chatty reply costs more than a long prompt.

(The token counts here come from the offline run, which estimates 4 characters per token. Real counts will differ a little.)

Notice there’s no temperature. Old tutorials pass it on every call, but claude-opus-5 rejects it with a 400 error. Article 01 explains what it is; on current flagship models you steer the output with the prompt instead.

The system prompt

The email is the input. The rules for handling it are the same for every email, so they go in the system prompt: standing instructions the model follows for the whole request, separate from the user’s text.

You are the support assistant for Kitebase, a project-tracking app.
Write a short, friendly first reply to the customer's email. Three sentences at most.
Don't promise anything the support team hasn't confirmed.

Put the role, the format and the hard rules here, and only the changing input in messages. Keeping them apart makes it clear which text is yours and which came from a customer. A customer who writes “ignore your instructions” is still a problem, though: that’s prompt injection, from article 01. Article 03 shows how to keep prompts like this in versioned templates.

The API is stateless: it remembers nothing between calls. For a back-and-forth chat, you send the whole conversation every time, alternating {"role": "user", ...} and {"role": "assistant", ...} messages.

The same call with OpenAI

Other providers look much the same. OpenAI’s version:

from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "system", "content": REPLY_PROMPT},
        {"role": "user", "content": EMAIL},
    ],
)
print(response.choices[0].message.content)

The system prompt is the first message instead of a separate parameter, the text is a plain string at choices[0].message.content, and the stop reason is choices[0].finish_reason ("stop", "length", …). Everything else here (streaming, timeouts, which errors to retry) carries over. Both official Python SDKs even share the same defaults: two automatic retries and a 10-minute timeout.

Streaming

A 60-token reply isn’t instant. With messages.create, your code gets nothing until the model has written the last token, and the customer watches a spinner the whole time.

Streaming sends the reply in pieces as the model writes it. The total time is the same, but the first words show up almost at once. It’s like reading a file line by line instead of waiting for read() to return the whole thing.

def stream_reply(client, email: str, on_text=lambda t: print(t, end="", flush=True)):
    with client.messages.stream(
        model=MODEL,
        max_tokens=1024,
        system=REPLY_PROMPT,
        messages=[{"role": "user", "content": email}],
    ) as stream:
        for text in stream.text_stream:
            on_text(text)
        final = stream.get_final_message()
    if final.stop_reason in CUT_OFF:
        on_text(" [cut off]")
    return final

text_stream yields only the text pieces. get_final_message() gives you the same response object as create, with stop_reason and usage, once the stream ends. Under the hood the API sends a series of events:

BROWSER YOUR SERVER CLAUDE API POST /reply (the email) messages.stream(...) message_start content_block_delta "Sorry you're" "Sorry you're" content_block_delta " locked out," " locked out," ... more deltas, a few tokens each ... message_delta stop_reason: "end_turn", output_tokens: 60 message_stop The reader sees text after the first delta. WITHOUT STREAMING Same work, same total time, but messages.create returns one JSON when the model is done. The browser shows a spinner until all 60 tokens are done, then the whole reply at once.
Streaming changes when the reader sees text, not how long the model takes.

In a web app, your server passes each piece on to the browser, usually as Server-Sent Events (SSE): one long HTTP response the server keeps writing lines to. Stream whenever a person is waiting for more than a sentence. Skip it for background jobs and for short JSON like the triage ticket, which your code can’t use until it’s complete.

There’s a second reason to stream long replies. The SDK’s timeout (covered below) is how long it waits for the next bytes, not for the whole answer. A streamed reply that keeps arriving never hits it. A non-streamed one that takes 40 seconds does.

What breaks, and what’s worth retrying

The SDK turns every failure into an exception. For each one, ask: would the same request work if you sent it again in a second?

The triage call, offline: two failures, then the ticket ATTEMPT 1 429 RateLimitError retry-after: 1 too many requests wait 1.0s server's hint ATTEMPT 2 529 OverloadedError no retry-after API is busy wait 0.9s 1.0s minus jitter ATTEMPT 3 200 OK {"category": "login", "urgent": true, ...} at 0s at 1.0s at 1.9s With no hint, the wait doubles, then stops growing at 8s fail 1 0.5s fail 2 1s fail 3 2s fail 4 4s fail 5+ 8s (the cap) Each wait loses up to 25% at random, so clients that failed together don't all retry together. RETRY: MIGHT WORK NEXT TIME 429 rate limited 500, 503 server error 529 overloaded timeout no reply in time, or network error DON'T RETRY: FAILS THE SAME WAY 400 bad request, or prompt too long 401 missing or wrong API key 403 key isn't allowed to do this 404 no such model, e.g. a typo refusal and max_tokens aren't errors at all: they arrive as 200 OK. Only stop_reason tells you.
The retry loop on the worked example, the backoff schedule, and which errors get a retry.
  • 429 rate limited. Your account has a cap on requests and tokens per minute, and you went over it. Waiting fixes it.
  • 500, 503, 529. Trouble on the provider’s side. 529 means the API is overloaded, which happens at busy times. Waiting usually fixes it.
  • Timeouts and connection errors. The network dropped, or no reply came in time. Worth another try.
  • 400, 401, 403, 404. Something is wrong with the request itself: a bad parameter (like temperature on claude-opus-5), a prompt longer than the model’s context window (the most text it can take in one request), a wrong key, a misspelled model name. Sending it again fails the same way, so fail fast and log it.

In code, that’s one small function:

def should_retry(error: Exception) -> bool:
    if isinstance(error, anthropic.APIConnectionError):  # network trouble, including timeouts
        return True
    if isinstance(error, anthropic.APIStatusError):
        return error.status_code == 429 or error.status_code >= 500
    return False

APITimeoutError is a subclass of APIConnectionError, and RateLimitError, OverloadedError and BadRequestError are subclasses of APIStatusError, so two checks cover them all.

Timeouts

The SDK’s default timeout is 10 minutes. That’s sensible for a long batch job and far too long for a person waiting on a support widget. Set your own when you create the client:

client = anthropic.Anthropic(timeout=30.0, max_retries=0)

Thirty seconds is a good default for anything interactive, as long as you stream anything longer than a sentence. For batch jobs, 60 to 120 seconds. You can also override it for one call with client.with_options(timeout=10.0).messages.create(...). The max_retries=0 is explained in the next section.

Retries with backoff

A retry that fires immediately usually fails again: a rate limit won’t have reset, and an overloaded API won’t have recovered. So you wait between attempts, and wait longer each time. That’s exponential backoff. Start at 0.5 seconds and double the wait after every failure, up to a cap:

def backoff_delay(attempt: int, rand=random.random) -> float:
    delay = min(BASE_DELAY * 2 ** attempt, MAX_DELAY)  # 0.5, 1, 2, 4, 8, 8...
    return delay * (1 - 0.25 * rand())

def with_retries(call, *, max_attempts=MAX_ATTEMPTS, sleep=time.sleep, log=print):
    for attempt in range(max_attempts):
        try:
            return call()
        except Exception as error:
            if not should_retry(error) or attempt == max_attempts - 1:
                raise
            hint = retry_after(error)
            delay = hint if hint is not None else backoff_delay(attempt)
            source = "retry-after header" if hint is not None else "backoff"
            log(f"  attempt {attempt + 1} failed: {describe(error)}. "
                f"Waiting {delay:.1f}s ({source}).")
            sleep(delay)

The rand() term is jitter: it shaves up to 25% off each wait at random. Without it, a hundred clients that got a 429 in the same second would all retry in the same second, and get rate limited together again. retry_after reads the server’s retry-after header when there is one, since the server knows better than your formula how long to wait. (sleep and log are parameters so the tests can check the delays without waiting.)

Offline, the fake API fails twice and then answers:

1. Triage (structured output, with retries)
  attempt 1 failed: 429 RateLimitError. Waiting 1.0s (retry-after header).
  attempt 2 failed: 529 OverloadedError. Waiting 0.9s (backoff).
  {"category":"login","urgent":true,"summary":"Customer is locked out, the reset email never arrives, and their team has a client demo at 3pm."}
  86 tokens in, 36 out, $0.0013

Here’s the gotcha: the SDK already does this. Out of the box it retries connection errors, 408, 409, 429 and 5xx twice, with backoff starting at 0.5 seconds, doubling up to 8, with jitter, honouring retry-after. The loop above is the same idea written out so you can see it. If you add your own loop and leave the SDK’s on, every one of your 4 attempts makes up to 3 requests, and one bad minute turns a single call into 12 requests. So pick one:

  • Default: the SDK’s retries. anthropic.Anthropic(max_retries=4, timeout=30.0) and no loop of your own.
  • Your own loop, with max_retries=0, when you want to log every attempt or do something the SDK can’t, like fall back to another model after the last failure. The example does this so you can watch it.
Is it safe to send the same request twice?

An operation is idempotent if doing it twice has the same effect as doing it once. Setting a user’s name to “Sam” is. Charging a card isn’t.

An LLM call is safe to repeat in the sense that nothing changes on the provider’s side, but it isn’t free. If a request times out on your end, the model may still have finished it, and you pay for both runs. That’s fine for triage. The step to protect is the one after the call: if your code sends the reply email, key that on the email’s ID so a retry can’t send it twice.

JSON you can trust

The triage result goes into a support queue, so your code needs fields, not prose. The obvious first attempt is to ask for JSON in the prompt and call json.loads. It works in testing. Then one day the model wraps its answer in a Markdown fence:

```json
{"category": "login", "urgent": true, "summary": "..."}
```
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)

Or it adds “Here’s the ticket:” first, or invents a category called "account access" that your queue doesn’t have. You can write cleanup code for each case, but you’ll never finish.

Structured output fixes this at the source. You describe the shape as a Pydantic model (a Python class that declares fields and their types, and validates data against them), and the API makes the model’s output match it as it writes:

class Ticket(BaseModel):
    category: Literal["login", "billing", "bug", "feature_request", "other"]
    urgent: bool
    summary: str

try:
    response = client.messages.parse(
        model=MODEL,
        max_tokens=1024,
        system=TRIAGE_PROMPT,
        messages=[{"role": "user", "content": email}],
        output_format=Ticket,
    )
except ValidationError as error:
    raise TriageFailed(f"JSON didn't match Ticket, probably cut off: {error}") from error
if response.stop_reason == "refusal":
    raise TriageFailed("Claude declined to triage this email")
ticket = response.parsed_output  # a Ticket, already validated

Literal[...] limits category to five values, so "account access" can’t come back. For the Kitebase email, the ticket is login and urgent, as the retry output above showed.

Two things can still go wrong, and the try and the if catch them. If the JSON is cut off at max_tokens, half an object isn’t valid JSON, so messages.parse raises Pydantic’s ValidationError. The same limit will fail the same way, so give it room: the ticket is about 40 tokens, and 1,024 leaves space for any thinking the model does first, which counts toward max_tokens too. And a refusal has no ticket to read. TriageFailed isn’t in should_retry, so with_retries passes it straight through: a 429 gets another try, a broken ticket goes to a person. Article 03 goes further with structured outputs, few-shot examples and prompt templates.

Try it yourself

The companion example is the whole triage helper: main.py has every function above, and fake_claude.py is the scripted API that fails on cue. With no key it runs offline. With ANTHROPIC_API_KEY set, it calls claude-opus-5 for real and prints the actual token counts and cost.

Download the runnable example (zip)

cd 02-first-llm-integration
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. python main.py --no-stream. The reply appears all at once after a pause, instead of word by word. Same text, same total time.
  2. In main(), change failures=["429", "529"] to ["400"]. The run fails on the first attempt with BadRequestError, with no retries. Then try ["529"] * 5 and watch it give up after 4 attempts.
  3. In triage, set max_tokens=20 and run it with a real key. The ticket JSON gets cut off and you get Triage failed: JSON didn't match Ticket, probably cut off, not a crash three functions later.

pip install pytest && pytest -q runs the offline tests. They need no key and no network.

Common beginner mistakes

  • Calling the API from the browser. Anyone can open dev tools and copy your key. Call it from your backend and send the browser only the text.
  • Ignoring stop_reason. A cut-off reply and a refusal both arrive as success. Check it before you show or parse anything.
  • Retrying everything, or retrying twice. Retrying a 400 just fails slower. Stacking your loop on the SDK’s multiplies requests. Retry 429, 5xx and timeouts, in one place.
  • Logging whole prompts and replies. Support emails contain names, addresses and account details. Log tokens, latency and IDs; redact or skip the text.

Questions you will face in production

“What timeout should we use?” Thirty seconds for anything a person is waiting on, with streaming so long replies don’t hit it. Sixty to 120 seconds for batch jobs. Never the 10-minute default for a user-facing request.

“What happens when the provider is down?” Retries cover a bad minute, not a bad hour. For a feature that must keep working, fall back after the last retry: to another model, another provider, or a plain “we’re on it” reply. A circuit breaker helps too: after, say, 5 failures in a row, stop calling for a minute and go straight to the fallback. A second provider means a second prompt to test, so add one only when the feature earns it.

“How do we stop a runaway bill?” Cap max_tokens on every call. Check input size before you send: client.messages.count_tokens(...) returns the count without running the model. Log usage on every call to see cost per feature. And set a spend limit in the provider’s console as the last line of defence.

“How do we keep the key safe?” In an environment variable, from a secrets manager in production and a .env file (listed in .gitignore) locally. Never log it, and rotate it when someone with access leaves.

Check your understanding

Triage works in testing, but in production about one email in a few hundred fails with TriageFailed: "probably cut off". What's going on, and what do you change?

Those are probably unusually long emails, where the model’s summary runs past max_tokens and the JSON is cut in half. Retrying won’t help. Raise max_tokens, or tell the model in the system prompt to keep summary to one short sentence, then check the failures stop.

A teammate sets max_retries=5 on the client and wraps every call in with_retries (4 attempts). During an outage, how many requests can one triage make?

Each of the 4 attempts makes up to 6 requests (the first try plus 5 SDK retries), so up to 24, with waits stacking up too. Keep retries in one place: either the SDK’s, or your loop with max_retries=0.

A request fails with a 400 saying the prompt is too long. Should with_retries retry it?

No. The same prompt will be too long the next time too, and should_retry correctly returns False for 400s. Shrink the input: drop old conversation turns, or send less context.

You move some code from an old tutorial onto claude-opus-5 and every call fails instantly, with no retries. What's the likely cause?

The old code passes temperature (or a prefilled assistant message), which claude-opus-5 rejects with a 400. No retries happen because a 400 isn’t worth retrying. Remove the parameter; use the prompt, or structured output, to control the result.

What to remember

  • A response is blocks plus stop_reason plus usage. Check stop_reason, join the text blocks, log the tokens.
  • Rules go in the system prompt, the changing input in messages. The API remembers nothing between calls.
  • Stream when a person is waiting. The total time is the same, but the reader sees text almost at once.
  • Retry 429, 5xx and timeouts with exponential backoff and jitter. Fail fast on other 4xx. Retry in one place only.
  • Set a 30-second timeout for interactive calls; the 10-minute default is for batch jobs.
  • For JSON, use structured output (messages.parse), and still handle a cut-off or a refusal.

What to study next

You can call the API and keep the call alive. Next is making the prompt itself dependable: article 03: Prompting as Code covers templates, structured outputs in more depth, few-shot examples and prompt versioning. When you get to logging every call in production, article 11: Production AI Observability picks up from the usage line here.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.