Plan, Act, Observe: The Agent Loop in Code

What an AI Agent Actually Is showed the agent loop in about 20 lines: ask the model, run the tool it picks, feed the result back, repeat. Point that sketch at a real API and it breaks in specific ways. Forget one line and the second call comes back as a 400 error. The model asks for two tools at once and you answer one. A tool throws and the whole run dies. A confused model keeps calling tools, and every turn is another paid request.

None of these are hard to fix once you can see the exact messages going back and forth. This article writes the loop by hand with the Claude SDK and prints every message, so you can.

What you’ll build: a Kitebase support agent that works ticket KITE-142 (“Customer locked out after SSO change”) with five tools, in one hand-written loop on the Claude SDK. It runs offline against a scripted stand-in for Claude and prints every turn.

The worked example: one ticket, five tools

Kitebase is the made-up project-tracking app used across this site. A support lead gives the agent one instruction:

Handle ticket KITE-142, and tell me if other customers have hit the same problem.

The agent can use five tools, functions your code runs when the model asks for them:

ToolWhat it does
get_ticket(ticket_id)Title, status, assignee and customer of one ticket
search_help(query)The 2 best-matching sections of the help center
search_tickets(query, status)Tickets whose title or customer contains every word
assign_ticket(ticket_id, assignee)Assign a ticket to a support engineer
reply_to_customer(ticket_id, message)Send the customer a reply

The model never runs any of these. It only sees a tool definition for each one: a name, a description it chooses tools by, and a JSON Schema describing the input. Here’s one of the five, from tools.py:

{"name": "get_ticket",
 "description": "Look up one support ticket by id: title, status, assignee and customer.",
 "input_schema": {"type": "object", "required": ["ticket_id"],
                  "properties": {"ticket_id": {"type": "string", "description": "e.g. KITE-142"}}}}

Writing definitions the model picks correctly is its own skill, covered in Giving an Agent Tools. This article is about the loop that runs them. Here it is, with the values from the worked example:

for turn in range(1, MAX_TURNS + 1): each pass is one turn and one model call MESSAGES [user: the goal] grows each turn 1 → 3 → 5 → 7 → 8 1. PLAN client.messages.create( system, tools=TOOLS, messages=messages) then append it as the assistant turn CHECK stop_reason "tool_use": go on anything else: stop stop ANSWER return the final text turn 4, "end_turn" "tool_use" 2. ACT for every tool_use block, your code runs run_tool(block.name, block.input) turn 2: search_help and search_tickets a failure becomes text with is_error: true 3. OBSERVE append ONE user message holding a tool_result for every tool_use id tool_result toolu_02, tool_result toolu_03 next turn TURN LIMIT 10 turns and still "tool_use": stop and return "Stopped: hit the 10-turn limit". The model can't count its own turns. Your code can, and every turn is a paid call.
One pass is one turn and one paid model call. The rest of this article walks round it once.

The loop has three steps and one guard. Each section below adds one of them.

Plan: send the tools and the whole conversation

A turn is one pass round the loop, and it starts with a normal Messages API call, the same one from Your First LLM Integration, with one new parameter:

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=4096,
    system=SYSTEM_PROMPT,
    tools=TOOLS,          # the five definitions, sent on every call
    messages=messages,    # the whole conversation so far
)

messages starts as a list with one entry, the goal. With tools in the request, the model can answer with text or by asking you to run a tool. The response’s content is a list of content blocks, and here’s what comes back on turn 1:

{
  "stop_reason": "tool_use",
  "content": [
    {"type": "text", "text": "I'll start by reading the ticket."},
    {"type": "tool_use", "id": "toolu_01", "name": "get_ticket",
     "input": {"ticket_id": "KITE-142"}}
  ]
}

The tool_use block is the model’s request: run get_ticket with this input. Its id is how you’ll say which request a result answers. Real ids are long random strings like toolu_01A09q90qw90lq917835lq9; the example uses short ones so you can follow them. input is already a Python dict, so there’s nothing to parse.

The text block before it is the model reasoning about its next step, a pattern called ReAct (reason, then act) after the paper that named it. Models do it on their own. When an agent does something odd, those lines tell you what it thought it was doing.

Read stop_reason, then keep the assistant turn

stop_reason says why the model stopped writing, and in an agent loop it decides what happens next:

  • "tool_use": it wants tools run. Keep looping.
  • "end_turn": it’s finished. The text blocks are the answer.
  • "max_tokens": it hit your max_tokens cap mid-reply. Treat it as a failure, not an answer.
  • "refusal": it declined, and content may be empty.

So the loop’s exit test is one line: if stop_reason isn’t "tool_use", you’re done. Don’t test “are there tool_use blocks?” instead. A reply cut off by max_tokens can end halfway through one, with an input that was never finished.

Before that test, append the reply to messages exactly as it came back:

messages.append({"role": "assistant", "content": response.content})

if response.stop_reason != "tool_use":
    answer = "".join(b.text for b in response.content if b.type == "text")
    ...  # flag it if stop_reason isn't "end_turn", then return

This is the line the 20-line sketch made easy to forget. Your tool results will point at toolu_01. If the assistant message holding toolu_01 isn’t in the list you send next, the API can’t tell what they answer and rejects the request with a 400. Append the whole response.content, text blocks included; don’t rebuild it or keep only the tool calls. Appending before the exit test also keeps the final answer in the list, so a follow-up question can continue the same conversation.

Act: run every tool the model asked for

On turn 2 the model has read the ticket and asks for two things at once:

"They lost access after an SSO change. I'll check the help center and look for other SSO tickets at the same time."
tool_use  toolu_02  search_help(query="locked out after SSO change")
tool_use  toolu_03  search_tickets(query="sso", status="any")

These are parallel tool calls: several tool_use blocks in one response, because the model saw two lookups that don’t depend on each other. It happens often, so “act” is a loop over every block, never response.content[0]:

results = []
for block in response.content:
    if block.type == "tool_use":
        output, is_error = run_tool(kb, block.name, block.input)
        result = {"type": "tool_result", "tool_use_id": block.id, "content": output}
        if is_error:
            result["is_error"] = True
        results.append(result)

kb is the in-memory support desk the tools act on. Running the calls one after another is fine here.

run_tool is where tool failures stop being crashes. Tools fail all the time: an id is mistyped, an API times out, the model invents a tool name. None of that should end the run.

def run_tool(kb: Kitebase, name: str, args: dict) -> tuple[str, bool]:
    if name not in TOOL_NAMES:
        return f"Error: there's no tool called {name!r}.", True
    try:
        return getattr(kb, name)(**args), False
    except Exception as e:  # a bad id, a missing argument, a timeout from a real API
        return f"Error: {e}", True

If the model had asked for get_ticket("KITE-999"), the result would be "Error: No ticket 'KITE-999'. Ids look like KITE-142." with is_error: true. The model reads that like any other result and decides what to do: search for the ticket, or tell the lead it couldn’t find it.

Why send errors to the model instead of retrying in code?

Retry in code when the error is a blip. A timeout or a 503 from the ticket API usually works on the second try, and the model can’t do anything useful with it.

Most tool errors aren’t blips, though; they tell the model to do something different. “No ticket KITE-999” means the id is wrong, so search for it. “Missing argument: ticket_id” means the call was malformed, so fix it. Retrying the same call in code gets the same error. The model is the only part of the system that can change its mind, so send those errors back. Retries, Timeouts, and Failure Handling covers where to draw the line.

Observe: every result goes back in one user message

A tool_result block is your answer to one tool_use block. It has three fields: tool_use_id (the id you’re answering), content (what the tool returned, as text) and optionally is_error. Results go back with role user: they come from your side of the conversation, even though no person typed them.

All the results for a turn go in one user message, one block per tool_use id:

messages.append({"role": "user", "content": results})

After turn 2 that message holds two blocks: toolu_02 with the two help-center sections, starting with [account-recovery#2] Locked out of your account: If your workspace uses single sign-on, and toolu_03 with three tickets, KITE-139, KITE-142 and KITE-144. The pairing rules are where most hand-written loops go wrong:

RIGHT: WHAT CALL 3 ENDS WITH [3] ASSISTANT appended exactly as it came back tool_use toolu_02 search_help tool_use toolu_03 search_tickets [4] USER one message, one result per id tool_result toolu_02 tool_result toolu_03 Every tool_use id is answered in the very next message. Call 3 goes through. FORGOT TO APPEND THE ASSISTANT TURN The list goes [2] user, then [3] user with tool_result toolu_02, tool_result toolu_03 400: those ids answer a question the API never saw. The fake client wouldn't notice. The tests would. ANSWERED ONLY ONE OF TWO The loop reads only the first tool_use block: [4] user: tool_result toolu_02 400: toolu_03 has no tool_result after it. Loop over every block, not content[0]. LET THE TOOL RAISE get_ticket("KITE-999") raises ValueError No result at all: the whole run crashes. Send it back: is_error: true, "No ticket 'KITE-999'..." The model reads it and tries search_tickets.
Every tool_use id gets exactly one tool_result, in the very next message.

If you add text of your own to that message, the tool_result blocks go first. And don’t split one turn’s results over several user messages; one message per turn is the documented shape for parallel calls.

The messages list is the agent’s state

Run the example offline and it prints each turn: how much it sent, what came back, and what it appended. Trimmed a little:

Turn 1: sent 1 message, about 452 tokens. stop_reason = 'tool_use'
  + assistant text         "I'll start by reading the ticket."
              tool_use     get_ticket(ticket_id="KITE-142")  id=toolu_01
  + user      tool_result  id=toolu_01  {"id": "KITE-142", "title": "Customer loc...
  messages now: 3

Turn 2: sent 3 messages, about 567 tokens. stop_reason = 'tool_use'
  + assistant text         "They lost access after an SSO change. I'll check the he..."
              tool_use     search_help(query="locked out after S...)  id=toolu_02
              tool_use     search_tickets(query="sso", status="any")  id=toolu_03
  + user      tool_result  id=toolu_02  [account-recovery#2] Locked out of your a...
              tool_result  id=toolu_03  [{"id": "KITE-139", "title": "SSO login l...
  messages now: 5

Turn 3: sent 5 messages, about 1,022 tokens. stop_reason = 'tool_use'
  + assistant tool_use     reply_to_customer(ticket_id="KITE-142", message="Hi, since your wor...)  id=toolu_04
  + user      tool_result  id=toolu_04  Reply sent to Northwind Studio on KITE-142.
  messages now: 7

Turn 4: sent 7 messages, about 1,150 tokens. stop_reason = 'end_turn'
  + assistant text         "Replied to Northwind Studio on KITE-142: after the SSO ..."
  messages now: 8

Final answer:
Replied to Northwind Studio on KITE-142: after the SSO change they sign in through their
identity provider, and their IT admin resets passwords [account-recovery#2]. Related:
KITE-139 (closed) was an SSO sign-in loop for the same customer, and KITE-144 (open,
unassigned) looks like the same problem at Harbor Pine.

Look at the first number on each turn: 1, 3, 5, 7. The model is stateless: it remembers nothing between calls. The only reason it knows on turn 3 what the help center said on turn 2 is that you sent turn 2 again. So the loop keeps no memory anywhere except this one list:

The messages list after the worked example Roles alternate. Each tool_result carries its tool_use id. call 1 ~452 call 2 ~567 call 3 ~1,022 call 4 ~1,150 [0] USER "Handle ticket KITE-142, and tell me if other customers..." [1] ASSISTANT "I'll start by reading the ticket." tool_use toolu_01 get_ticket [2] USER tool_result toolu_01 KITE-142, priya, Northwind [3] ASSISTANT "They lost access after an SSO change. I'll check..." tool_use toolu_02 search_help tool_use toolu_03 search_tickets [4] USER tool_result toolu_02 [account-recovery#2]... tool_result toolu_03 KITE-139, 142, 144 [5] ASSISTANT tool_use toolu_04 reply_to_customer [6] USER tool_result toolu_04 Reply sent to Northwind [7] ASSISTANT "Replied to Northwind Studio on KITE-142: ..." Bars: the messages each call sent, and about how many tokens that was, counting the system prompt and the five tool definitions that ride along every time. Arrows: the reply each call added. The last call sent about 1,150 tokens, but the four together sent about 3,190.
The list after the run. Each call sends everything above the reply it produces.

A lot follows from that:

  • Cost grows faster than the conversation. The token counts (estimated at 4 characters per token) include the system prompt and the tool definitions, which ride along on every call. The last call sent about 1,150 tokens, but the four together sent about 3,190. At claude-opus-5’s $5 per million input tokens that’s under 2 cents; a 30-turn run with big tool results is a different bill. The real count is in response.usage.input_tokens.
  • Saving an agent means saving the list. Store it as JSON after each turn; resuming is loading it and calling the model again. python main.py --messages prints exactly what you’d store.
  • Debugging means reading the list. When the agent does something odd on turn 5, the messages sent on call 5 are everything it knew. There’s nowhere else to look.
  • It only grows. Past a point the model tracks early turns less well, and eventually the list overflows the context window, the most tokens a model can take in one call. Trimming it is the other half of State and Memory.

The turn limit

Now suppose the model never says it’s done. It searches, doesn’t like the result, searches again with other words, and keeps going. Every turn is a paid call that re-sends the whole growing list.

So the outer loop isn’t while True. It counts:

MAX_TURNS = 10

for turn in range(1, max_turns + 1):
    ...  # plan, check, act, observe

return f"Stopped: hit the {max_turns}-turn limit before the agent finished.", messages

python main.py --max-turns 2 shows it: the run stops after the two lookups, before any reply is sent, and the answer is the “Stopped” message instead of a summary.

Start at 10 for a task like this one, which takes 4, and raise it when real tasks need more. Return an explicit error, never an empty string, so “the agent gave up” can’t be mistaken for “the agent had nothing to say”.

Why can't the model just stop itself?

It doesn’t know how long it’s been going. Each call sees the conversation, not a turn counter, and a model that’s stuck doesn’t feel stuck. It sees one more search that might help.

The limit lives in your code, outside the model, which is where a safety rail belongs. It’s a timeout for the loop: set it once and it holds however the model behaves. Observability and Cost Control for Agents adds a budget in dollars on top.

The whole loop

Put the pieces together and this is run_agent from main.py, minus the printing:

def run_agent(client, kb: Kitebase, goal: str, max_turns: int = MAX_TURNS):
    messages = [{"role": "user", "content": goal}]

    for turn in range(1, max_turns + 1):
        response = client.messages.create(model=MODEL, max_tokens=4096,
            system=SYSTEM_PROMPT, tools=TOOLS, messages=messages)
        messages.append({"role": "assistant", "content": response.content})

        if response.stop_reason != "tool_use":
            answer = "".join(b.text for b in response.content if b.type == "text")
            if response.stop_reason != "end_turn":
                answer += f"\n[stopped early: stop_reason={response.stop_reason}]"
            return answer, messages

        results = []
        for block in response.content:
            if block.type == "tool_use":
                output, is_error = run_tool(kb, block.name, block.input)
                result = {"type": "tool_result", "tool_use_id": block.id, "content": output}
                if is_error:
                    result["is_error"] = True
                results.append(result)
        messages.append({"role": "user", "content": results})

    return f"Stopped: hit the {max_turns}-turn limit before the agent finished.", messages

That’s the whole agent: one list, one API call per turn, and your code in between. It takes client as a parameter, so the offline example passes a fake and the real run passes anthropic.Anthropic().

The fake hands back the four canned responses you’ve seen, in order. It doesn’t read tool results, so it can’t test whether the agent is smart. It’s exactly right for testing the loop, and that’s what test_main.py does: messages alternate user and assistant, every tool_use id comes back as a tool_result in the next message, the loop stops on end_turn, and it gives up at the limit.

The shortcut: the SDK’s tool runner

Once you understand this loop, you don’t have to keep writing it. The Python SDK’s tool runner does it for you: decorate plain functions, and it builds the definitions from their type hints and docstrings, then runs the same loop:

from anthropic import Anthropic, beta_tool

@beta_tool
def get_ticket(ticket_id: str) -> str:
    """Look up one support ticket by id: title, status, assignee and customer.

    Args:
        ticket_id: e.g. KITE-142
    """
    return kb.get_ticket(ticket_id)

runner = Anthropic().beta.messages.tool_runner(
    model="claude-opus-5", max_tokens=4096, system=SYSTEM_PROMPT,
    tools=[get_ticket, ...],  # the other four, decorated the same way
    messages=[{"role": "user", "content": GOAL}],
    max_iterations=10,
)
final = runner.until_done()  # the last assistant message

It sends a failing tool back with is_error for you, and max_iterations is your turn limit. One gotcha: when it hits max_iterations it just stops and hands you the last message, so check final.stop_reason. If it’s still "tool_use", the agent didn’t finish.

Anthropic’s SDK docs recommend it as the default for your own tools. It’s beta, as the name says, so pin your SDK version. Either way, you now know what it does under the hood, and that’s what you’ll be reading when you debug it.

Try it yourself

The companion example is the Kitebase agent from this article. Offline, it runs the full loop against the scripted client and prints every turn. With ANTHROPIC_API_KEY set, the same loop runs against claude-opus-5 (add --scripted to force the offline run).

Download the runnable example (zip)

cd 02-agent-loop-in-code
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. python main.py --max-turns 2. The run stops after the two lookups and returns “Stopped: hit the 2-turn limit”. No reply goes to the customer.
  2. In main.py, delete the line messages.append({"role": "assistant", "content": response.content}) and run pytest -q. The scripted client doesn’t care, but the tests fail on exactly the rule the real API enforces with a 400.
  3. In scripted_client.py, change the first get_ticket call to ticket_id="KITE-999" and run it. The result comes back with is_error and the loop carries on instead of crashing. (The scripted model doesn’t read results, so it carries on with its script. A real model would read the error and search for the ticket.)

pip install pytest && pytest -q runs the offline tests. They need no key and no network.

Common beginner mistakes

  • Not appending the assistant turn. The results point at ids the API has never seen, and the request fails with a 400. Append response.content as is, before anything else.
  • Answering only the first tool call. Models batch calls. Loop over every tool_use block, or the extra ones go unanswered and the next request is rejected.
  • Letting a tool exception escape. One bad call kills the run. Catch it and send the error back with is_error: true.
  • while True. The model can’t be trusted to stop. Count turns and return an explicit “stopped” error at the limit.
  • Treating max_tokens as done. A reply cut off by max_tokens isn’t an answer. Check stop_reason before trusting anything.

Questions you will face in production

“Should I run a turn’s tool calls in parallel?” If they’re read-only and independent, like the two lookups on turn 2, yes: it saves time on slow APIs. Run them one after another when one writes something another reads. Either way, all the results go back in one message.

“What should the caller get when the turn limit hits?” An explicit failure it can act on, such as a “stopped” status that routes the ticket to a person. Log the full message list with it: it tells you whether the task was too hard, a tool kept failing, or the limit was too low.

“How long can the messages list get?” Watch tokens, not message count: one big tool result can outweigh twenty short turns. Keep tool outputs small (the help search returns 2 sections, not the whole article), and look at response.usage.input_tokens per turn. When it heads toward the context window, it’s time for the trimming in State and Memory.

Check your understanding

Your loop works for one tool call per turn. The first time the model asks for two, the next request fails with a 400. What's the likely bug?

The loop answered only one of them, probably by reading response.content[0] or returning after the first tool_use block. The API wants a tool_result for every tool_use id in the very next message. Loop over every block and put all the results in one user message.

A reply comes back with stop_reason "max_tokens" and ends partway through a tool_use block. What should the loop do?

Stop and treat it as a failure, not run the tool. The input may be incomplete, so running it could do the wrong thing. Testing stop_reason != "tool_use" rather than “are there tool_use blocks?” handles this for free. If it keeps happening on real tasks, raise max_tokens.

A teammate wants to save money by sending only the last two messages each turn. What breaks?

The model forgets everything else. On turn 4 it would see the reply being sent but not the goal, so it wouldn’t know it was also asked about related tickets. Shrinking the list is possible, but as a deliberate design: keep the goal, summarise old turns, and never separate a tool_use from its tool_result. State and Memory covers how.

Your agent hits the 10-turn limit on 5% of tickets. What do you look at first?

The saved message lists for those runs. Look for the same tool called with slightly different input over and over (a search that never finds anything), or an error the model keeps retrying. Usually the fix is a better tool or a clearer error message, not a higher limit.

What to remember

  • An agent is one loop: call the model with the tools and the whole conversation, run what it asks for, send the results back, repeat.
  • stop_reason drives the loop. "tool_use" means keep going; anything else means stop, and only "end_turn" means it finished properly.
  • Append every assistant turn exactly as it came back, then answer every tool_use id with a tool_result in one user message.
  • Tool failures go back to the model as is_error results, not up your stack as exceptions.
  • The messages list is the agent’s entire state, and you re-send all of it on every call.
  • Cap the turns in your code. The model can’t count them for you.

What to study next

The loop runs whatever tools you hand it, and it’s only as good as those tools. Giving an Agent Tools covers writing definitions the model picks correctly, keeping tool results small, and checking the arguments the model sends before anything with side effects runs, like reply_to_customer.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.