Plan, Act, Observe: The Agent Loop in Code
What an AI Agent Actually Is showed the agent loop in about 20 lines: ask the model, run the tool it picks, feed the result back, repeat. Point that sketch at a real API and it breaks in specific ways. Forget one line and the second call comes back as a 400 error. The model asks for two tools at once and you answer one. A tool throws and the whole run dies. A confused model keeps calling tools, and every turn is another paid request.
None of these are hard to fix once you can see the exact messages going back and forth. This article writes the loop by hand with the Claude SDK and prints every message, so you can.
What you’ll build: a Kitebase support agent that works ticket KITE-142 (“Customer locked out after SSO change”) with five tools, in one hand-written loop on the Claude SDK. It runs offline against a scripted stand-in for Claude and prints every turn.
The worked example: one ticket, five tools
Kitebase is the made-up project-tracking app used across this site. A support lead gives the agent one instruction:
Handle ticket KITE-142, and tell me if other customers have hit the same problem.
The agent can use five tools, functions your code runs when the model asks for them:
| Tool | What it does |
|---|---|
get_ticket(ticket_id) | Title, status, assignee and customer of one ticket |
search_help(query) | The 2 best-matching sections of the help center |
search_tickets(query, status) | Tickets whose title or customer contains every word |
assign_ticket(ticket_id, assignee) | Assign a ticket to a support engineer |
reply_to_customer(ticket_id, message) | Send the customer a reply |
The model never runs any of these. It only sees a tool definition for each one: a name, a description it chooses tools by, and a JSON Schema describing the input. Here’s one of the five, from tools.py:
{"name": "get_ticket",
"description": "Look up one support ticket by id: title, status, assignee and customer.",
"input_schema": {"type": "object", "required": ["ticket_id"],
"properties": {"ticket_id": {"type": "string", "description": "e.g. KITE-142"}}}}
Writing definitions the model picks correctly is its own skill, covered in Giving an Agent Tools. This article is about the loop that runs them. Here it is, with the values from the worked example:
The loop has three steps and one guard. Each section below adds one of them.
Plan: send the tools and the whole conversation
A turn is one pass round the loop, and it starts with a normal Messages API call, the same one from Your First LLM Integration, with one new parameter:
response = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
system=SYSTEM_PROMPT,
tools=TOOLS, # the five definitions, sent on every call
messages=messages, # the whole conversation so far
)
messages starts as a list with one entry, the goal. With tools in the request, the model can answer with text or by asking you to run a tool. The response’s content is a list of content blocks, and here’s what comes back on turn 1:
{
"stop_reason": "tool_use",
"content": [
{"type": "text", "text": "I'll start by reading the ticket."},
{"type": "tool_use", "id": "toolu_01", "name": "get_ticket",
"input": {"ticket_id": "KITE-142"}}
]
}
The tool_use block is the model’s request: run get_ticket with this input. Its id is how you’ll say which request a result answers. Real ids are long random strings like toolu_01A09q90qw90lq917835lq9; the example uses short ones so you can follow them. input is already a Python dict, so there’s nothing to parse.
The text block before it is the model reasoning about its next step, a pattern called ReAct (reason, then act) after the paper that named it. Models do it on their own. When an agent does something odd, those lines tell you what it thought it was doing.
Read stop_reason, then keep the assistant turn
stop_reason says why the model stopped writing, and in an agent loop it decides what happens next:
"tool_use": it wants tools run. Keep looping."end_turn": it’s finished. The text blocks are the answer."max_tokens": it hit yourmax_tokenscap mid-reply. Treat it as a failure, not an answer."refusal": it declined, andcontentmay be empty.
So the loop’s exit test is one line: if stop_reason isn’t "tool_use", you’re done. Don’t test “are there tool_use blocks?” instead. A reply cut off by max_tokens can end halfway through one, with an input that was never finished.
Before that test, append the reply to messages exactly as it came back:
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
answer = "".join(b.text for b in response.content if b.type == "text")
... # flag it if stop_reason isn't "end_turn", then return
This is the line the 20-line sketch made easy to forget. Your tool results will point at toolu_01. If the assistant message holding toolu_01 isn’t in the list you send next, the API can’t tell what they answer and rejects the request with a 400. Append the whole response.content, text blocks included; don’t rebuild it or keep only the tool calls. Appending before the exit test also keeps the final answer in the list, so a follow-up question can continue the same conversation.
Act: run every tool the model asked for
On turn 2 the model has read the ticket and asks for two things at once:
"They lost access after an SSO change. I'll check the help center and look for other SSO tickets at the same time."
tool_use toolu_02 search_help(query="locked out after SSO change")
tool_use toolu_03 search_tickets(query="sso", status="any")
These are parallel tool calls: several tool_use blocks in one response, because the model saw two lookups that don’t depend on each other. It happens often, so “act” is a loop over every block, never response.content[0]:
results = []
for block in response.content:
if block.type == "tool_use":
output, is_error = run_tool(kb, block.name, block.input)
result = {"type": "tool_result", "tool_use_id": block.id, "content": output}
if is_error:
result["is_error"] = True
results.append(result)
kb is the in-memory support desk the tools act on. Running the calls one after another is fine here.
run_tool is where tool failures stop being crashes. Tools fail all the time: an id is mistyped, an API times out, the model invents a tool name. None of that should end the run.
def run_tool(kb: Kitebase, name: str, args: dict) -> tuple[str, bool]:
if name not in TOOL_NAMES:
return f"Error: there's no tool called {name!r}.", True
try:
return getattr(kb, name)(**args), False
except Exception as e: # a bad id, a missing argument, a timeout from a real API
return f"Error: {e}", True
If the model had asked for get_ticket("KITE-999"), the result would be "Error: No ticket 'KITE-999'. Ids look like KITE-142." with is_error: true. The model reads that like any other result and decides what to do: search for the ticket, or tell the lead it couldn’t find it.
Why send errors to the model instead of retrying in code?
Retry in code when the error is a blip. A timeout or a 503 from the ticket API usually works on the second try, and the model can’t do anything useful with it.
Most tool errors aren’t blips, though; they tell the model to do something different. “No ticket KITE-999” means the id is wrong, so search for it. “Missing argument: ticket_id” means the call was malformed, so fix it. Retrying the same call in code gets the same error. The model is the only part of the system that can change its mind, so send those errors back. Retries, Timeouts, and Failure Handling covers where to draw the line.
Observe: every result goes back in one user message
A tool_result block is your answer to one tool_use block. It has three fields: tool_use_id (the id you’re answering), content (what the tool returned, as text) and optionally is_error. Results go back with role user: they come from your side of the conversation, even though no person typed them.
All the results for a turn go in one user message, one block per tool_use id:
messages.append({"role": "user", "content": results})
After turn 2 that message holds two blocks: toolu_02 with the two help-center sections, starting with [account-recovery#2] Locked out of your account: If your workspace uses single sign-on, and toolu_03 with three tickets, KITE-139, KITE-142 and KITE-144. The pairing rules are where most hand-written loops go wrong:
If you add text of your own to that message, the tool_result blocks go first. And don’t split one turn’s results over several user messages; one message per turn is the documented shape for parallel calls.
The messages list is the agent’s state
Run the example offline and it prints each turn: how much it sent, what came back, and what it appended. Trimmed a little:
Turn 1: sent 1 message, about 452 tokens. stop_reason = 'tool_use'
+ assistant text "I'll start by reading the ticket."
tool_use get_ticket(ticket_id="KITE-142") id=toolu_01
+ user tool_result id=toolu_01 {"id": "KITE-142", "title": "Customer loc...
messages now: 3
Turn 2: sent 3 messages, about 567 tokens. stop_reason = 'tool_use'
+ assistant text "They lost access after an SSO change. I'll check the he..."
tool_use search_help(query="locked out after S...) id=toolu_02
tool_use search_tickets(query="sso", status="any") id=toolu_03
+ user tool_result id=toolu_02 [account-recovery#2] Locked out of your a...
tool_result id=toolu_03 [{"id": "KITE-139", "title": "SSO login l...
messages now: 5
Turn 3: sent 5 messages, about 1,022 tokens. stop_reason = 'tool_use'
+ assistant tool_use reply_to_customer(ticket_id="KITE-142", message="Hi, since your wor...) id=toolu_04
+ user tool_result id=toolu_04 Reply sent to Northwind Studio on KITE-142.
messages now: 7
Turn 4: sent 7 messages, about 1,150 tokens. stop_reason = 'end_turn'
+ assistant text "Replied to Northwind Studio on KITE-142: after the SSO ..."
messages now: 8
Final answer:
Replied to Northwind Studio on KITE-142: after the SSO change they sign in through their
identity provider, and their IT admin resets passwords [account-recovery#2]. Related:
KITE-139 (closed) was an SSO sign-in loop for the same customer, and KITE-144 (open,
unassigned) looks like the same problem at Harbor Pine.
Look at the first number on each turn: 1, 3, 5, 7. The model is stateless: it remembers nothing between calls. The only reason it knows on turn 3 what the help center said on turn 2 is that you sent turn 2 again. So the loop keeps no memory anywhere except this one list:
A lot follows from that:
- Cost grows faster than the conversation. The token counts (estimated at 4 characters per token) include the system prompt and the tool definitions, which ride along on every call. The last call sent about 1,150 tokens, but the four together sent about 3,190. At
claude-opus-5’s $5 per million input tokens that’s under 2 cents; a 30-turn run with big tool results is a different bill. The real count is inresponse.usage.input_tokens. - Saving an agent means saving the list. Store it as JSON after each turn; resuming is loading it and calling the model again.
python main.py --messagesprints exactly what you’d store. - Debugging means reading the list. When the agent does something odd on turn 5, the messages sent on call 5 are everything it knew. There’s nowhere else to look.
- It only grows. Past a point the model tracks early turns less well, and eventually the list overflows the context window, the most tokens a model can take in one call. Trimming it is the other half of State and Memory.
The turn limit
Now suppose the model never says it’s done. It searches, doesn’t like the result, searches again with other words, and keeps going. Every turn is a paid call that re-sends the whole growing list.
So the outer loop isn’t while True. It counts:
MAX_TURNS = 10
for turn in range(1, max_turns + 1):
... # plan, check, act, observe
return f"Stopped: hit the {max_turns}-turn limit before the agent finished.", messages
python main.py --max-turns 2 shows it: the run stops after the two lookups, before any reply is sent, and the answer is the “Stopped” message instead of a summary.
Start at 10 for a task like this one, which takes 4, and raise it when real tasks need more. Return an explicit error, never an empty string, so “the agent gave up” can’t be mistaken for “the agent had nothing to say”.
Why can't the model just stop itself?
It doesn’t know how long it’s been going. Each call sees the conversation, not a turn counter, and a model that’s stuck doesn’t feel stuck. It sees one more search that might help.
The limit lives in your code, outside the model, which is where a safety rail belongs. It’s a timeout for the loop: set it once and it holds however the model behaves. Observability and Cost Control for Agents adds a budget in dollars on top.
The whole loop
Put the pieces together and this is run_agent from main.py, minus the printing:
def run_agent(client, kb: Kitebase, goal: str, max_turns: int = MAX_TURNS):
messages = [{"role": "user", "content": goal}]
for turn in range(1, max_turns + 1):
response = client.messages.create(model=MODEL, max_tokens=4096,
system=SYSTEM_PROMPT, tools=TOOLS, messages=messages)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
answer = "".join(b.text for b in response.content if b.type == "text")
if response.stop_reason != "end_turn":
answer += f"\n[stopped early: stop_reason={response.stop_reason}]"
return answer, messages
results = []
for block in response.content:
if block.type == "tool_use":
output, is_error = run_tool(kb, block.name, block.input)
result = {"type": "tool_result", "tool_use_id": block.id, "content": output}
if is_error:
result["is_error"] = True
results.append(result)
messages.append({"role": "user", "content": results})
return f"Stopped: hit the {max_turns}-turn limit before the agent finished.", messages
That’s the whole agent: one list, one API call per turn, and your code in between. It takes client as a parameter, so the offline example passes a fake and the real run passes anthropic.Anthropic().
The fake hands back the four canned responses you’ve seen, in order. It doesn’t read tool results, so it can’t test whether the agent is smart. It’s exactly right for testing the loop, and that’s what test_main.py does: messages alternate user and assistant, every tool_use id comes back as a tool_result in the next message, the loop stops on end_turn, and it gives up at the limit.
The shortcut: the SDK’s tool runner
Once you understand this loop, you don’t have to keep writing it. The Python SDK’s tool runner does it for you: decorate plain functions, and it builds the definitions from their type hints and docstrings, then runs the same loop:
from anthropic import Anthropic, beta_tool
@beta_tool
def get_ticket(ticket_id: str) -> str:
"""Look up one support ticket by id: title, status, assignee and customer.
Args:
ticket_id: e.g. KITE-142
"""
return kb.get_ticket(ticket_id)
runner = Anthropic().beta.messages.tool_runner(
model="claude-opus-5", max_tokens=4096, system=SYSTEM_PROMPT,
tools=[get_ticket, ...], # the other four, decorated the same way
messages=[{"role": "user", "content": GOAL}],
max_iterations=10,
)
final = runner.until_done() # the last assistant message
It sends a failing tool back with is_error for you, and max_iterations is your turn limit. One gotcha: when it hits max_iterations it just stops and hands you the last message, so check final.stop_reason. If it’s still "tool_use", the agent didn’t finish.
Anthropic’s SDK docs recommend it as the default for your own tools. It’s beta, as the name says, so pin your SDK version. Either way, you now know what it does under the hood, and that’s what you’ll be reading when you debug it.
Try it yourself
The companion example is the Kitebase agent from this article. Offline, it runs the full loop against the scripted client and prints every turn. With ANTHROPIC_API_KEY set, the same loop runs against claude-opus-5 (add --scripted to force the offline run).
Download the runnable example (zip)
cd 02-agent-loop-in-code
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
python main.py --max-turns 2. The run stops after the two lookups and returns “Stopped: hit the 2-turn limit”. No reply goes to the customer.- In
main.py, delete the linemessages.append({"role": "assistant", "content": response.content})and runpytest -q. The scripted client doesn’t care, but the tests fail on exactly the rule the real API enforces with a 400. - In
scripted_client.py, change the firstget_ticketcall toticket_id="KITE-999"and run it. The result comes back withis_errorand the loop carries on instead of crashing. (The scripted model doesn’t read results, so it carries on with its script. A real model would read the error and search for the ticket.)
pip install pytest && pytest -q runs the offline tests. They need no key and no network.
Common beginner mistakes
- Not appending the assistant turn. The results point at ids the API has never seen, and the request fails with a 400. Append
response.contentas is, before anything else. - Answering only the first tool call. Models batch calls. Loop over every
tool_useblock, or the extra ones go unanswered and the next request is rejected. - Letting a tool exception escape. One bad call kills the run. Catch it and send the error back with
is_error: true. while True. The model can’t be trusted to stop. Count turns and return an explicit “stopped” error at the limit.- Treating
max_tokensas done. A reply cut off bymax_tokensisn’t an answer. Checkstop_reasonbefore trusting anything.
Questions you will face in production
“Should I run a turn’s tool calls in parallel?” If they’re read-only and independent, like the two lookups on turn 2, yes: it saves time on slow APIs. Run them one after another when one writes something another reads. Either way, all the results go back in one message.
“What should the caller get when the turn limit hits?” An explicit failure it can act on, such as a “stopped” status that routes the ticket to a person. Log the full message list with it: it tells you whether the task was too hard, a tool kept failing, or the limit was too low.
“How long can the messages list get?”
Watch tokens, not message count: one big tool result can outweigh twenty short turns. Keep tool outputs small (the help search returns 2 sections, not the whole article), and look at response.usage.input_tokens per turn. When it heads toward the context window, it’s time for the trimming in State and Memory.
Check your understanding
Your loop works for one tool call per turn. The first time the model asks for two, the next request fails with a 400. What's the likely bug?
The loop answered only one of them, probably by reading response.content[0] or returning after the first tool_use block. The API wants a tool_result for every tool_use id in the very next message. Loop over every block and put all the results in one user message.
A reply comes back with stop_reason "max_tokens" and ends partway through a tool_use block. What should the loop do?
Stop and treat it as a failure, not run the tool. The input may be incomplete, so running it could do the wrong thing. Testing stop_reason != "tool_use" rather than “are there tool_use blocks?” handles this for free. If it keeps happening on real tasks, raise max_tokens.
A teammate wants to save money by sending only the last two messages each turn. What breaks?
The model forgets everything else. On turn 4 it would see the reply being sent but not the goal, so it wouldn’t know it was also asked about related tickets. Shrinking the list is possible, but as a deliberate design: keep the goal, summarise old turns, and never separate a tool_use from its tool_result. State and Memory covers how.
Your agent hits the 10-turn limit on 5% of tickets. What do you look at first?
The saved message lists for those runs. Look for the same tool called with slightly different input over and over (a search that never finds anything), or an error the model keeps retrying. Usually the fix is a better tool or a clearer error message, not a higher limit.
What to remember
- An agent is one loop: call the model with the tools and the whole conversation, run what it asks for, send the results back, repeat.
stop_reasondrives the loop."tool_use"means keep going; anything else means stop, and only"end_turn"means it finished properly.- Append every assistant turn exactly as it came back, then answer every
tool_useid with atool_resultin one user message. - Tool failures go back to the model as
is_errorresults, not up your stack as exceptions. - The messages list is the agent’s entire state, and you re-send all of it on every call.
- Cap the turns in your code. The model can’t count them for you.
What to study next
The loop runs whatever tools you hand it, and it’s only as good as those tools. Giving an Agent Tools covers writing definitions the model picks correctly, keeping tool results small, and checking the arguments the model sends before anything with side effects runs, like reply_to_customer.
Further reading
- Anthropic: Tool use with Claude. The API reference for tool definitions,
tool_useandtool_resultblocks, parallel calls and the pairing rules. - Anthropic: Building effective agents. Practical guidance on when to use an agent loop at all, and on keeping it simple.
- ReAct: Synergizing Reasoning and Acting in Language Models. The paper behind the reason-then-act pattern in the text blocks before each tool call.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.