What an AI Agent Actually Is
A ticket lands in Kitebase’s support queue: KITE-142, “Customer locked out after SSO change”. Kitebase is the small project-tracking app used across this site, and someone wants AI to handle tickets like this one. The first suggestion in the meeting is “build an agent”.
That word gets stuck on chatbots, cron jobs, RAG pipelines and anything else with a model inside, so it tells you little about what to build. This article gives it a mechanical meaning you can check against code, and answers the question that matters more: does this ticket need an agent at all?
What you’ll build: the same Kitebase ticket handled three ways (a single model call, a fixed workflow and an agent loop), printing every step, how many model calls each one made and what it cost. It runs offline with a scripted stand-in for Claude; an API key swaps in the real model.
An agent is a model in a loop with tools
Here’s the whole idea in one sentence: an agent is a program where the model decides the next step, runs in a loop, and keeps going until it decides the task is done.
It has three parts:
- A model. On its own it’s a stateless function: text in, text out. It remembers nothing between calls and can’t touch the outside world.
- Tools. A tool is a function in your code that you describe to the model: a name, a one-line description and the arguments it takes. The model can’t run it. It can only ask, and your code runs it.
- A loop. Your code calls the model, runs whatever tool it asked for, hands back the result and calls the model again. The loop ends when the model answers without asking for a tool.
The support agent in this article has five tools, and the other articles in this topic use the same set:
| Tool | What it does |
|---|---|
get_ticket(ticket_id) | Title, status, assignee and customer for one ticket |
search_tickets(query, status) | Tickets whose title or customer matches |
search_help(query) | The two best-matching sections of the help center |
assign_ticket(ticket_id, assignee) | Give the ticket to an engineer |
reply_to_customer(ticket_id, message) | Send the customer a reply |
A tool definition is plain data. Here’s get_ticket as the API receives it:
{"name": "get_ticket",
"description": "Look up one support ticket by id: title, status, assignee and customer.",
"input_schema": {"type": "object", "required": ["ticket_id"],
"properties": {"ticket_id": {"type": "string", "description": "e.g. KITE-142"}}}}
input_schema is a JSON Schema: a standard way to describe the shape of some JSON, here “an object with a string called ticket_id”. When the model wants the tool, its reply doesn’t end in a finished answer. It ends in a request:
stop_reason: "tool_use"
content: [
{"type": "text", "text": "I'll look up the ticket first."},
{"type": "tool_use", "id": "toolu_01", "name": "get_ticket", "input": {"ticket_id": "KITE-142"}}
]
stop_reason says why the model stopped writing, as in Your First LLM Integration. "tool_use" means “run this and tell me what happened”. Your code runs get_ticket("KITE-142") and sends the output back as a tool result, tagged with the same id so the model knows which request it answers.
Three ways to handle one ticket
The companion code handles KITE-142 three ways. What changes between them is who picks the next step.
1. A single call
Paste the ticket title into one prompt and ask for a reply:
def single_call(client, kb, ticket_id, run):
title = kb.tickets[ticket_id]["title"]
response = call_model(client, run, system=WRITER_PROMPT,
messages=[{"role": "user", "content": f'Reply to the customer on this ticket: "{title}"'}])
return text_of(response)
call_model is client.messages.create plus a counter for calls and tokens. The output:
1. SINGLE CALL: one prompt, one reply, no tools
[model] writes a reply from the ticket title alone
reply: Sorry you're locked out! Click Forgot password? on the sign-in page...
sent to customer: no assignee after: priya
model calls: 1 tool calls: 0 tokens: 78 in, 38 out cost: $0.0013
That reply is wrong. The customer’s company signs in with SSO (single sign-on: they log in through their company’s identity provider, such as Okta or Google Workspace, not with a Kitebase password). Kitebase’s help center says Kitebase doesn’t store SSO passwords, so a reset link can’t help. The model never saw the help center, so it wrote the most likely-sounding answer, the same hallucination problem RAG Explained opens with. Nothing got sent either: a single call returns text, and doing anything with it is up to your code.
A single call is still right for plenty of jobs: summarising, classifying or rewriting a ticket needs nothing beyond the prompt.
2. A fixed workflow
A workflow is code you write that calls the model at fixed points. You decide the steps and their order; the model fills in one of them. For support tickets: look up the ticket, search the help center for its title, have the model draft a reply from the article it found, send it.
def workflow(client, kb, ticket_id, run):
ticket = json.loads(call_tool(kb, run, "get_ticket", {"ticket_id": ticket_id})[0])
help_text, _ = call_tool(kb, run, "search_help", {"query": ticket["title"]})
prompt = f'Ticket {ticket_id}: "{ticket["title"]}"\n\n<help>\n{help_text}\n</help>\n\nWrite the reply.'
reply = text_of(call_model(client, run, system=WRITER_PROMPT, messages=[{"role": "user", "content": prompt}]))
call_tool(kb, run, "reply_to_customer", {"ticket_id": ticket_id, "message": reply})
return reply
(Trimmed: the version in main.py also logs each step.) The output:
2. FIXED WORKFLOW: your code picks every step
[code] get_ticket("KITE-142")
[code] search_help("Customer locked out after SSO change") -> account-recovery#2
[model] drafts the reply from that article
[code] reply_to_customer("KITE-142", "Your workspace signs in wit...")
reply: Your workspace signs in with single sign-on (SSO), so Kitebase does...
sent to customer: yes assignee after: priya
model calls: 1 tool calls: 3 tokens: 274 in, 63 out cost: $0.0029
The search found account-recovery#2, the help section for SSO workspaces, and the reply tells the customer to ask their IT admin. Right answer, one model call, the same four steps every run. The tools are the same functions the agent uses, but your code calls them; the model never sees them.
3. An agent loop
Now hand the model the task, the five tools and a system prompt (instructions that apply to the whole conversation): look the ticket up, find the help article, reply, and assign it if nobody owns it. Then loop:
def run_agent(client, kb, task, run, max_turns=MAX_TURNS):
messages = [{"role": "user", "content": task}]
for turn in range(1, max_turns + 1):
response = call_model(client, run, system=AGENT_PROMPT, tools=TOOLS, messages=messages)
messages.append({"role": "assistant", "content": response.content})
log_turn(run, turn, response)
if response.stop_reason != "tool_use": # no tool requested: the model says it's done
return text_of(response)
results = []
for block in response.content:
if block.type == "tool_use":
output, is_error = call_tool(kb, run, block.name, block.input)
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": output, "is_error": is_error})
messages.append({"role": "user", "content": results}) # what the model sees next turn
return f"Stopped after {max_turns} turns without finishing."
That’s a complete agent. Nothing in it mentions tickets or the order to do things in. Your code only decides “run what was asked” and “stop when nothing was asked, or after max_turns”. (log_turn just records the trace below.) The output:
3. AGENT LOOP: the model picks every step
turn 1 [model] "I'll look up the ticket first."
-> get_ticket(ticket_id="KITE-142")
turn 2 [model] "Now the help article."
-> search_help(query="Customer locked out aft...")
turn 3 [model] "I have what I need to reply."
-> reply_to_customer(ticket_id="KITE-142", message="Your workspace signs in...")
turn 4 [model] done: "Replied to Northwind Studio on KITE-142."
reply: Your workspace signs in with single sign-on (SSO), so Kitebase does...
sent to customer: yes assignee after: priya
model calls: 4 tool calls: 3 tokens: 2,853 in, 199 out cost: $0.0192
input tokens per call: 450, 569, 842, 992
Same three tools, same reply as the workflow, four model calls instead of one and 6.6 times the cost. For KITE-142 the agent bought nothing. Hold that thought; the next ticket is different.
Offline, “the model” is scripted_model.py, a stand-in that returns the SDK’s response shape but picks steps with a few fixed rules. With ANTHROPIC_API_KEY set, claude-opus-5 makes the choices, and it may pick a different order or ask for two tools in one turn. The inner for loop handles that: every tool_use block gets run, and all the results go back in one message.
The loop at a glance
Every agent, however it’s dressed up, runs this loop. You’ll often see it described as plan, act, observe: the model plans the next step, your code acts on it, and the result goes back so the model can observe what happened.
Three things in this picture matter more than they look:
- The messages list is the agent’s memory. The model is stateless, so every turn your code re-sends the whole conversation: the task, every tool call, every result. That’s how turn 3 “knows” what turn 2 found.
- Your code does all the acting. The model only returns text and requests.
reply_to_customerruns because your loop chose to run it, which is also where you can stop it (Guardrails and Human in the Loop). - The step limit lives in your code. Nothing in the model stops a confused one from asking for tools forever, and every turn is a paid call.
MAX_TURNS = 10is the safety rail.
Why a loop, and not one big prompt with the tools listed?
The model can’t know a tool’s result before the tool runs. To decide whether to assign KITE-144, it has to see that nobody owns it, and it can’t see that until get_ticket has actually run.
Ask for the whole plan in one prompt and the model has to guess each result, and it will happily write a plausible one that’s wrong. The loop grounds every decision in real output: one step, the real result, then the next decision.
Where an agent earns its keep: KITE-144
Run the same code on a different ticket: python main.py KITE-144, “SSO users can’t reset their password”, from Harbor Pine. Nobody is assigned to it. The workflow does what it always does:
2. FIXED WORKFLOW: your code picks every step
[code] get_ticket("KITE-144")
[code] search_help("SSO users can't reset their password") -> account-recovery#2
[model] drafts the reply from that article
[code] reply_to_customer("KITE-144", "Your workspace signs in wit...")
sent to customer: yes assignee after: nobody
The customer gets the right answer, and the ticket still belongs to no one, so nobody follows up. The agent’s run, with no change to its code:
3. AGENT LOOP: the model picks every step
turn 1 [model] "I'll look up the ticket first."
-> get_ticket(ticket_id="KITE-144")
turn 2 [model] "Nobody owns this. I'll check who handled similar 'sso' ti..."
-> search_tickets(query="sso", status="any")
turn 3 [model] "priya handled the similar ones."
-> assign_ticket(ticket_id="KITE-144", assignee="priya")
turn 4 [model] "Now the help article."
-> search_help(query="SSO users can't reset t...")
turn 5 [model] "I have what I need to reply."
-> reply_to_customer(ticket_id="KITE-144", message="Your workspace signs in...")
turn 6 [model] done: "Replied to Harbor Pine on KITE-144. KITE-144 assigned..."
sent to customer: yes assignee after: priya
model calls: 6 tool calls: 5 tokens: 5,078 in, 300 out cost: $0.0329
input tokens per call: 450, 564, 774, 865, 1,138, 1,287
Turn 2 is the one to look at. After seeing "assignee": null in turn 1’s result, the model took a step it didn’t take for KITE-142. It found KITE-139 and KITE-142, both SSO tickets handled by priya, and gave her this one. Same loop, different ticket, different path. That’s the definition in action: the path depends on what the model finds along the way.
Anthropic’s Building effective agents draws the same line. Workflows run tools “through predefined code paths”; in agents, the model directs its own process and tool use.
Now the honest part. In scripted_model.py, the stand-in’s “decision” at turn 2 is an if: nobody assigned, so search for similar tickets. If you can write a rule down, you can put it in the workflow, and the third exercise in “Try it yourself” does: one model call, assigned to priya, about a tenth of the cost. An agent is worth it when the rules are too many or too fuzzy to write: a refund, a bug report, a question back to the customer, an escalation, or three of those in an order you can’t predict.
What the loop costs
The agent costs more in three ways, and all three grow with every turn.
Money. Models bill per token, the unit they read text in (about 4 characters of English). At claude-opus-5’s $5 per million input tokens and $25 per million output, the agent’s run on KITE-144 is $0.0329 and the workflow’s is $0.0029. At 1,000 tickets a day that’s roughly $33 against $3 (using the stand-in’s token estimates). The gap comes from the messages list: each call re-sends everything before it, and the five tool definitions ride along every time.
Time. Each model call is a full round trip to the API, and turn 3 can’t start until turn 2’s result is back. Six calls take roughly six times as long as one, and the customer or the support engineer waits for all of them.
Reliability. Every turn is a decision the model can get wrong: the wrong tool, the right tool with a bad argument, a search that misses. The mistakes compound. If the model picked the right step 95% of the time (a made-up number, just for the arithmetic), six picks in a row would all be right 0.95⁶ ≈ 74% of the time. A workflow makes one model decision and the rest is your code, which does the same thing every run.
The failure modes show up the first day you run one:
- It never stops, calling tools in circles. That’s why
MAX_TURNSexists. - It picks the wrong tool or passes bad arguments, then acts on the result (Giving an Agent Tools).
- A tool fails and it retries the same broken call. Errors go back as tool results with
is_errorso it can change course (Retries, Timeouts, and Failure Handling). - It loses the thread as the messages list grows (State and Memory).
The loop is the easy part. The reliability work around it is most of the job.
Which one to build
Start from the simplest thing that works and move up only when it can’t do the job:
- Single call when everything the model needs fits in the prompt and nothing needs doing afterwards: summarise, classify, rewrite, extract.
- Fixed workflow when you can write the steps down, even with a few branches. It’s cheaper, faster and does the same thing every run. Plenty of useful AI features are exactly this.
- Agent when the next step genuinely depends on what the last one found, and the branches are too many to write. Accept that you’re paying for it in tokens, time and testing.
For KITE-142 and KITE-144, build the workflow with an if assignee is None branch. Mixing is normal too: a workflow can hand one fuzzy step, say “work out what this customer actually needs”, to a small agent and take back control afterwards.
Try it yourself
The companion example runs all three approaches on any ticket in data/tickets.json, printing each step and the counts. It runs offline; set ANTHROPIC_API_KEY to make the same calls to claude-opus-5.
Download the runnable example (zip)
cd 01-what-is-an-agent
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py # KITE-142
python main.py KITE-144 # nobody assigned
Then try these:
python main.py KITE-143(“Invite emails going to spam”). No help article covers it and nobody has handled a similar ticket, so the workflow and the agent send the same holding reply. The agent takes 5 model calls to get there and costs about 13 times as much.- Set
MAX_TURNS = 3inmain.pyand run KITE-144. The agent assigns the ticket, then gets cut off before it replies to the customer. The safety rail worked, and it left a half-done job, so picking the limit takes some thought. - Give the workflow the branch it’s missing. In
workflow(), afterget_ticket, add: ifticket["assignee"]isNone, callsearch_ticketswith the first word of the title andassign_ticketto the most common assignee in the results. Run KITE-144 again: assigned to priya, one model call.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- Building an agent when the steps are known. If you could draw the flowchart, write the workflow.
- No step limit.
while Truearound a model call is a runaway bill waiting for a confused turn. Always count turns. - Dropping the history. Sending only the latest tool result instead of the whole messages list. The model is stateless, so it forgets what it already did and does it again.
- Treating tool calls as trusted. The model writes the arguments. Check them before running anything that changes data or talks to a customer.
- Assuming the demo path is the path. A real model may take a different route on the next run. Test the outcome (ticket assigned, reply sent), not the exact sequence of calls.
Questions you will face in production
“How do I stop an agent from looping forever?” Count turns in your code and stop at a limit, 10 to 20 for most tasks. Add a token or time budget on top if a single run could get expensive. When the limit hits, return a clear “didn’t finish” result instead of pretending it worked.
“Do I need a framework?”
Not to start. The loop above is under 20 lines, and writing it yourself shows you what every framework does underneath. Once it’s familiar, the Anthropic SDK’s tool runner (client.beta.messages.tool_runner) runs the same loop for you.
“Where does MCP fit?”
MCP is a standard way to package tools so any app can use them. The Kitebase tools here are plain Python functions; Setting Up Your First MCP Server builds get_ticket and search_tickets as an MCP server, and What Is MCP? explains when that’s worth doing. The agent loop is the same either way.
Check your understanding
Your script calls the model to classify a ticket, then your code routes it to billing or support based on the answer. Is that an agent?
No, it’s a workflow. The model fills in one step, but your code decided there would be exactly one classification followed by a route. It would be an agent if the model chose what to do next and the loop kept going on its choices.
An agent handles KITE-142 correctly in 4 calls, and a teammate proposes the same agent for every ticket. What do you check first?
Whether the tickets actually need different paths. If most take the same few steps, a workflow gives the same result for about a sixth of the cost here, runs faster and is easier to test. Look at a sample of real tickets and try to write the steps down before choosing.
Your agent's bill is five times what you estimated from "about 500 tokens per call". Why?
Each call re-sends the whole conversation, so calls get bigger every turn: 450 tokens on turn 1 and 1,287 by turn 6 for KITE-144. Estimate from the total across all turns, not the first call. Long tool results and many tool definitions make it grow faster.
A customer got the same reply three times. Your logs show the agent called reply_to_customer on turns 3, 5 and 7. What probably went wrong?
The model didn’t see its earlier replies, most likely because the loop dropped the history or the tool results never made it back into the messages list. Check that every assistant turn and every tool result is appended. A guard in your code that refuses a second reply on the same ticket stops it happening even when the model slips.
What to remember
- An agent is a model in a loop with tools: it picks the next step, your code runs it, the result goes back, until the model stops asking for tools.
- The difference from a workflow is who picks the steps. In a workflow it’s your code; in an agent it’s the model.
- The messages list is the agent’s only memory, and it’s re-sent every turn, so each call costs more than the last.
- Agents cost more money, time and reliability than a workflow for the same result. Pay that only when the path really depends on what the model finds.
- Always cap the loop with a turn limit in your code.
What to study next
You’ve seen the loop at its smallest. Plan, Act, Observe: The Agent Loop in Code turns it into one you’d run for real: a proper stop condition, several tool calls per turn, and tool errors fed back to the model instead of crashing the loop. After that, Giving an Agent Tools covers designing tools the model picks correctly.
Further reading
- Anthropic: Building effective agents. A practical breakdown of workflows versus agents and when to use each. Start here.
- ReAct: Synergizing Reasoning and Acting in Language Models. The paper behind the plan, act, observe pattern most agents use.
- Anthropic: Tool use documentation. How tool calls work at the API level, the mechanism under every agent.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.