Giving an Agent Tools
You’ve got the loop from Plan, Act, Observe: The Agent Loop in Code running. Now you give it five tools for Kitebase, the made-up project-tracking app, and the support lead types: “Northwind Studio says a user is still locked out after their SSO change. It’s ticket 142. Check what the help center says and reply to the customer with the fix.”
On its first turn Claude asks for two tools at once, and one of them gets "142" where your code expects "KITE-142". Whether that ends in the right reply or a crashed run is decided mostly by things you write: the tool definitions, and what your code sends back when a call works or fails.
What you’ll build: five Kitebase support tools with strict schemas, a validator, and a run of the agent where Claude asks for two tools in one turn, gets an error back for one of them, and recovers. It runs offline with a scripted stand-in for Claude; with an API key it runs against claude-opus-5.
A tool, from Claude’s side
A tool is a function in your code that the model can ask you to run. The model never runs it: it reads a description, sends back the tool’s name and arguments, and your code does the work.
The description you send is the tool definition, and it has three parts:
name: what the model calls it, likeget_ticket.description: plain text saying what it does and when to use it.input_schema: a JSON Schema for the arguments. JSON Schema is a standard JSON format for describing what other JSON may look like: which fields, which types, which are required.
Here’s get_ticket from the companion code’s tools.py:
{
"name": "get_ticket",
"description": (
"Get one Kitebase support ticket by id: its title, status, assignee and customer. "
"If you only know the customer or the topic, use search_tickets first to find the id."
),
"strict": True,
"input_schema": {
"type": "object",
"properties": {"ticket_id": {"type": "string", "description": "A ticket id such as KITE-142."}},
"required": ["ticket_id"],
"additionalProperties": False,
},
}
(strict gets its own section below.) You pass the definitions on every request with tools=TOOLS. When Claude wants a tool, the response has stop_reason: "tool_use" and one tool_use block per call: an id, the tool name and an input dict. You answer with a tool_result block carrying the same id. The agent loop article covers that round trip; this one is about what goes into the definitions and the results. The five tools:
| Tool | Reads or writes | What it does |
|---|---|---|
search_help(query) | reads | Up to 2 help-center sections that match |
get_ticket(ticket_id) | reads | One ticket’s title, status, assignee, customer |
search_tickets(query, status) | reads | Up to 5 ticket summaries |
assign_ticket(ticket_id, assignee) | writes | Sets the ticket’s owner |
reply_to_customer(ticket_id, message) | writes | Emails the ticket’s customer |
The description decides which tool gets picked
When the model chooses a tool, it has the conversation and your definitions. Not your code, not the comment above the function. If two definitions look like they fit, it guesses.
Kitebase has two searches, and both could match an SSO lockout: there are SSO tickets and an SSO help section. Compare what two versions of search_help give the model to go on:
Weak: "Searches help."
Good: "Search Kitebase's public help-center articles for how-to guidance, such as how to
sign in with SSO. Use it to find the fix to send a customer. Returns up to 2 matching
sections, each with an id like account-recovery#2 and at most 300 characters of text.
For customer tickets, use search_tickets instead."
The good one says what it searches, when to use it, what comes back and how much, and which neighbor to use instead. MCP Server Best Practices covers names and descriptions in depth; MCP tools are the same three fields, so the rules carry over.
The gotcha: you can’t tell from your code whether a description works. Write 20 or so real requests with the tool you expect for each, run them against the API, and count the wrong picks. LLM Evaluation Pipelines shows how.
Steering with tool_choice
tool_choice controls whether the model may use tools. Use the default, {"type": "auto"}, which lets it decide. Older examples force a tool with {"type": "any"} or {"type": "tool", "name": "get_ticket"}, but some newer Claude models reject those with a 400 error. If a request must use a tool, say so in the prompt (“Look the ticket up with get_ticket before you answer”), and check the response contains a tool_use block before you rely on it.
strict: true for arguments that match the schema
On its own, the schema is a strong suggestion. A model can leave out a required field, send "in progress" where the allowed value is in_progress, or add a field you never defined, and your code crashes on a KeyError or quietly does the wrong thing.
Set "strict": True and the API constrains the model’s output to your schema. It’s structured outputs (JSON guaranteed to match a schema, from Prompting as Code) applied to tool arguments: required fields present, right types, enum values respected, no extra fields.
Strict schemas have rules. Every object needs "additionalProperties": False, and length limits (minLength, maxLength), number ranges (minimum, maximum) and regex pattern aren’t supported. Put what you can in an enum, and check the rest in code. assign_ticket:
TEAM = ["lena", "priya", "sam"]
"input_schema": {
"type": "object",
"properties": {
"ticket_id": TICKET_ID,
"assignee": {"type": "string", "enum": TEAM, "description": "The teammate's username."},
},
"required": ["ticket_id", "assignee"],
"additionalProperties": False,
},
With the enum, "Priya" or "priyanka" can’t happen: the model picks one of three usernames. Build the list from your real team at startup.
The gotcha: strict guarantees the shape, not the meaning. "142" is a valid string, and "KITE-999" is a well-formed id for a ticket that doesn’t exist. Strict also can’t help when a response was cut off at max_tokens or refused. So your code still checks.
If strict mode guarantees the schema, why does the example validate the input again?
Strict mode isn’t available on every model, provider or account, and code outlives the config it was written for. A schema can also drift from the function it describes; a validator catches that the first time. And it’s cheap: a dict check on a few fields.
In a real project, use the jsonschema package. The example writes its own 20 lines so you can see what gets checked.
Check the input, then check what it means
Every argument is model output. Treat it like the body of a public HTTP request: probably fine, never trusted. Check it in two layers before anything runs.
The first layer is the schema. validate in tools.py handles the parts of JSON Schema these five tools use:
def validate(schema: dict, value, path: str = "input") -> list[str]:
expected = JSON_TYPES[schema["type"]]
if not isinstance(value, expected) or (isinstance(value, bool) and schema["type"] != "boolean"):
return [f"{path} should be a {schema['type']}, got {json.dumps(value)}"]
if "enum" in schema and value not in schema["enum"]:
return [f"{path} must be one of {', '.join(schema['enum'])}; got {json.dumps(value)}"]
problems = []
if schema["type"] == "object":
props = schema.get("properties", {})
problems += [f"{path}.{key} is required" for key in schema.get("required", []) if key not in value]
if schema.get("additionalProperties") is False:
problems += [f"{path}.{key} isn't a known field" for key in value if key not in props]
for key, sub in props.items():
if key in value:
problems += validate(sub, value[key], f"{path}.{key}")
return problems
Given a few bad inputs for assign_ticket, it returns one readable line per problem:
{"ticket_id": "KITE-142"} input.assignee is required
{"ticket_id": 142, "assignee": "priya"} input.ticket_id should be a string, got 142
{"ticket_id": "KITE-142", "assignee": "Priya"} input.assignee must be one of lena, priya, sam; got "Priya"
{"ticket_id": "KITE-142", "assignee": "priya", "notify": true} input.notify isn't a known field
The second layer is the tool itself, for rules that need your data or a regex:
def get_ticket(self, ticket_id: str) -> dict:
ticket_id = ticket_id.strip().upper()
if not re.fullmatch(r"KITE-\d{1,6}", ticket_id):
raise ToolError(f"{ticket_id!r} isn't a ticket id. Ids look like KITE-142. "
"If you only know the customer, use search_tickets.")
if ticket_id not in self.tickets:
raise ToolError(f"No ticket {ticket_id}. Use search_tickets to find the id.")
return dict(self.tickets[ticket_id])
ToolError is the example’s exception for “a failure the model should read”. It’s forgiving where that’s harmless (" kite-142 " works) and strict where a guess does damage: "142" isn’t turned into KITE-142, because a wrong guess replies on someone else’s ticket.
Return errors as results, with is_error
When a tool fails, the worst thing your code can do is raise the exception out of the loop: the run crashes and the model never learns what happened. The error message is all the model has to decide its next move.
So every outcome becomes a tool_result. A failure gets "is_error": True, which tells the model the call failed, rather than succeeded with text that mentions an error. run_tool is the one place that turns a tool_use into a tool_result:
def run_tool(kb: Kitebase, tool_use_id: str, name: str, args: dict) -> dict:
def result(content: str, is_error: bool = False) -> dict:
block = {"type": "tool_result", "tool_use_id": tool_use_id, "content": content}
return {**block, "is_error": True} if is_error else block
if name not in SCHEMAS:
return result(f"There's no tool called {name!r}. Available: {', '.join(SCHEMAS)}.", is_error=True)
if problems := validate(SCHEMAS[name], args):
return result(f"Invalid input for {name}: {'; '.join(problems)}.", is_error=True)
try:
output = getattr(kb, name)(**args)
except ToolError as e:
return result(str(e), is_error=True)
except Exception as e: # a bug or an outage: say so plainly, don't let the model guess what happened
return result(f"{name} failed unexpectedly ({type(e).__name__}). It may not have run.", is_error=True)
return result(json.dumps(output, separators=(",", ":"))) # compact JSON: fewer tokens, same facts
For the worked example’s bad id, the block that goes back to Claude is:
{
"type": "tool_result",
"tool_use_id": "toolu_01",
"content": "'142' isn't a ticket id. Ids look like KITE-142. If you only know the customer, use search_tickets.",
"is_error": true
}
A good error says what was wrong, what a valid value looks like, and what to try next. From “Error executing tool” a model can only give up or retry blindly; from this one it can fix the call. MCP Server Best Practices has more examples.
Note that the catch-all leaves out the exception’s message: ConnectionError("db down at 10.0.3.7") becomes “search_tickets failed unexpectedly (ConnectionError)”. A tool result can end up in a reply to a customer, so hostnames, stack traces and secrets belong in your logs.
Several tools in one turn
The model can ask for several tools in one response when the calls don’t depend on each other. In the worked example it wants the ticket and the help section, and neither needs the other’s answer:
This is a parallel tool call: two tool_use blocks in one assistant turn. The rules on your side:
- Answer every call. One
tool_resultpertool_useid, all in the next user message. Leave one out and the API rejects the next request. - Keep Claude’s order. The results match the
tool_useblocks one for one, error or not. - Results first. If you add your own text to that user message, put it after the
tool_resultblocks, not before. - Only run reads at the same time. Two lookups can’t affect each other. A write can: an assign and a reply in the same turn should happen one after the other, in the model’s order.
run_tool_calls does all four:
def run_tool_calls(kb: Kitebase, content: list) -> list[dict]:
calls = [block for block in content if block.type == "tool_use"]
if any(call.name in WRITE_TOOLS for call in calls):
# Anything that changes state runs one at a time, in the model's order.
return [run_tool(kb, c.id, c.name, c.input) for c in calls]
with ThreadPoolExecutor() as pool: # reads don't affect each other, so run them at the same time
return list(pool.map(lambda c: run_tool(kb, c.id, c.name, c.input), calls)) # map keeps the order
pool.map returns results in input order, not finishing order. The companion tests check it with two reads that each sleep 0.3 seconds: they finish in about 0.3 seconds, not 0.6, and the results still come back as toolu_01, toolu_02, toolu_03.
The first turn of python main.py:
Request 1: about 803 input tokens (803 so far)
Claude asks for 2 tools:
toolu_01 get_ticket({"ticket_id": "142"})
ERROR, about 24 tokens: '142' isn't a ticket id. Ids look like KITE-142. If you only know the cu...
toolu_02 search_help({"query": "locked out sso"})
ok, about 168 tokens: {"sections":[{"id":"account-recovery#2","title":"Locked out of your acco...
Request 2: about 1,040 input tokens (1,843 so far)
Claude asks for 1 tool:
toolu_03 get_ticket({"ticket_id": "KITE-142"})
If parallel calls cause trouble (tools sharing a rate limit, say), "disable_parallel_tool_use": True in tool_choice caps it at one call per response, at the cost of an extra model call per tool.
Keep results small
Everything a tool returns goes into the conversation, and every later request sends the whole conversation again, tool definitions included. The worked example’s four requests:
The five definitions, about 701 tokens sent four times, are two thirds of the run. And a result isn’t paid for once: the search_help result arrives in request 2 and rides along in every request after it. So it caps its output at 2 sections of at most 300 characters, each with an id to cite. In full, about 168 tokens:
{"sections":[{"id":"account-recovery#2","title":"Locked out of your account: If your workspace uses single sign-on","text":"If you sign in through your company's identity provider (SSO), Kitebase doesn't store your password. Ask your IT admin to reset it there."},{"id":"account-recovery#1","title":"Locked out of your account: If you never set a recovery email","text":"Without a recovery email, Kitebase can't send you a reset link. Instead, click **Forgot password?** and then **I can't access my email**. Fill in the form with your workspace URL and the date of your last invoice. Support checks the details and emails a one-time sign-in link to the address you gi..."}]}
Without the caps it returns all 5 sections that share a word with the query, 413 tokens, when the reply needs only the first. On a real help center, where “locked out” matches dozens of pages, whole articles mean thousands of tokens per search, resent every turn, with the right paragraph buried. MCP Server Best Practices measured a raw ticket search at about 11,310 tokens, against 114 for five-field summaries.
The defaults that keep results small:
- Pick the fields.
get_ticketreturns five fields, not the whole database row. - Cap lists and long text, and say the cap in the description, so the model knows there may be more.
- Report the total when you cut a list:
search_ticketsreturns"total"next to the tickets. - Compact JSON.
separators=(",", ":")drops the spaces; the model reads it just as well.
If the model needs more, add a narrower tool (get_help_section(section_id)) rather than raising the cap for everyone. State and Memory covers what to do when the conversation grows too big anyway.
Read tools and write tools
Three of the five tools only read. Two change things: assign_ticket changes a ticket’s owner, and reply_to_customer sends an email a customer will read.
Read tools are cheap to get wrong: a bad search costs a few tokens and a retry. They can run at the same time, repeat safely, and need nobody’s approval.
Write tools need more care, most of it in the definition:
- Say what it changes, and when to call it: “This sends a real email, so only call it when the user asked you to reply.”
- Narrow the target.
reply_to_customer(ticket_id, message)can only reach that ticket’s customer. Asend_email(to, subject, body)tool could email anyone, including an address planted in a ticket comment. - Make repeats harmless.
assign_ticketsets the owner, so calling it twice withpriyaleavespriya. A retry ofreply_to_customeris a second email; MCP Server Best Practices shows how an idempotency key stops that. - Run them one at a time, in the model’s order, as
run_tool_callsdoes. - Leave out what nobody needs. Kitebase’s API can delete tickets. None of the five tools can.
A definition can’t stop a confident, wrong write. For that, a person approves writes before they run. Guardrails and Human in the Loop builds that step.
Try it yourself
The companion example has the five tools, the validator, run_tool, run_tool_calls, the loop, and a scripted stand-in for Claude that plays the worked example’s four turns.
Download the runnable example (zip)
cd 03-giving-an-agent-tools
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
With ANTHROPIC_API_KEY set, python main.py --live sends the same request to claude-opus-5. Its moves won’t match the script exactly.
Then try these:
- In
tools.py, setMAX_HELP_SECTIONS = 13andSECTION_CHARS = 10_000and runmain.py. Thesearch_helpresult goes from about 168 tokens to 413, and every request after it grows by the same 245. - In
scripted_model.py, change the firstget_ticketcall toticket_id=142(a number).validatecatches it before the tool runs:input.ticket_id should be a string, got 142. With strict mode on, a real model can’t send that. - Add a sixth tool: copy the
get_ticketdefinition inTOOLSand rename it. The “Tool definitions” line at the top grows by about 104 tokens, and so does every one of the four requests, though the new tool is never called.
pip install pytest && pytest -q runs the 10 offline tests: strict-ready schemas, is_error results for bad inputs and failing tools, and parallel results in order.
Common beginner mistakes
- One-word descriptions. “Searches help.” gives the model nothing to choose with. Say what, when, what comes back, and which neighbor to use instead.
- Raising instead of returning. An exception out of a tool ends the run. Catch it and send an
is_errorresult the model can act on. - Trusting strict mode for meaning. It guarantees
"142"is a string, not that it’s a ticket. Check the values in code. - Returning whatever the API returned. Full records and whole articles cost tokens on every later turn and bury what matters.
- Running writes in parallel. Two writes in one turn can depend on each other. Run them in order, one at a time.
Questions you will face in production
“How many tools is too many?” There’s no fixed number, but every definition is sent on every request, and more similar tools mean more wrong picks. Start with the five to ten tasks users actually ask for. If two tools keep getting confused, fix their descriptions or merge them.
“Should I use the SDK’s tool runner instead of writing the loop?”
Once you understand the loop, yes. client.beta.messages.tool_runner with the @beta_tool decorator builds schemas from your type hints and runs the loop. The descriptions, small results and useful errors are still yours to write.
“The tool definitions are two thirds of my input tokens. Can I cache them?” Yes. They come first in every request and don’t change, a good fit for prompt caching, where cache reads cost a tenth of the input price. Caching for LLM Apps shows how.
“Do I need MCP to give an agent tools?” No. Tools defined in your code are fine for one app. MCP exposes the same tools to many hosts (Claude Desktop, an IDE, someone else’s agent) without rewriting them; What Is MCP? explains when that’s worth it. An MCP tool is the same name, description and input schema.
Check your understanding
Claude calls search_tickets for "How do I fix SSO sign-in for a customer?" when you wanted search_help. What do you change first?
The descriptions. Make search_help say it’s for how-to guidance and fixes, and make search_tickets say “for how-to guides, use search_help instead”. Then add the request to your test set with the tool you expect.
Your get_ticket tool raises KeyError for an unknown id, and the agent crashes. What should it do instead, and what should the message say?
Catch it and return a tool_result with "is_error": true and the same tool_use_id, saying what was wrong and what to do: “No ticket KITE-999. Use search_tickets to find the id.” The model can recover on the next turn.
You turned on strict: true for every tool. Can you delete your input validation?
Not safely. Strict guarantees the shape (fields, types, enum values), not whether a ticket exists or an id is well formed, and it isn’t available everywhere. Keep the cheap shape check and the checks that need your data.
In one turn, Claude asks for assign_ticket and reply_to_customer on the same ticket. Should your code run them at the same time?
No. Both are writes, and the reply might depend on the new owner, so run them in the order Claude asked. Parallel is for reads. Either way, both results go back in one user message, in the order of the calls.
What to remember
- Claude picks tools from the name, description and input schema alone. Say what a tool does, when to use it instead of its neighbor, and what it returns.
strict: truemakes arguments match the schema. It doesn’t check meaning, so validate values in code.- Every outcome is a
tool_result. Failures getis_error: trueand a message that says how to fix the call. - Answer every
tool_usein one user message, in Claude’s order. Run reads in parallel, writes one at a time. - Results and definitions are resent on every later request. Pick the fields, cap the size, and say the cap.
What to study next
Every result here stayed in the conversation, and the conversation only grows. State and Memory covers what an agent should keep, what to trim, and memory that outlives a run. To make these tools usable from any MCP host, MCP Server Best Practices builds the same five as an MCP server.
Further reading
- Anthropic: Tool use with Claude. Tool definitions,
tool_useandtool_resultblocks,is_error, parallel calls andtool_choice. - Anthropic: Structured outputs. What
strict: trueguarantees and which JSON Schema features it supports. - JSON Schema: Understanding JSON Schema. The format behind
input_schema, includingenumandrequired. - Anthropic: Building effective agents. Its appendix on prompt-engineering your tools is the best short read on tool descriptions.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.