Giving an Agent Tools

You’ve got the loop from Plan, Act, Observe: The Agent Loop in Code running. Now you give it five tools for Kitebase, the made-up project-tracking app, and the support lead types: “Northwind Studio says a user is still locked out after their SSO change. It’s ticket 142. Check what the help center says and reply to the customer with the fix.”

On its first turn Claude asks for two tools at once, and one of them gets "142" where your code expects "KITE-142". Whether that ends in the right reply or a crashed run is decided mostly by things you write: the tool definitions, and what your code sends back when a call works or fails.

What you’ll build: five Kitebase support tools with strict schemas, a validator, and a run of the agent where Claude asks for two tools in one turn, gets an error back for one of them, and recovers. It runs offline with a scripted stand-in for Claude; with an API key it runs against claude-opus-5.

A tool, from Claude’s side

A tool is a function in your code that the model can ask you to run. The model never runs it: it reads a description, sends back the tool’s name and arguments, and your code does the work.

The description you send is the tool definition, and it has three parts:

  • name: what the model calls it, like get_ticket.
  • description: plain text saying what it does and when to use it.
  • input_schema: a JSON Schema for the arguments. JSON Schema is a standard JSON format for describing what other JSON may look like: which fields, which types, which are required.

Here’s get_ticket from the companion code’s tools.py:

{
    "name": "get_ticket",
    "description": (
        "Get one Kitebase support ticket by id: its title, status, assignee and customer. "
        "If you only know the customer or the topic, use search_tickets first to find the id."
    ),
    "strict": True,
    "input_schema": {
        "type": "object",
        "properties": {"ticket_id": {"type": "string", "description": "A ticket id such as KITE-142."}},
        "required": ["ticket_id"],
        "additionalProperties": False,
    },
}

(strict gets its own section below.) You pass the definitions on every request with tools=TOOLS. When Claude wants a tool, the response has stop_reason: "tool_use" and one tool_use block per call: an id, the tool name and an input dict. You answer with a tool_result block carrying the same id. The agent loop article covers that round trip; this one is about what goes into the definitions and the results. The five tools:

ToolReads or writesWhat it does
search_help(query)readsUp to 2 help-center sections that match
get_ticket(ticket_id)readsOne ticket’s title, status, assignee, customer
search_tickets(query, status)readsUp to 5 ticket summaries
assign_ticket(ticket_id, assignee)writesSets the ticket’s owner
reply_to_customer(ticket_id, message)writesEmails the ticket’s customer

The description decides which tool gets picked

When the model chooses a tool, it has the conversation and your definitions. Not your code, not the comment above the function. If two definitions look like they fit, it guesses.

"Northwind Studio says a user is still locked out after their SSO change. It's ticket 142. Check what the help center says and reply to the customer with the fix." What Claude sees: name + description + schema, nothing else search_help(query) Help-center how-to guidance. Use it to find the fix. get_ticket(ticket_id) One ticket by id. Only know the customer? search_tickets. search_tickets(query, status) Tickets by title or customer. For how-to, search_help. assign_ticket(ticket_id, assignee) Changes the owner. Only when the user asked. reply_to_customer(ticket_id, message) Sends a real email. Only when the user asked. CLAUDE matches the request to the text What comes back TOOL_USE toolu_01 get_ticket {"ticket_id": "142"} "ticket 142" in the request became "142", not KITE-142 TOOL_USE toolu_02 search_help {"query": "locked out sso"} picked for "find the fix" SAME TOOLS, VAGUE DESCRIPTIONS search_help "Searches help." search_tickets "Searches tickets." Both fit "locked out after SSO change". The model has to guess. Your Python is never sent. These five definitions, about 700 tokens, are the whole interface Claude has.
The first turn of the worked example. Claude picks from the text of the definitions, nothing else.

Kitebase has two searches, and both could match an SSO lockout: there are SSO tickets and an SSO help section. Compare what two versions of search_help give the model to go on:

Weak:  "Searches help."

Good:  "Search Kitebase's public help-center articles for how-to guidance, such as how to
        sign in with SSO. Use it to find the fix to send a customer. Returns up to 2 matching
        sections, each with an id like account-recovery#2 and at most 300 characters of text.
        For customer tickets, use search_tickets instead."

The good one says what it searches, when to use it, what comes back and how much, and which neighbor to use instead. MCP Server Best Practices covers names and descriptions in depth; MCP tools are the same three fields, so the rules carry over.

The gotcha: you can’t tell from your code whether a description works. Write 20 or so real requests with the tool you expect for each, run them against the API, and count the wrong picks. LLM Evaluation Pipelines shows how.

Steering with tool_choice

tool_choice controls whether the model may use tools. Use the default, {"type": "auto"}, which lets it decide. Older examples force a tool with {"type": "any"} or {"type": "tool", "name": "get_ticket"}, but some newer Claude models reject those with a 400 error. If a request must use a tool, say so in the prompt (“Look the ticket up with get_ticket before you answer”), and check the response contains a tool_use block before you rely on it.

strict: true for arguments that match the schema

On its own, the schema is a strong suggestion. A model can leave out a required field, send "in progress" where the allowed value is in_progress, or add a field you never defined, and your code crashes on a KeyError or quietly does the wrong thing.

Set "strict": True and the API constrains the model’s output to your schema. It’s structured outputs (JSON guaranteed to match a schema, from Prompting as Code) applied to tool arguments: required fields present, right types, enum values respected, no extra fields.

Strict schemas have rules. Every object needs "additionalProperties": False, and length limits (minLength, maxLength), number ranges (minimum, maximum) and regex pattern aren’t supported. Put what you can in an enum, and check the rest in code. assign_ticket:

TEAM = ["lena", "priya", "sam"]

"input_schema": {
    "type": "object",
    "properties": {
        "ticket_id": TICKET_ID,
        "assignee": {"type": "string", "enum": TEAM, "description": "The teammate's username."},
    },
    "required": ["ticket_id", "assignee"],
    "additionalProperties": False,
},

With the enum, "Priya" or "priyanka" can’t happen: the model picks one of three usernames. Build the list from your real team at startup.

The gotcha: strict guarantees the shape, not the meaning. "142" is a valid string, and "KITE-999" is a well-formed id for a ticket that doesn’t exist. Strict also can’t help when a response was cut off at max_tokens or refused. So your code still checks.

If strict mode guarantees the schema, why does the example validate the input again?

Strict mode isn’t available on every model, provider or account, and code outlives the config it was written for. A schema can also drift from the function it describes; a validator catches that the first time. And it’s cheap: a dict check on a few fields.

In a real project, use the jsonschema package. The example writes its own 20 lines so you can see what gets checked.

Check the input, then check what it means

Every argument is model output. Treat it like the body of a public HTTP request: probably fine, never trusted. Check it in two layers before anything runs.

The first layer is the schema. validate in tools.py handles the parts of JSON Schema these five tools use:

def validate(schema: dict, value, path: str = "input") -> list[str]:
    expected = JSON_TYPES[schema["type"]]
    if not isinstance(value, expected) or (isinstance(value, bool) and schema["type"] != "boolean"):
        return [f"{path} should be a {schema['type']}, got {json.dumps(value)}"]
    if "enum" in schema and value not in schema["enum"]:
        return [f"{path} must be one of {', '.join(schema['enum'])}; got {json.dumps(value)}"]
    problems = []
    if schema["type"] == "object":
        props = schema.get("properties", {})
        problems += [f"{path}.{key} is required" for key in schema.get("required", []) if key not in value]
        if schema.get("additionalProperties") is False:
            problems += [f"{path}.{key} isn't a known field" for key in value if key not in props]
        for key, sub in props.items():
            if key in value:
                problems += validate(sub, value[key], f"{path}.{key}")
    return problems

Given a few bad inputs for assign_ticket, it returns one readable line per problem:

{"ticket_id": "KITE-142"}                                  input.assignee is required
{"ticket_id": 142, "assignee": "priya"}                    input.ticket_id should be a string, got 142
{"ticket_id": "KITE-142", "assignee": "Priya"}             input.assignee must be one of lena, priya, sam; got "Priya"
{"ticket_id": "KITE-142", "assignee": "priya", "notify": true}   input.notify isn't a known field

The second layer is the tool itself, for rules that need your data or a regex:

def get_ticket(self, ticket_id: str) -> dict:
    ticket_id = ticket_id.strip().upper()
    if not re.fullmatch(r"KITE-\d{1,6}", ticket_id):
        raise ToolError(f"{ticket_id!r} isn't a ticket id. Ids look like KITE-142. "
                        "If you only know the customer, use search_tickets.")
    if ticket_id not in self.tickets:
        raise ToolError(f"No ticket {ticket_id}. Use search_tickets to find the id.")
    return dict(self.tickets[ticket_id])

ToolError is the example’s exception for “a failure the model should read”. It’s forgiving where that’s harmless (" kite-142 " works) and strict where a guess does damage: "142" isn’t turned into KITE-142, because a wrong guess replies on someone else’s ticket.

Return errors as results, with is_error

When a tool fails, the worst thing your code can do is raise the exception out of the loop: the run crashes and the model never learns what happened. The error message is all the model has to decide its next move.

So every outcome becomes a tool_result. A failure gets "is_error": True, which tells the model the call failed, rather than succeeded with text that mentions an error. run_tool is the one place that turns a tool_use into a tool_result:

def run_tool(kb: Kitebase, tool_use_id: str, name: str, args: dict) -> dict:
    def result(content: str, is_error: bool = False) -> dict:
        block = {"type": "tool_result", "tool_use_id": tool_use_id, "content": content}
        return {**block, "is_error": True} if is_error else block

    if name not in SCHEMAS:
        return result(f"There's no tool called {name!r}. Available: {', '.join(SCHEMAS)}.", is_error=True)
    if problems := validate(SCHEMAS[name], args):
        return result(f"Invalid input for {name}: {'; '.join(problems)}.", is_error=True)
    try:
        output = getattr(kb, name)(**args)
    except ToolError as e:
        return result(str(e), is_error=True)
    except Exception as e:  # a bug or an outage: say so plainly, don't let the model guess what happened
        return result(f"{name} failed unexpectedly ({type(e).__name__}). It may not have run.", is_error=True)
    return result(json.dumps(output, separators=(",", ":")))  # compact JSON: fewer tokens, same facts

For the worked example’s bad id, the block that goes back to Claude is:

{
  "type": "tool_result",
  "tool_use_id": "toolu_01",
  "content": "'142' isn't a ticket id. Ids look like KITE-142. If you only know the customer, use search_tickets.",
  "is_error": true
}

A good error says what was wrong, what a valid value looks like, and what to try next. From “Error executing tool” a model can only give up or retry blindly; from this one it can fix the call. MCP Server Best Practices has more examples.

Note that the catch-all leaves out the exception’s message: ConnectionError("db down at 10.0.3.7") becomes “search_tickets failed unexpectedly (ConnectionError)”. A tool result can end up in a reply to a customer, so hostnames, stack traces and secrets belong in your logs.

Several tools in one turn

The model can ask for several tools in one response when the calls don’t depend on each other. In the worked example it wants the ticket and the help section, and neither needs the other’s answer:

1. CLAUDE'S RESPONSE stop_reason: "tool_use" text "I'll look up both at once." tool_use toolu_01 get_ticket {"ticket_id": "142"} tool_use toolu_02 search_help {"query": "locked out sso"} 2. YOUR CODE Validate each input against its schema. Both are reads, so run them at the same time. toolu_01 get_ticket ToolError: bad id toolu_02 search_help 2 sections Wrap every outcome as a tool_result with the same id, in the order Claude asked. 3. ONE USER MESSAGE role: "user" tool_result toolu_01 is_error: true "'142' isn't a ticket id. Ids look like KITE-142." about 24 tokens tool_result toolu_02 {"sections": [ account-recovery#2, account-recovery#1]} about 168 tokens 4. NEXT TURN: CLAUDE READS THE ERROR "The id needs its KITE- prefix." tool_use toolu_03 get_ticket {"ticket_id": "KITE-142"} RESULT, ABOUT 34 TOKENS {"id":"KITE-142","status":"in_progress", "assignee":"priya", "customer":"Northwind Studio", ...} The error was a result, not an exception. The model read it and fixed the id. Every result stays in the conversation, sent again with each later request.
One failed call doesn't spoil the turn. The other result still arrives, and the model fixes the one that failed.

This is a parallel tool call: two tool_use blocks in one assistant turn. The rules on your side:

  • Answer every call. One tool_result per tool_use id, all in the next user message. Leave one out and the API rejects the next request.
  • Keep Claude’s order. The results match the tool_use blocks one for one, error or not.
  • Results first. If you add your own text to that user message, put it after the tool_result blocks, not before.
  • Only run reads at the same time. Two lookups can’t affect each other. A write can: an assign and a reply in the same turn should happen one after the other, in the model’s order.

run_tool_calls does all four:

def run_tool_calls(kb: Kitebase, content: list) -> list[dict]:
    calls = [block for block in content if block.type == "tool_use"]
    if any(call.name in WRITE_TOOLS for call in calls):
        # Anything that changes state runs one at a time, in the model's order.
        return [run_tool(kb, c.id, c.name, c.input) for c in calls]
    with ThreadPoolExecutor() as pool:  # reads don't affect each other, so run them at the same time
        return list(pool.map(lambda c: run_tool(kb, c.id, c.name, c.input), calls))  # map keeps the order

pool.map returns results in input order, not finishing order. The companion tests check it with two reads that each sleep 0.3 seconds: they finish in about 0.3 seconds, not 0.6, and the results still come back as toolu_01, toolu_02, toolu_03.

The first turn of python main.py:

Request 1: about 803 input tokens (803 so far)
  Claude asks for 2 tools:
    toolu_01  get_ticket({"ticket_id": "142"})
              ERROR, about 24 tokens: '142' isn't a ticket id. Ids look like KITE-142. If you only know the cu...
    toolu_02  search_help({"query": "locked out sso"})
              ok, about 168 tokens: {"sections":[{"id":"account-recovery#2","title":"Locked out of your acco...
Request 2: about 1,040 input tokens (1,843 so far)
  Claude asks for 1 tool:
    toolu_03  get_ticket({"ticket_id": "KITE-142"})

If parallel calls cause trouble (tools sharing a rate limit, say), "disable_parallel_tool_use": True in tool_choice caps it at one call per response, at the cost of an extra model call per tool.

Keep results small

Everything a tool returns goes into the conversation, and every later request sends the whole conversation again, tool definitions included. The worked example’s four requests:

What each request to Claude carries in the worked example Estimated input tokens, 4 characters per token. Each request resends everything before it. 5 tool definitions: 701 Request 1 803 Request 2 1,040 Request 3 1,095 Request 4 1,217 search_help: 168 4 requests: 4,155 tool definitions system prompt + request Claude's tool calls tool results error result DEFINITIONS RIDE ON EVERY REQUEST 701 tokens x 4 requests = 2,804, two thirds of the whole run. A 6th tool costs you on every call, used or not. SO DOES EVERY RESULT AFTER IT ARRIVES search_help's 168 tokens go out 3 times. Uncapped (5 sections, 413) it's 1,239. Cap results, and say the cap in the description.
Token counts are the example's 4-characters-per-token estimates. The API also adds a fixed amount for its own tool-use instructions.

The five definitions, about 701 tokens sent four times, are two thirds of the run. And a result isn’t paid for once: the search_help result arrives in request 2 and rides along in every request after it. So it caps its output at 2 sections of at most 300 characters, each with an id to cite. In full, about 168 tokens:

{"sections":[{"id":"account-recovery#2","title":"Locked out of your account: If your workspace uses single sign-on","text":"If you sign in through your company's identity provider (SSO), Kitebase doesn't store your password. Ask your IT admin to reset it there."},{"id":"account-recovery#1","title":"Locked out of your account: If you never set a recovery email","text":"Without a recovery email, Kitebase can't send you a reset link. Instead, click **Forgot password?** and then **I can't access my email**. Fill in the form with your workspace URL and the date of your last invoice. Support checks the details and emails a one-time sign-in link to the address you gi..."}]}

Without the caps it returns all 5 sections that share a word with the query, 413 tokens, when the reply needs only the first. On a real help center, where “locked out” matches dozens of pages, whole articles mean thousands of tokens per search, resent every turn, with the right paragraph buried. MCP Server Best Practices measured a raw ticket search at about 11,310 tokens, against 114 for five-field summaries.

The defaults that keep results small:

  • Pick the fields. get_ticket returns five fields, not the whole database row.
  • Cap lists and long text, and say the cap in the description, so the model knows there may be more.
  • Report the total when you cut a list: search_tickets returns "total" next to the tickets.
  • Compact JSON. separators=(",", ":") drops the spaces; the model reads it just as well.

If the model needs more, add a narrower tool (get_help_section(section_id)) rather than raising the cap for everyone. State and Memory covers what to do when the conversation grows too big anyway.

Read tools and write tools

Three of the five tools only read. Two change things: assign_ticket changes a ticket’s owner, and reply_to_customer sends an email a customer will read.

Read tools are cheap to get wrong: a bad search costs a few tokens and a retry. They can run at the same time, repeat safely, and need nobody’s approval.

Write tools need more care, most of it in the definition:

  • Say what it changes, and when to call it: “This sends a real email, so only call it when the user asked you to reply.”
  • Narrow the target. reply_to_customer(ticket_id, message) can only reach that ticket’s customer. A send_email(to, subject, body) tool could email anyone, including an address planted in a ticket comment.
  • Make repeats harmless. assign_ticket sets the owner, so calling it twice with priya leaves priya. A retry of reply_to_customer is a second email; MCP Server Best Practices shows how an idempotency key stops that.
  • Run them one at a time, in the model’s order, as run_tool_calls does.
  • Leave out what nobody needs. Kitebase’s API can delete tickets. None of the five tools can.

A definition can’t stop a confident, wrong write. For that, a person approves writes before they run. Guardrails and Human in the Loop builds that step.

Try it yourself

The companion example has the five tools, the validator, run_tool, run_tool_calls, the loop, and a scripted stand-in for Claude that plays the worked example’s four turns.

Download the runnable example (zip)

cd 03-giving-an-agent-tools
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

With ANTHROPIC_API_KEY set, python main.py --live sends the same request to claude-opus-5. Its moves won’t match the script exactly.

Then try these:

  1. In tools.py, set MAX_HELP_SECTIONS = 13 and SECTION_CHARS = 10_000 and run main.py. The search_help result goes from about 168 tokens to 413, and every request after it grows by the same 245.
  2. In scripted_model.py, change the first get_ticket call to ticket_id=142 (a number). validate catches it before the tool runs: input.ticket_id should be a string, got 142. With strict mode on, a real model can’t send that.
  3. Add a sixth tool: copy the get_ticket definition in TOOLS and rename it. The “Tool definitions” line at the top grows by about 104 tokens, and so does every one of the four requests, though the new tool is never called.

pip install pytest && pytest -q runs the 10 offline tests: strict-ready schemas, is_error results for bad inputs and failing tools, and parallel results in order.

Common beginner mistakes

  • One-word descriptions. “Searches help.” gives the model nothing to choose with. Say what, when, what comes back, and which neighbor to use instead.
  • Raising instead of returning. An exception out of a tool ends the run. Catch it and send an is_error result the model can act on.
  • Trusting strict mode for meaning. It guarantees "142" is a string, not that it’s a ticket. Check the values in code.
  • Returning whatever the API returned. Full records and whole articles cost tokens on every later turn and bury what matters.
  • Running writes in parallel. Two writes in one turn can depend on each other. Run them in order, one at a time.

Questions you will face in production

“How many tools is too many?” There’s no fixed number, but every definition is sent on every request, and more similar tools mean more wrong picks. Start with the five to ten tasks users actually ask for. If two tools keep getting confused, fix their descriptions or merge them.

“Should I use the SDK’s tool runner instead of writing the loop?” Once you understand the loop, yes. client.beta.messages.tool_runner with the @beta_tool decorator builds schemas from your type hints and runs the loop. The descriptions, small results and useful errors are still yours to write.

“The tool definitions are two thirds of my input tokens. Can I cache them?” Yes. They come first in every request and don’t change, a good fit for prompt caching, where cache reads cost a tenth of the input price. Caching for LLM Apps shows how.

“Do I need MCP to give an agent tools?” No. Tools defined in your code are fine for one app. MCP exposes the same tools to many hosts (Claude Desktop, an IDE, someone else’s agent) without rewriting them; What Is MCP? explains when that’s worth it. An MCP tool is the same name, description and input schema.

Check your understanding

Claude calls search_tickets for "How do I fix SSO sign-in for a customer?" when you wanted search_help. What do you change first?

The descriptions. Make search_help say it’s for how-to guidance and fixes, and make search_tickets say “for how-to guides, use search_help instead”. Then add the request to your test set with the tool you expect.

Your get_ticket tool raises KeyError for an unknown id, and the agent crashes. What should it do instead, and what should the message say?

Catch it and return a tool_result with "is_error": true and the same tool_use_id, saying what was wrong and what to do: “No ticket KITE-999. Use search_tickets to find the id.” The model can recover on the next turn.

You turned on strict: true for every tool. Can you delete your input validation?

Not safely. Strict guarantees the shape (fields, types, enum values), not whether a ticket exists or an id is well formed, and it isn’t available everywhere. Keep the cheap shape check and the checks that need your data.

In one turn, Claude asks for assign_ticket and reply_to_customer on the same ticket. Should your code run them at the same time?

No. Both are writes, and the reply might depend on the new owner, so run them in the order Claude asked. Parallel is for reads. Either way, both results go back in one user message, in the order of the calls.

What to remember

  • Claude picks tools from the name, description and input schema alone. Say what a tool does, when to use it instead of its neighbor, and what it returns.
  • strict: true makes arguments match the schema. It doesn’t check meaning, so validate values in code.
  • Every outcome is a tool_result. Failures get is_error: true and a message that says how to fix the call.
  • Answer every tool_use in one user message, in Claude’s order. Run reads in parallel, writes one at a time.
  • Results and definitions are resent on every later request. Pick the fields, cap the size, and say the cap.

What to study next

Every result here stayed in the conversation, and the conversation only grows. State and Memory covers what an agent should keep, what to trim, and memory that outlives a run. To make these tools usable from any MCP host, MCP Server Best Practices builds the same five as an MCP server.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.