Multi-Agent Systems: When and When Not

Ten support tickets came into Kitebase overnight, and someone wants them triaged before the support engineers log on. The proposal: a lead agent that hands the batch to specialist agents, one for sign-in problems, one for billing, one for bugs. It sounds like how a real support team works.

Before building it, run the numbers. This article triages the same ten tickets three ways and compares what each costs, how long it takes and what goes wrong. The answer it lands on is the one to start from: use one agent until you can show it isn’t enough.

What you’ll build: ten tickets triaged by one agent, by the same agent in a for loop, and by a lead agent with three cheaper workers, printing the calls, tokens and cost of each and catching the conflict the workers create. It runs offline with scripted stand-ins for Claude; an API key swaps in the real models.

The batch and the one-agent baseline

Here’s what came in overnight. All ten are open and nobody owns them yet:

TicketTitleCustomer
KITE-146Password reset link never arrivesBlue Fern Labs
KITE-147Lost my phone, can’t get a two-factor codeNorthwind Studio
KITE-148Where do I download last month’s invoice?Blue Fern Labs
KITE-149Password reset fails for our SSO teamHarbor Pine
KITE-150Can’t see invoices since switching to single sign-onHarbor Pine
KITE-151How do we downgrade our plan?Northwind Studio
KITE-152Account locked after too many sign-in attemptsBlue Fern Labs
KITE-153CSV export drops the last row againNorthwind Studio
KITE-154Dark mode text unreadable on the settings pageHarbor Pine
KITE-155Board view freezes with 500+ cardsBlue Fern Labs

Triage here means: look the ticket up, find the help section that fixes it (or a similar past ticket, for bugs), assign an owner and send a first reply. Sign-in tickets go to priya, billing to lena, and bugs to whoever fixed a similar one, otherwise sam.

The agent is the one from What an AI Agent Actually Is: claude-opus-5 in a loop with the five Kitebase tools (search_help, get_ticket, search_tickets, assign_ticket, reply_to_customer). The simplest thing is to give it the whole batch as one task, in one conversation:

1. ONE AGENT, ONE CONVERSATION (claude-opus-5, 5 tools)
  replies sent: 10 to 10 tickets   assigned: 10 of 10
  model calls: 31 (31 in a row)   tokens: 89,825 in, 1,570 out   cost: $0.4884
  input tokens: 575 on the first call, 4,968 on the biggest

It works: every ticket gets one reply and one owner. But look at the last line. Every call re-sends the whole messages list (Plan, Act, Observe), and by ticket 10 that list holds nine finished tickets. Ticket 1’s three calls sent 2,168 input tokens. Ticket 10’s sent 14,181, for the same work.

It’s slow, too: 31 calls, each waiting for the one before. And the history is noise: the model re-reads nine closed tickets while working on the tenth, which is how agents lose the thread (State and Memory).

Fix one: a fresh conversation per ticket

The tickets don’t depend on each other. Triaging KITE-153 needs nothing from KITE-146. So there’s no reason for them to share a messages list. Run the same agent once per ticket, each time with an empty one:

def one_agent_per_ticket(client, kb, run):
    for ticket_id in BATCH:  # a for loop hands each ticket to exactly one run, by construction
        run_agent(client, run, AGENT_MODEL, AGENT_PROMPT, list(SCHEMAS.values()),
                  batch_task(kb, [ticket_id]), lambda name, args: run_tool(kb, name, args))

(Trimmed: the version in main.py also counts calls.) run_agent is the same loop as before; the only change is that each ticket starts a new conversation. The output:

2. ONE AGENT PER TICKET, A FOR LOOP (claude-opus-5, 5 tools)
  replies sent: 10 to 10 tickets   assigned: 10 of 10
  model calls: 40 (4 in a row)   tokens: 26,637 in, 1,750 out   cost: $0.1769
  input tokens: 458 on the first call, 995 on the biggest

More calls (each run needs its own “done” turn) but under a third of the tokens, because no call carries more than one ticket. “4 in a row” is the longest chain of calls that must wait for each other: the runs are independent, so a thread pool or asyncio can start them all at once. (The example runs them one after another so its output repeats.)

This is already the simplest multi-agent pattern, fan-out: the same agent run on many independent items at once. There’s no lead and no coordination, because your for loop decides who handles what, and it can’t hand one ticket to two runs.

A lead agent and workers

The design proposed at the start is the orchestrator-worker pattern, as Anthropic’s Building effective agents names it. A multi-agent system is any setup where more than one agent loop works on the same job. In this one:

  • The lead agent (the orchestrator) reads the job, splits it into pieces and hands each piece out. Here it’s claude-opus-5, and it splits the batch by area.
  • Each worker (also called a sub-agent) is a separate agent loop with its own system prompt, its own tools and its own fresh context: a messages list that starts empty and holds only its own piece of the work.
  • The workers report back, and something merges their results into one outcome.

The mechanism is less exotic than it sounds. The lead has one tool, start_worker, and the body of that tool is another agent loop:

def start_worker(name: str, args: dict) -> tuple[str, bool]:
    area, ids = args["area"], args["ticket_ids"]
    allowed = {n: SCHEMAS[n] for n in WORKER_TOOLS[area]}  # two read-only tools
    final = run_agent(client, run, WORKER_MODEL, WORKER_PROMPT, list(allowed.values()),
                      args["brief"], lambda n, a: run_tool(kb, n, a, allowed))
    found = [dict(p, worker=area) for p in json.loads(final)]
    found = [p for p in found if p["ticket_id"] in ids]  # outside its brief: not its call to make
    proposals.extend(found)
    return json.dumps(found), False

(Trimmed: main.py also counts calls and handles a worker that doesn’t return JSON.) When the lead asks for start_worker three times in one turn, your code runs three complete agent loops, and each one’s final answer becomes a tool result for the lead. The lead never sees the workers’ lookups, only their conclusions.

10 OVERNIGHT TICKETS KITE-146 to KITE-155 sign-in, billing and bug reports LEAD AGENT: CLAUDE-OPUS-5 Splits the batch by area and calls start_worker once per brief. 2 calls: $0.0185 of the run's $0.0456 SIGN-IN WORKER claude-haiku-4-5 KITE-146 KITE-147 KITE-149 KITE-150 KITE-152 tools: get_ticket, search_help 11 calls, own messages list BILLING WORKER claude-haiku-4-5 KITE-148 KITE-150 KITE-151 tools: get_ticket, search_help 7 calls, own messages list BUGS WORKER claude-haiku-4-5 KITE-153 KITE-154 KITE-155 tools: get_ticket, search_tickets 7 calls, own messages list 11 proposals for 10 tickets, as JSON. Nothing written yet. MERGE, IN YOUR CODE One proposal for a ticket: apply it. Two that disagree: hold both. APPLIED: 9 TICKETS assign_ticket and reply_to_customer, one reply each HELD: KITE-150 sign-in worker: priya, "ask your IT admin" billing worker: lena, "ask an admin to check your role" Harbor Pine gets nothing until someone picks one.
The worked run. Workers read and propose; one place in your code writes.

A narrow brief and a small tool set

A worker knows only what its brief tells it. Its messages list starts with the brief and nothing else: not the other tickets, not the lead’s reasoning, not the other workers. Here’s the brief the lead writes for the sign-in worker:

Area: sign-in. Tickets: KITE-146, KITE-147, KITE-149, KITE-150, KITE-152.
For each ticket: get_ticket, then search_help with the title, and pick the sign-in section that fixes it. Propose assignee priya and a reply based only on that section.
Return only the JSON list. Don't reply to customers or assign anything.

That covers the four things Anthropic’s engineers found each subagent needs in their multi-agent research system: an objective, an output format, which tools to use, and clear boundaries. With vague briefs, they write, agents duplicate work, leave gaps or miss what they needed.

The tool set is just as narrow. The sign-in and billing workers get get_ticket and search_help; the bugs worker gets get_ticket and search_tickets. None of them gets assign_ticket or reply_to_customer. That buys three things:

  • Fewer wrong picks. The model picks tools from their descriptions (Giving an Agent Tools); two tools are fewer ways to pick wrong.
  • Smaller calls. The five tool definitions are about 328 tokens and ride along on every call. Two are about 124.
  • No writes. A worker can look things up and propose. It can’t touch a ticket or a customer. Your code double-checks that: run_tool refuses any tool that isn’t in the worker’s allowed set, and proposals for tickets outside the brief are dropped.

That last point, workers read, one place writes, is what makes the rest of this design safe, including the next choice.

A cheaper model for the workers

The workers run on claude-haiku-4-5, at $1 per million input tokens and $5 per million output, against $5 and $25 for claude-opus-5. That’s a fair trade for this job:

  • Each worker’s task is small and fixed: two tools, one area, the same steps per ticket, a strict output format.
  • Its output is a proposal your code checks before anything reaches a customer.
  • The judgment call, deciding who gets which ticket, stays on claude-opus-5 in the lead.

For your own tasks, run your evaluation set (LLM Evaluation Pipelines) on the cheaper model first and switch only if its scores hold.

The lead still costs real money: its two calls are $0.0185, about 40% of this run’s $0.0456, for splitting ten titles into three lists. And a cheaper model isn’t a multi-agent feature. The for loop runs on Haiku too, for about $0.035. Keep the model choice and the architecture choice apart when you compare.

Duplicated work and conflicting writes

Look at the briefs again. KITE-150, “Can’t see invoices since switching to single sign-on”, is about sign-in and invoices, so the lead put it in both. A real model might well do the same. Now two workers triage one ticket, each with no idea the other exists.

That’s duplicated work: one ticket looked up and searched twice, paid for twice. Worse is what they propose. The sign-in worker reads the SSO help section and proposes priya with “ask your IT admin”. The billing worker reads the invoices section and proposes lena with “ask an admin to check your role”. Both are reasonable from what each saw. Applied in order, they’re a conflicting write: two writes to one record that disagree, where the second silently replaces the first.

SIGN-IN WORKER proposal applied 1st TICKET KITE-150 assignee: null HARBOR PINE the customer's inbox BILLING WORKER proposal applied 2nd 1. assign_ticket: priya 2. reply: "ask your IT admin" 3. assign_ticket: lena 4. reply: "ask an admin to check your role" RESULT: EVERY WRITE LANDED The ticket ends up with lena and priya's triage is silently overwritten. Harbor Pine gets two replies that disagree. Each worker was right about what it saw; neither saw the other.
What happens to KITE-150 when every proposal is applied. Set CHECK_CONFLICTS = False to see it.

Cognition’s engineers put it in one line in Don’t Build Multi-Agents: “Actions carry implicit decisions, and conflicting decisions carry bad results.” Each worker’s reply is a decision about what the ticket is really about, and they decided differently.

Because the workers only propose, there’s one place to catch it. The merge step groups proposals by ticket and refuses to guess:

def merge_and_apply(kb, proposals):
    by_ticket = defaultdict(list)
    for p in proposals:
        by_ticket[p["ticket_id"]].append(p)
    held = []
    for ticket_id, ps in by_ticket.items():
        agree = len({(p["assignee"], p["reply"]) for p in ps}) == 1
        if CHECK_CONFLICTS and not agree:
            held.extend(ps)  # a human (or the lead) picks one; the customer hears nothing yet
            continue
        for p in (ps[:1] if agree else ps):
            kb.assign_ticket(ticket_id, p["assignee"])
            kb.reply_to_customer(ticket_id, p["reply"])
    return held

Two proposals that agree are only wasted work, so it applies one. Two that disagree are held. The output:

3. LEAD + WORKERS (lead claude-opus-5, workers claude-haiku-4-5)
  worker sign-in  5 tickets  11 calls  tools: get_ticket, search_help
  worker billing  3 tickets   7 calls  tools: get_ticket, search_help
  worker bugs     3 tickets   7 calls  tools: get_ticket, search_tickets
  merge: 11 proposals for 10 tickets
  replies sent: 9 to 9 tickets   assigned: 9 of 10
  model calls: 27 (13 in a row)   tokens: 22,657 in, 1,623 out   cost: $0.0456
  input tokens: 326 on the first call, 1,766 on the biggest
  CONFLICT KITE-150: held, nothing written
    sign-in  worker: assign priya reply "Your workspace signs in with SSO, so Kitebas..."
    billing  worker: assign lena  reply "Invoices are under Admin > Billing > Invoice..."

The check is an if in your code, not a model call, on purpose: comparing two proposals is exact and free.

You can also stop the duplicate before any worker runs: with CHECK_SPLIT = True, start_worker refuses a ticket that already went to another worker, and the lead restarts the billing worker without KITE-150. But KITE-150 then gets the sign-in answer, and the billing one was probably better. A check stops double handling; deciding which area owns a ticket is still a judgment call.

The for loop never had this bug. It hands out one ticket per run, so runs can’t collide. The conflict came from letting a model draw the lines between workers.

Why not let the workers talk to each other and sort it out?

Because then they share state, with the usual problems of concurrent programs. Worker A reads “unassigned”, worker B reads “unassigned”, both write. That’s a race condition, and chat between agents adds more of them.

The usual fixes apply: give each worker its inputs up front, have it return results, and put every write in one place. If workers need a lot of each other’s context, the work shouldn’t be split.

What it cost, side by side

The last lines of python main.py:

                                     calls in a row  tokens in     cost  all on opus
1. ONE AGENT, ONE CONVERSATION          31       31     89,825  $0.4884      $0.4884
2. ONE AGENT PER TICKET, A FOR LOOP     40        4     26,637  $0.1769      $0.1769
3. LEAD + WORKERS                       27       13     22,657  $0.0456      $0.1539

“All on opus” prices every call at claude-opus-5 rates, so you can compare the designs without the model choice mixed in.

Input tokens sent on every model call, all three setups on one scale Each bar is one call. A call re-sends its whole messages list, so bar height is what that call paid for. 1. ONE CONVERSATION 31 calls, 89,825 tokens in $0.4884 575 4,968 on call 31 Ticket 10 re-sends tickets 1 to 9. 2. PER TICKET 40 calls, 26,637 tokens in $0.1769 Every ticket starts again at about 460 tokens. 3. LEAD + WORKERS 27 calls, 22,657 tokens in $0.0456 the lead reads 11 proposals sign-in worker billing bugs claude-opus-5 call ($5 per million input tokens) claude-haiku-4-5 call ($1 per million)
Each bar is what one call paid for. One long conversation is the expensive shape.

Three things stand out.

The big saving came from fresh contexts, not from more agents. Setups 2 and 3 both keep every messages list short, and both send under a third of setup 1’s input tokens. Setup 2 does it with a for loop.

Every agent re-pays its overhead on every call. That’s where multi-agent costs multiply: each agent’s system prompt and tool definitions go out again with each of its calls. In setup 2 the five tool definitions alone are about 13,000 of the 26,637 input tokens. Setup 3 sends fewer per call (two tools per worker) but adds the lead’s calls and the duplicated work on KITE-150. At the same model the two land within 15% of each other.

Here the multiplication stays small, because each ticket needs the same small amount of work however you split it. It gets large when workers explore. In Anthropic’s research system each subagent runs its own searches, and they report multi-agent runs using about 15 times the tokens of a chat, against about 4 times for a single agent (source). That pays for hard research questions. For ticket triage it wouldn’t.

The lead is on the critical path. Setup 3 needs 13 calls in a row: the lead’s first turn, then the slowest worker’s 11, then the lead again. The parallel for loop needs 4.

So for this batch the winner is setup 2: the agent you already had, run once per ticket, on Haiku if your evals allow it. The lead and workers bought nothing it doesn’t have, and brought a merge step and a bug.

When multiple agents earn their place

Default to one agent, and move down this list only when the step above fails in a way you can point to:

  1. One agent, one task. Most features stop here.
  2. Fan-out: the same agent over many independent items, split in code. A batch of tickets, 200 documents to summarise, a list of accounts to check. The test: can item B start without item A’s result? If yes, a for loop (or a thread pool) gives each one a fresh context, with no coordination to get wrong.
  3. A lead and workers, when at least one of these is true:
    • The split needs judgment. You can’t list the sub-tasks up front: “research why churn went up last quarter” becomes sub-questions only once someone reads the data. If code can do the split, let code do it.
    • Workers read a lot and return a little. Each one searches through thousands of tokens and hands back a paragraph, so the lead’s context stays small. Anthropic’s research subagents work this way.
    • Sub-tasks need genuinely different tools. Not “it feels cleaner”: a tool list per worker that’s actually smaller and doesn’t overlap.

Don’t split when the pieces share context or depend on each other. Anthropic says as much about their own system, and names most coding tasks as less parallel than research. Two workers writing the same records are the KITE-150 bug waiting to happen.

Try it yourself

The companion example runs all three setups on the overnight batch and prints the numbers above. It runs offline; set ANTHROPIC_API_KEY to make the same calls to claude-opus-5 and claude-haiku-4-5.

Download the runnable example (zip)

cd 09-multi-agent-systems
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. Set CHECK_CONFLICTS = False in main.py. Harbor Pine gets two different replies on KITE-150, and the ticket ends up with lena, because the billing proposal was applied last.
  2. Set CHECK_SPLIT = True. The lead gets an error for KITE-150, restarts the billing worker without it, and every ticket gets one reply. It costs one more lead call ($0.0530 instead of $0.0456).
  3. Set AGENT_MODEL = "claude-haiku-4-5". The plain for loop drops to about $0.035, cheaper than the lead and workers, with no merge step and nothing to conflict.

pip install pytest && pytest -q runs the offline tests. They need no keys and no network.

Common beginner mistakes

  • Reaching for a team of agents first. It looks like an org chart, so it feels right. Build one agent, measure it, and split only when you can name what the split buys.
  • Letting workers write. Two workers with reply_to_customer will eventually both reply. Give workers read tools, have them return proposals, and write in one place.
  • Vague briefs. “Handle the sign-in tickets” leaves the worker guessing the steps, the owner and the output. Say the objective, the tools, the output format and what not to do.
  • Comparing against a strawman. Beating one agent with ten tickets in one conversation proves little. Compare with the same agent in a loop.
  • No limits on the workers. Each worker is a full agent loop and needs its own turn limit and timeout (Retries, Timeouts, and Failure Handling). A stuck worker burns money inside a tool call the lead is waiting on.

Questions you will face in production

“The framework makes multi-agent easy. Why not just use it?” Easy to wire up isn’t cheap to run or easy to debug. The framework hides the coordination, not its cost. Decide you need several agents from numbers like the ones above, then use the framework if it saves you code.

“How do I debug a multi-agent run that gave a wrong answer?” Log every brief and every worker result, tagged with the run and the worker. The bug is usually at a hand-off: a brief that left something out, or a result the lead misread. Observability and Cost Control for Agents covers tracing a run.

Check your understanding

A teammate shows that lead + workers cost $0.05 against $0.49 for one agent, and wants to ship it. What do you ask?

What the one agent was doing. If it had all ten tickets in one conversation, most of that $0.49 is re-sent history; compare with the same agent run once per ticket. Then ask how much of the saving is the cheaper worker model, which one agent could use too.

A worker returns a proposal for KITE-160, which wasn't in its brief. What should your code do, and why?

Drop it and log it. Either another worker owns that ticket, so you’d get two proposals for it, or nobody does, and a model acted outside its job. That’s why start_worker keeps only proposals for the brief’s tickets.

Two workers both proposed "assign to priya" with the same reply for KITE-150. Is that a conflict?

No, it’s duplicated work: you paid twice, but applying one proposal gives the right result, which is why the merge compares assignee and reply. It still means the split overlapped.

You're building an agent that refactors one module across 30 files, and someone suggests one worker per file. What's the risk?

The files depend on each other: a rename in one breaks imports in the others. Each worker decides about shared names without seeing the others’ decisions, which is the KITE-150 conflict in code. Keep it in one agent.

What to remember

  • Start with one agent. Split only when you can name what the split buys: time, context size or different tools.
  • For many independent items, run the same agent per item with a fresh context. A for loop can’t hand one item to two runs.
  • A worker is a tool whose body is another agent loop. It sees only its brief, so the brief needs the objective, tools, output format and boundaries.
  • Workers read and propose; one place in your code writes and checks for conflicts first.
  • A cheaper worker model suits narrow, checked, read-only work. Prove it with evals, and compare designs at the same model.

What to study next

That’s the end of the AI Agents topic. The natural next step is agents you already use every day: How AI Coding Tools Actually Work takes the loop from What an AI Agent Actually Is and shows what changes when the tools are your file system, your terminal and your test suite.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.