Multi-Agent Systems: When and When Not
Ten support tickets came into Kitebase overnight, and someone wants them triaged before the support engineers log on. The proposal: a lead agent that hands the batch to specialist agents, one for sign-in problems, one for billing, one for bugs. It sounds like how a real support team works.
Before building it, run the numbers. This article triages the same ten tickets three ways and compares what each costs, how long it takes and what goes wrong. The answer it lands on is the one to start from: use one agent until you can show it isn’t enough.
What you’ll build: ten tickets triaged by one agent, by the same agent in a for loop, and by a lead agent with three cheaper workers, printing the calls, tokens and cost of each and catching the conflict the workers create. It runs offline with scripted stand-ins for Claude; an API key swaps in the real models.
The batch and the one-agent baseline
Here’s what came in overnight. All ten are open and nobody owns them yet:
| Ticket | Title | Customer |
|---|---|---|
| KITE-146 | Password reset link never arrives | Blue Fern Labs |
| KITE-147 | Lost my phone, can’t get a two-factor code | Northwind Studio |
| KITE-148 | Where do I download last month’s invoice? | Blue Fern Labs |
| KITE-149 | Password reset fails for our SSO team | Harbor Pine |
| KITE-150 | Can’t see invoices since switching to single sign-on | Harbor Pine |
| KITE-151 | How do we downgrade our plan? | Northwind Studio |
| KITE-152 | Account locked after too many sign-in attempts | Blue Fern Labs |
| KITE-153 | CSV export drops the last row again | Northwind Studio |
| KITE-154 | Dark mode text unreadable on the settings page | Harbor Pine |
| KITE-155 | Board view freezes with 500+ cards | Blue Fern Labs |
Triage here means: look the ticket up, find the help section that fixes it (or a similar past ticket, for bugs), assign an owner and send a first reply. Sign-in tickets go to priya, billing to lena, and bugs to whoever fixed a similar one, otherwise sam.
The agent is the one from What an AI Agent Actually Is: claude-opus-5 in a loop with the five Kitebase tools (search_help, get_ticket, search_tickets, assign_ticket, reply_to_customer). The simplest thing is to give it the whole batch as one task, in one conversation:
1. ONE AGENT, ONE CONVERSATION (claude-opus-5, 5 tools)
replies sent: 10 to 10 tickets assigned: 10 of 10
model calls: 31 (31 in a row) tokens: 89,825 in, 1,570 out cost: $0.4884
input tokens: 575 on the first call, 4,968 on the biggest
It works: every ticket gets one reply and one owner. But look at the last line. Every call re-sends the whole messages list (Plan, Act, Observe), and by ticket 10 that list holds nine finished tickets. Ticket 1’s three calls sent 2,168 input tokens. Ticket 10’s sent 14,181, for the same work.
It’s slow, too: 31 calls, each waiting for the one before. And the history is noise: the model re-reads nine closed tickets while working on the tenth, which is how agents lose the thread (State and Memory).
Fix one: a fresh conversation per ticket
The tickets don’t depend on each other. Triaging KITE-153 needs nothing from KITE-146. So there’s no reason for them to share a messages list. Run the same agent once per ticket, each time with an empty one:
def one_agent_per_ticket(client, kb, run):
for ticket_id in BATCH: # a for loop hands each ticket to exactly one run, by construction
run_agent(client, run, AGENT_MODEL, AGENT_PROMPT, list(SCHEMAS.values()),
batch_task(kb, [ticket_id]), lambda name, args: run_tool(kb, name, args))
(Trimmed: the version in main.py also counts calls.) run_agent is the same loop as before; the only change is that each ticket starts a new conversation. The output:
2. ONE AGENT PER TICKET, A FOR LOOP (claude-opus-5, 5 tools)
replies sent: 10 to 10 tickets assigned: 10 of 10
model calls: 40 (4 in a row) tokens: 26,637 in, 1,750 out cost: $0.1769
input tokens: 458 on the first call, 995 on the biggest
More calls (each run needs its own “done” turn) but under a third of the tokens, because no call carries more than one ticket. “4 in a row” is the longest chain of calls that must wait for each other: the runs are independent, so a thread pool or asyncio can start them all at once. (The example runs them one after another so its output repeats.)
This is already the simplest multi-agent pattern, fan-out: the same agent run on many independent items at once. There’s no lead and no coordination, because your for loop decides who handles what, and it can’t hand one ticket to two runs.
A lead agent and workers
The design proposed at the start is the orchestrator-worker pattern, as Anthropic’s Building effective agents names it. A multi-agent system is any setup where more than one agent loop works on the same job. In this one:
- The lead agent (the orchestrator) reads the job, splits it into pieces and hands each piece out. Here it’s
claude-opus-5, and it splits the batch by area. - Each worker (also called a sub-agent) is a separate agent loop with its own system prompt, its own tools and its own fresh context: a messages list that starts empty and holds only its own piece of the work.
- The workers report back, and something merges their results into one outcome.
The mechanism is less exotic than it sounds. The lead has one tool, start_worker, and the body of that tool is another agent loop:
def start_worker(name: str, args: dict) -> tuple[str, bool]:
area, ids = args["area"], args["ticket_ids"]
allowed = {n: SCHEMAS[n] for n in WORKER_TOOLS[area]} # two read-only tools
final = run_agent(client, run, WORKER_MODEL, WORKER_PROMPT, list(allowed.values()),
args["brief"], lambda n, a: run_tool(kb, n, a, allowed))
found = [dict(p, worker=area) for p in json.loads(final)]
found = [p for p in found if p["ticket_id"] in ids] # outside its brief: not its call to make
proposals.extend(found)
return json.dumps(found), False
(Trimmed: main.py also counts calls and handles a worker that doesn’t return JSON.) When the lead asks for start_worker three times in one turn, your code runs three complete agent loops, and each one’s final answer becomes a tool result for the lead. The lead never sees the workers’ lookups, only their conclusions.
A narrow brief and a small tool set
A worker knows only what its brief tells it. Its messages list starts with the brief and nothing else: not the other tickets, not the lead’s reasoning, not the other workers. Here’s the brief the lead writes for the sign-in worker:
Area: sign-in. Tickets: KITE-146, KITE-147, KITE-149, KITE-150, KITE-152.
For each ticket: get_ticket, then search_help with the title, and pick the sign-in section that fixes it. Propose assignee priya and a reply based only on that section.
Return only the JSON list. Don't reply to customers or assign anything.
That covers the four things Anthropic’s engineers found each subagent needs in their multi-agent research system: an objective, an output format, which tools to use, and clear boundaries. With vague briefs, they write, agents duplicate work, leave gaps or miss what they needed.
The tool set is just as narrow. The sign-in and billing workers get get_ticket and search_help; the bugs worker gets get_ticket and search_tickets. None of them gets assign_ticket or reply_to_customer. That buys three things:
- Fewer wrong picks. The model picks tools from their descriptions (Giving an Agent Tools); two tools are fewer ways to pick wrong.
- Smaller calls. The five tool definitions are about 328 tokens and ride along on every call. Two are about 124.
- No writes. A worker can look things up and propose. It can’t touch a ticket or a customer. Your code double-checks that:
run_toolrefuses any tool that isn’t in the worker’sallowedset, and proposals for tickets outside the brief are dropped.
That last point, workers read, one place writes, is what makes the rest of this design safe, including the next choice.
A cheaper model for the workers
The workers run on claude-haiku-4-5, at $1 per million input tokens and $5 per million output, against $5 and $25 for claude-opus-5. That’s a fair trade for this job:
- Each worker’s task is small and fixed: two tools, one area, the same steps per ticket, a strict output format.
- Its output is a proposal your code checks before anything reaches a customer.
- The judgment call, deciding who gets which ticket, stays on
claude-opus-5in the lead.
For your own tasks, run your evaluation set (LLM Evaluation Pipelines) on the cheaper model first and switch only if its scores hold.
The lead still costs real money: its two calls are $0.0185, about 40% of this run’s $0.0456, for splitting ten titles into three lists. And a cheaper model isn’t a multi-agent feature. The for loop runs on Haiku too, for about $0.035. Keep the model choice and the architecture choice apart when you compare.
Duplicated work and conflicting writes
Look at the briefs again. KITE-150, “Can’t see invoices since switching to single sign-on”, is about sign-in and invoices, so the lead put it in both. A real model might well do the same. Now two workers triage one ticket, each with no idea the other exists.
That’s duplicated work: one ticket looked up and searched twice, paid for twice. Worse is what they propose. The sign-in worker reads the SSO help section and proposes priya with “ask your IT admin”. The billing worker reads the invoices section and proposes lena with “ask an admin to check your role”. Both are reasonable from what each saw. Applied in order, they’re a conflicting write: two writes to one record that disagree, where the second silently replaces the first.
Cognition’s engineers put it in one line in Don’t Build Multi-Agents: “Actions carry implicit decisions, and conflicting decisions carry bad results.” Each worker’s reply is a decision about what the ticket is really about, and they decided differently.
Because the workers only propose, there’s one place to catch it. The merge step groups proposals by ticket and refuses to guess:
def merge_and_apply(kb, proposals):
by_ticket = defaultdict(list)
for p in proposals:
by_ticket[p["ticket_id"]].append(p)
held = []
for ticket_id, ps in by_ticket.items():
agree = len({(p["assignee"], p["reply"]) for p in ps}) == 1
if CHECK_CONFLICTS and not agree:
held.extend(ps) # a human (or the lead) picks one; the customer hears nothing yet
continue
for p in (ps[:1] if agree else ps):
kb.assign_ticket(ticket_id, p["assignee"])
kb.reply_to_customer(ticket_id, p["reply"])
return held
Two proposals that agree are only wasted work, so it applies one. Two that disagree are held. The output:
3. LEAD + WORKERS (lead claude-opus-5, workers claude-haiku-4-5)
worker sign-in 5 tickets 11 calls tools: get_ticket, search_help
worker billing 3 tickets 7 calls tools: get_ticket, search_help
worker bugs 3 tickets 7 calls tools: get_ticket, search_tickets
merge: 11 proposals for 10 tickets
replies sent: 9 to 9 tickets assigned: 9 of 10
model calls: 27 (13 in a row) tokens: 22,657 in, 1,623 out cost: $0.0456
input tokens: 326 on the first call, 1,766 on the biggest
CONFLICT KITE-150: held, nothing written
sign-in worker: assign priya reply "Your workspace signs in with SSO, so Kitebas..."
billing worker: assign lena reply "Invoices are under Admin > Billing > Invoice..."
The check is an if in your code, not a model call, on purpose: comparing two proposals is exact and free.
You can also stop the duplicate before any worker runs: with CHECK_SPLIT = True, start_worker refuses a ticket that already went to another worker, and the lead restarts the billing worker without KITE-150. But KITE-150 then gets the sign-in answer, and the billing one was probably better. A check stops double handling; deciding which area owns a ticket is still a judgment call.
The for loop never had this bug. It hands out one ticket per run, so runs can’t collide. The conflict came from letting a model draw the lines between workers.
Why not let the workers talk to each other and sort it out?
Because then they share state, with the usual problems of concurrent programs. Worker A reads “unassigned”, worker B reads “unassigned”, both write. That’s a race condition, and chat between agents adds more of them.
The usual fixes apply: give each worker its inputs up front, have it return results, and put every write in one place. If workers need a lot of each other’s context, the work shouldn’t be split.
What it cost, side by side
The last lines of python main.py:
calls in a row tokens in cost all on opus
1. ONE AGENT, ONE CONVERSATION 31 31 89,825 $0.4884 $0.4884
2. ONE AGENT PER TICKET, A FOR LOOP 40 4 26,637 $0.1769 $0.1769
3. LEAD + WORKERS 27 13 22,657 $0.0456 $0.1539
“All on opus” prices every call at claude-opus-5 rates, so you can compare the designs without the model choice mixed in.
Three things stand out.
The big saving came from fresh contexts, not from more agents. Setups 2 and 3 both keep every messages list short, and both send under a third of setup 1’s input tokens. Setup 2 does it with a for loop.
Every agent re-pays its overhead on every call. That’s where multi-agent costs multiply: each agent’s system prompt and tool definitions go out again with each of its calls. In setup 2 the five tool definitions alone are about 13,000 of the 26,637 input tokens. Setup 3 sends fewer per call (two tools per worker) but adds the lead’s calls and the duplicated work on KITE-150. At the same model the two land within 15% of each other.
Here the multiplication stays small, because each ticket needs the same small amount of work however you split it. It gets large when workers explore. In Anthropic’s research system each subagent runs its own searches, and they report multi-agent runs using about 15 times the tokens of a chat, against about 4 times for a single agent (source). That pays for hard research questions. For ticket triage it wouldn’t.
The lead is on the critical path. Setup 3 needs 13 calls in a row: the lead’s first turn, then the slowest worker’s 11, then the lead again. The parallel for loop needs 4.
So for this batch the winner is setup 2: the agent you already had, run once per ticket, on Haiku if your evals allow it. The lead and workers bought nothing it doesn’t have, and brought a merge step and a bug.
When multiple agents earn their place
Default to one agent, and move down this list only when the step above fails in a way you can point to:
- One agent, one task. Most features stop here.
- Fan-out: the same agent over many independent items, split in code. A batch of tickets, 200 documents to summarise, a list of accounts to check. The test: can item B start without item A’s result? If yes, a
forloop (or a thread pool) gives each one a fresh context, with no coordination to get wrong. - A lead and workers, when at least one of these is true:
- The split needs judgment. You can’t list the sub-tasks up front: “research why churn went up last quarter” becomes sub-questions only once someone reads the data. If code can do the split, let code do it.
- Workers read a lot and return a little. Each one searches through thousands of tokens and hands back a paragraph, so the lead’s context stays small. Anthropic’s research subagents work this way.
- Sub-tasks need genuinely different tools. Not “it feels cleaner”: a tool list per worker that’s actually smaller and doesn’t overlap.
Don’t split when the pieces share context or depend on each other. Anthropic says as much about their own system, and names most coding tasks as less parallel than research. Two workers writing the same records are the KITE-150 bug waiting to happen.
Try it yourself
The companion example runs all three setups on the overnight batch and prints the numbers above. It runs offline; set ANTHROPIC_API_KEY to make the same calls to claude-opus-5 and claude-haiku-4-5.
Download the runnable example (zip)
cd 09-multi-agent-systems
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
- Set
CHECK_CONFLICTS = Falseinmain.py. Harbor Pine gets two different replies on KITE-150, and the ticket ends up with lena, because the billing proposal was applied last. - Set
CHECK_SPLIT = True. The lead gets an error for KITE-150, restarts the billing worker without it, and every ticket gets one reply. It costs one more lead call ($0.0530 instead of $0.0456). - Set
AGENT_MODEL = "claude-haiku-4-5". The plainforloop drops to about $0.035, cheaper than the lead and workers, with no merge step and nothing to conflict.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- Reaching for a team of agents first. It looks like an org chart, so it feels right. Build one agent, measure it, and split only when you can name what the split buys.
- Letting workers write. Two workers with
reply_to_customerwill eventually both reply. Give workers read tools, have them return proposals, and write in one place. - Vague briefs. “Handle the sign-in tickets” leaves the worker guessing the steps, the owner and the output. Say the objective, the tools, the output format and what not to do.
- Comparing against a strawman. Beating one agent with ten tickets in one conversation proves little. Compare with the same agent in a loop.
- No limits on the workers. Each worker is a full agent loop and needs its own turn limit and timeout (Retries, Timeouts, and Failure Handling). A stuck worker burns money inside a tool call the lead is waiting on.
Questions you will face in production
“The framework makes multi-agent easy. Why not just use it?” Easy to wire up isn’t cheap to run or easy to debug. The framework hides the coordination, not its cost. Decide you need several agents from numbers like the ones above, then use the framework if it saves you code.
“How do I debug a multi-agent run that gave a wrong answer?” Log every brief and every worker result, tagged with the run and the worker. The bug is usually at a hand-off: a brief that left something out, or a result the lead misread. Observability and Cost Control for Agents covers tracing a run.
Check your understanding
A teammate shows that lead + workers cost $0.05 against $0.49 for one agent, and wants to ship it. What do you ask?
What the one agent was doing. If it had all ten tickets in one conversation, most of that $0.49 is re-sent history; compare with the same agent run once per ticket. Then ask how much of the saving is the cheaper worker model, which one agent could use too.
A worker returns a proposal for KITE-160, which wasn't in its brief. What should your code do, and why?
Drop it and log it. Either another worker owns that ticket, so you’d get two proposals for it, or nobody does, and a model acted outside its job. That’s why start_worker keeps only proposals for the brief’s tickets.
Two workers both proposed "assign to priya" with the same reply for KITE-150. Is that a conflict?
No, it’s duplicated work: you paid twice, but applying one proposal gives the right result, which is why the merge compares assignee and reply. It still means the split overlapped.
You're building an agent that refactors one module across 30 files, and someone suggests one worker per file. What's the risk?
The files depend on each other: a rename in one breaks imports in the others. Each worker decides about shared names without seeing the others’ decisions, which is the KITE-150 conflict in code. Keep it in one agent.
What to remember
- Start with one agent. Split only when you can name what the split buys: time, context size or different tools.
- For many independent items, run the same agent per item with a fresh context. A
forloop can’t hand one item to two runs. - A worker is a tool whose body is another agent loop. It sees only its brief, so the brief needs the objective, tools, output format and boundaries.
- Workers read and propose; one place in your code writes and checks for conflicts first.
- A cheaper worker model suits narrow, checked, read-only work. Prove it with evals, and compare designs at the same model.
What to study next
That’s the end of the AI Agents topic. The natural next step is agents you already use every day: How AI Coding Tools Actually Work takes the loop from What an AI Agent Actually Is and shows what changes when the tools are your file system, your terminal and your test suite.
Further reading
- Anthropic: Building effective agents. Names the orchestrator-workers and parallelization patterns and argues for the simplest design that works. Start here.
- Anthropic: How we built our multi-agent research system. An engineering account of where a lead and subagents paid off, what the briefs needed and what the token bill looked like.
- Cognition: Don’t Build Multi-Agents. The case against splitting, from the team behind a coding agent: why shared context matters and how parallel decisions conflict.
- OpenAI: A practical guide to building agents. Vendor-neutral advice on when to move from one agent to several.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.