Observability and Cost Control for Agents

The Kitebase support agent from article 01 now works tickets on its own: it reads a ticket, searches the help center, and replies to the customer or hands the ticket to a person. One morning it picks up KITE-140, “Invoice shows the wrong VAT number”. The help center has nothing on VAT numbers, so the agent searches, rephrases, and searches again. Nothing crashes. Every model call returns 200 OK at a normal price. The run as a whole costs more than five tickets that went well, and the customer never gets a reply.

Article 11 of AI Engineering showed how to trace a single model call, and article 12 how to price one. Both still apply. What changes with an agent is that one ticket is many model calls, the number isn’t fixed, and each one costs more than the one before.

What you’ll build: the Kitebase agent loop, instrumented with one trace per run and a per-run budget of turns, tokens and dollars. A scripted fake model runs it over four tickets, one of which loops, and a small report prints cost per run and per resolved ticket and points at the loop. It runs offline.

One run, one trace

With article 11’s logging, the four tickets in this example produce 26 model-call records, each fine on its own. None says which ticket it belonged to or what the tools returned, so the loop on KITE-140 is invisible: twelve ordinary-looking calls among fourteen others.

The unit to trace is the run: one pass of the agent loop from the task (“Handle ticket KITE-142”) to the end, whether that’s a final answer or a stop. Each run gets one trace, the record of one piece of work, made of spans, one per step. Every span carries the run’s trace_id and points at the span it ran inside with parent_id, like a call stack written down. For an agent:

  • A root span, agent.run, for the whole run: the ticket, the outcome, the totals.
  • A model.turn span for each model call: tokens in and out, the cost of that call, the run’s cost so far, and stop_reason (why the model stopped; "tool_use" means it wants a tool, "end_turn" that it’s done).
  • A tool.call span for each tool the model asked for: the name, the arguments, and what came back.

Here’s the trace for KITE-142, “Customer locked out after SSO change” (python report.py --trace run-01 prints it as text):

Trace run-01: KITE-142, "Customer locked out after SSO change" span what it recorded cost of this span run total AGENT.RUN ticket KITE-142, outcome resolved, 4 turns, 3 tool calls $0.0196 MODEL.TURN 1 443 in, 39 out, tool_use $0.0032 $0.0032 TOOL.CALL get_ticket("KITE-142"), 145 chars no tokens MODEL.TURN 2 562 in, 44 out, tool_use $0.0039 $0.0071 TOOL.CALL search_help("locked out sso"), 762 chars no tokens MODEL.TURN 3 836 in, 108 out, tool_use $0.0069 $0.0140 TOOL.CALL reply_to_customer("KITE-142", 41 words) no tokens MODEL.TURN 4 994 in, 26 out, end_turn $0.0056 $0.0196 Every span below the root has trace_id run-01 and parent_id run-01.1. Tool calls cost no tokens themselves. Their results do, on every turn after them.
One run, read top to bottom. Each model turn costs more than the last, though it does about the same work.

The loop is the one from article 02 with a span around each model call, using article 11’s Trace class: a context manager that writes the span when the block ends, even if it raised. Trimmed from main.py:

with trace.span("model.turn", turn=turn) as span:
    response = client.messages.create(model=MODEL, max_tokens=1024, system=SYSTEM_PROMPT,
                                      tools=TOOLS, messages=messages)
    u = response.usage
    cost = cost_usd(MODEL, u.input_tokens, u.output_tokens)
    tokens, usd = tokens + u.input_tokens + u.output_tokens, usd + cost
    span.update(input_tokens=u.input_tokens, output_tokens=u.output_tokens,
                cost_usd=round(cost, 6), run_cost_usd=round(usd, 6),
                stop_reason=response.stop_reason)

cost_usd multiplies tokens by the price table: $5 per million input tokens and $25 per million output for claude-opus-5. run_cost_usd is the running total at that turn. Tool calls get their own spans:

for block in (b for b in response.content if b.type == "tool_use"):
    with trace.span("tool.call", turn=turn, tool=block.name, args=block.input) as span:
        output, is_error = run_tool(kb, block.name, block.input)
        # The model acts on the raw result, so keep enough of it to see what it saw.
        span.update(is_error=is_error, result_chars=len(output), result=output[:200])

The gotcha is the tool result. When an agent does something odd, the reason is usually in what a tool returned: an empty list, an error it read as data, a search result that looked relevant and wasn’t. Log only the model’s choices and you can’t tell why it made them, so keep the first few hundred characters and the full size. Tool results hold customer data (this one names Northwind Studio, and a real ticket body holds whatever the customer wrote), so they follow article 11’s rules for question and answer text: redact, and keep them for days, not months.

Is there a standard for agent spans, or do I make up names?

There’s one, still marked experimental. OpenTelemetry’s generative AI conventions (the standard names for LLM telemetry that Datadog, Honeycomb, Langfuse and others read) define an invoke_agent operation for a whole agent run, chat for a model call and execute_tool for a tool call, plus attributes like gen_ai.usage.input_tokens and gen_ai.tool.name.

The example uses shorter names. In production, use the standard ones, so hosted tools draw your runs without setup.

Why an agent’s cost is per run, not per call

Turn 1 of KITE-142 cost $0.0032. The run took four turns and cost $0.0196: six times as much, not four. Each turn does about the same work, so why does each cost more?

Because the Messages API is stateless: it keeps nothing between calls. To let the model “remember” that it already read the ticket, the loop resends the whole conversation every turn: system prompt, five tool definitions, the task, every earlier reply and every tool result. The new part of each request is small. The resent part grows.

Input tokens sent on each model turn, run-01 (KITE-142) resent: already sent on an earlier turn new since the last turn turn 1 443: system, 5 tools, task 443 in turn 2 443 +119 562 in turn 3 562 +274: search result 836 in turn 4 836 +158 994 in WHAT YOU PAY FOR 443 + 562 + 836 + 994 = 2,835 in 2,835 × $5 / 1M = $0.0142 The conversation at the end is 994 tokens. You paid for 2,835: 1,841 of them (65%) were resent. The first 443 went four times. A longer run resends more: run-04 is 85%.
Each turn's request starts with the whole previous request. You pay for the prefix again every time.

The conversation at the end is 994 tokens, and you paid for 2,835 input tokens to get there. If every turn adds about the same amount, turn 20 sends 20 turns’ worth of history, so a run’s total grows roughly with the square of its turns. The KITE-140 script, stopped after different numbers of turns:

12 turns    19,942 tokens   $0.1099
20 turns    51,681 tokens   $0.2754
50 turns   301,638 tokens   $1.5509

Going from 20 turns to 50, two and a half times the turns, costs 5.6 times as much.

So price the run, not the call. The average, the worst case and the budget are all per run, summed from its model.turn spans. Per-call numbers can’t tell you what a ticket cost.

The same fact is the biggest saving. Most of each turn’s input is a prefix the API has already seen, which is what prompt caching discounts: reading a cached prefix costs a tenth of the input price, $0.50 per million instead of $5 on claude-opus-5. Article 10 sets it up; for a growing conversation like an agent loop, put the cache breakpoint on the last message each turn. The fake model here doesn’t simulate caching, so these numbers are the uncached worst case.

Cost per resolved ticket

The report adds up every run:

$ python report.py
run     ticket    turns tools  tokens     cost  resent  outcome
run-01  KITE-142      4     3   3,052  $0.0196    65%  resolved
run-02  KITE-144      6     5   5,759  $0.0342    75%  resolved
run-03  KITE-143      4     3   2,790  $0.0171    67%  handed_off
run-04  KITE-140     12    12  19,942  $0.1099    85%  stopped

4 runs, $0.1808 in total, 2 tickets resolved.
Cost per run:             $0.0452
Cost per resolved ticket: $0.0904
Stopped runs: 1, $0.1099 (61% of the spend, and nothing to show for it)

Cost per run, $0.0452, is easy to compute and hides the problem. Cost per resolved ticket is total spend divided by tickets the agent actually closed out: $0.0904, because runs that resolved nothing still have to be paid for by the ones that did. KITE-143 was handed to a person, correctly, for $0.0171. KITE-140 spent $0.1099 and produced nothing. Together they double what a resolved ticket costs.

It’s the number that moves when something goes wrong: loops, handoffs and wasted turns all push it up while each call looks normal. It’s the agent version of article 12’s cost per resolved conversation, and it judges changes the same way: a cheaper model that resolves fewer tickets can cost more per resolved ticket.

The gotcha is the word “resolved”. The example counts a run as resolved when the agent sent a reply and finished, but a reply that doesn’t fix the problem isn’t a resolution. In production, take it from the ticket later: closed, and not reopened within, say, seven days. Store the trace id on the ticket so you can join the two.

A per-run budget that stops a runaway

Back to KITE-140. The loop from article 02 has a turn limit, so it can’t run forever, but with max_turns = 20 it stops at $0.2754, fourteen times a normal run. A turn limit bounds how many calls, not how much they cost.

A budget is a per-run limit on what the run may spend, checked before every model call: a timeout measured in tokens and dollars. The example’s is a small dataclass:

@dataclass
class Budget:
    max_turns: int = 20
    max_tokens: int = 40_000  # input + output, summed over the run
    max_usd: float = 0.10

    def exceeded(self, tokens: int, usd: float) -> str | None:
        if tokens >= self.max_tokens:
            return f"token budget ({tokens:,} of {self.max_tokens:,})"
        if usd >= self.max_usd:
            return f"dollar budget (${usd:.4f} of ${self.max_usd:.2f})"
        return None

And the check, at the top of each turn, before the model is called:

for turn in range(1, budget.max_turns + 1):
    if reason := budget.exceeded(tokens, usd):
        run.update(outcome="stopped", stopped_by=reason)
        break
    ...
else:
    run.update(outcome="stopped", stopped_by=f"turn limit ({budget.max_turns})")

python main.py runs all four tickets with this budget:

run-01  KITE-142   4 turns   3,052 tokens  $0.0196  resolved
run-02  KITE-144   6 turns   5,759 tokens  $0.0342  resolved
run-03  KITE-143   4 turns   2,790 tokens  $0.0171  handed_off
run-04  KITE-140  12 turns  19,942 tokens  $0.1099  stopped  dollar budget ($0.1099 of $0.10)
run-04, KITE-140 "Invoice shows the wrong VAT number": run total after each turn $0 $0.20 1 12 20 model turn $0.10 budget: max_usd = 0.10 stopped: $0.1099 turn limit only: $0.2754 WHAT THE TRACE SHOWS 1 get_ticket KITE-140 2 "invoice vat number" 3 "change vat number on…" 4 "billing vat" 5 "vat id invoices" 6 "invoice vat number" 7 "invoice vat number" … the same 4, again 11 of 12 calls: search_help, one query 5 times, no reply, no handoff. A normal run costs $0.0196.
The budget stops the damage at a known cost. The trace beside it says why the run needed stopping.

Three details matter.

It’s checked before the call, so it overshoots by one turn. The run had spent $0.0945 before turn 12, under the limit, so turn 12 ran and took it to $0.1099. If a tight cap matters, estimate the next call first: you know exactly what you’re about to send, and client.messages.count_tokens counts it without running it.

Dollars and tokens catch different things. Dollars are what you care about. Tokens don’t depend on the price table, and they also cap latency and how close the run gets to the model’s context limit. Switch MODEL to claude-haiku-4-5, five times cheaper per token, and the runaway never reaches $0.10; the token budget stops it at 18 turns and $0.0454.

A stop is an outcome, not a crash. The customer is still waiting, so decide in code what happens next: hand the ticket to a person, with the trace id in the note (article 06 covers handing work to people). Don’t retry the run automatically; the same ticket and prompt will most likely loop the same way.

To pick the numbers, start with a dollar budget about five times a typical resolved run ($0.10 against run-01’s $0.0196 here). After a week of traces, move it to just above the most expensive run that resolved its ticket.

Why not tell the model its budget and let it stop itself?

It doesn’t know what it’s spending: it sees the conversation, not your token counts or prices. Adding “You have 3 turns left” to the last few requests can help it wrap up with a handoff instead of being cut off.

That’s a courtesy, not the limit. A confused model can ignore it like any other instruction, so the hard stop stays in your code.

Spotting loops and wasted turns

The budget limits the damage but doesn’t say what went wrong. python report.py --trace run-04 does: one get_ticket, then eleven search_help calls cycling through four phrasings, “invoice vat number” five times. The system prompt says to assign the ticket to sam when the help center doesn’t cover it. The model never did, because every search returned something: the word-matching search always finds a section that mentions invoices, so each result looked close enough to try once more.

These signals find runs like this without reading every trace:

  • The same tool with the same arguments. A read-only tool gives the same answer twice, so the repeat teaches the model nothing.
  • Far more turns than the task usually takes. KITE-142 took 4; twelve on a similar ticket is suspicious.
  • One tool dominating the run. 11 of 12 calls to search_help means it’s stuck on one step.
  • The same tool error over and over. The model is retrying something your code should handle (article 05 covers which failures to retry in code).

The first one is a few lines over the tool spans:

def repeated_calls(spans: list[dict]) -> dict[str, int]:
    seen = Counter(f"{s['tool']}({json.dumps(s['args'], sort_keys=True)})" for s in tool_calls(spans))
    return {call: n for call, n in seen.items() if n > 1}

The report runs it on every run:

Repeated tool calls (same tool, same arguments):
  run-02  2x  get_ticket({"ticket_id": "KITE-144"})
  run-02  2x  search_help({"query": "sso reset password"})
  run-04  5x  search_help({"query": "invoice vat number"})
  run-04  2x  search_help({"query": "change vat number on invoice"})
  ...

run-02 is the quieter case. KITE-144 was resolved, but the agent ran the same search twice and read the ticket twice. Those two turns cost $0.0117, and their results rode along in every later turn, so the run cost $0.0342 against run-01’s $0.0196 for the same kind of ticket. The outcome looks fine; only the trace shows it.

For run-04, make search_help say “No matches.” when nothing scores well (the minimum score from the RAG article), so the model gets a clear signal to hand off. Or catch the loop in code: on the third identical call, return a tool result saying “You’ve run this search 3 times with the same result. The help center doesn’t cover this; assign the ticket.” Add the ticket to your evaluation set too (article 08), so a later prompt change can’t bring the loop back.

Alerts worth having

Article 11’s rule still holds: alert on rates and per-unit costs, not totals, and page someone only for things that need fixing now. For an agent the units are runs, so the report checks three things:

ALERT run-04 KITE-140: stopped by dollar budget ($0.1099 of $0.10) after 12 turns, 11 of 12 tool calls were search_help.
ALERT 25% of runs hit a limit (threshold 5%).
ALERT cost per resolved ticket $0.0904 is above the $0.05 target.
  • Each budget stop, as a message to a channel with the trace id, not a page. An occasional stuck run is expected; someone should read it.
  • The share of runs hitting a limit over the last hour. Above about 5%, something changed: a tool started failing, a prompt shipped, or a new kind of ticket arrived. That’s worth a page, since every stuck run costs up to the budget and helps no one.
  • Cost per resolved ticket, daily, against a target or the week before (article 11’s 30% jump works here). It catches the slow drift: more wasted turns, more handoffs, a longer prompt.

Put turns per resolved run and each tool’s error rate on the daily dashboard too. A creeping turn count shows wasted turns before the bill does.

Try it yourself

The example is the agent loop with its trace and budget, a scripted fake model for four tickets, and the report. It runs offline. With ANTHROPIC_API_KEY set, python main.py --live KITE-142 runs one ticket against claude-opus-5 with the same tools, trace and budget.

Download the runnable example (zip)

cd 08-observability-and-cost
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py                   # four runs, spans to logs/spans.jsonl
python report.py                 # cost per run and per resolved ticket
python report.py --trace run-04  # the runaway, span by span

The fake counts tokens as a quarter of the characters the loop sends, so input grows each turn exactly as the history does. Real counts come out higher (the API adds its own tool-use instructions) but grow the same way.

Then try these:

  1. Set max_usd in Budget to 1.00 and run both scripts again. The token budget stops the runaway instead, at 18 turns and $0.2270. Set max_tokens to 1_000_000 too, and only the turn limit is left: 20 turns, $0.2754.
  2. Change MODEL in main.py to "claude-haiku-4-5". Every run is five times cheaper, the dollar budget never trips, and the token budget stops the runaway at 18 turns.
  3. Compare python report.py --trace run-02 with --trace run-01 and find the two wasted turns.

pip install pytest && pytest -q runs the offline tests: span structure, the cost maths, and that each limit stops the runaway before the call it guards.

Common beginner mistakes

  • Tracing calls, not runs. Per-call records hide which ticket a call belonged to and in what order. Give every span in a run the same trace id.
  • Logging the model’s choices without the tool results. The model acts on what the tools returned. Without it, you can see what it did and never why.
  • Quoting cost per call. Later turns resend everything earlier, so a 12-turn run costs far more than 12 first turns. Sum the turns into a cost per run.
  • A turn limit and nothing else. It bounds the number of calls, not their size. Add a token and a dollar budget.
  • Averaging over runs instead of dividing by resolved tickets. Failed and stopped runs still cost money; cost per resolved ticket is where they show up.

Questions you will face in production

“Do I build this or use a hosted tool?” If you already run Datadog, Honeycomb or Grafana, send OpenTelemetry spans with the generative AI names. If not, a hosted LLM tracing tool (LangSmith, Langfuse, Phoenix) shows agent runs as nested spans on day one. Either way, check it records tool arguments and results, tokens per turn, the running cost, and why the run ended.

“A legitimate task needs 40 turns. Doesn’t the budget kill it?” Give that task type its own budget, set from its own traces. And read those traces first: often a better tool (one that returns what the model needs in one call) or trimming old tool results from the history makes the run shorter and cheaper.

“What happens to the trace when a run is resumed after a crash?” Store the trace id and the running totals with the checkpoint (article 07 covers resuming runs). Then the spans before and after the crash land in one trace, and the budget doesn’t restart from zero.

Check your understanding

Average cost per run has been flat for a week, but cost per resolved ticket went up 40%. What's the most likely cause, and where do you look?

More runs are ending without resolving anything, as handoffs or budget stops, so the same spend is spread over fewer resolved tickets. Group the week’s runs by outcome and read traces from the group that grew, starting with the budget stops.

A teammate sets max_turns = 50 because "turns are cheap, about half a cent each". What's wrong with that reasoning?

Only the early turns cost half a cent. Every turn resends the whole history, so late turns cost many times the first, and on the KITE-140 script 50 turns cost $1.55, about 80 normal runs. Keep a dollar or token budget too.

The budget is $0.10, but a stopped run's trace says it spent $0.1099. Is the budget broken?

No. It’s checked before each call, and the run was at $0.0945 before its last turn, so that turn was allowed and took it over. If one turn of overshoot matters, estimate the next call with count_tokens first.

A resolved run's trace shows search_help("sso reset password") twice in a row. The ticket was answered correctly. Is it worth fixing?

Yes, if it’s common. The repeat costs its own turn, and its result rides along in every later turn: that’s why run-02 cost 75% more than run-01. Count how many runs repeat an identical read-only call, and if it’s many, fix the prompt or tool description, or answer the repeat with a note from your code.

What to remember

  • Trace the run, not the call: one trace per run, a span per model turn and per tool call, all sharing a trace id. Log tool arguments and results, since that’s what the model acted on.
  • Every turn resends the whole history, so later turns cost more and a run’s cost grows faster than its turn count. Price runs, and use prompt caching for the resent prefix.
  • Cost per resolved ticket (total spend over tickets resolved) is the number that moves when runs loop, waste turns or hand off.
  • Give every run a budget of turns, tokens and dollars, checked before each model call. A stop is an outcome your code handles, usually a handoff to a person.
  • Loops show up in traces as repeated identical tool calls, one tool dominating, and far more turns than the task needs. Alert on the share of runs hitting a limit, not on each one.

What to study next

You can now see inside a run and bound what it costs. Article 09: Multi-Agent Systems covers splitting work across several agents, where one task becomes runs inside runs: each sub-agent needs its spans under the parent’s trace and its own share of the budget. For the per-token levers behind every number here, keep article 12: Cost Optimization handy.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.