Cost Optimization in Production AI

The Kitebase support bot from article 04 has grown up. It answers about 10,000 questions a day. Every question goes to claude-opus-5 with a 2,400-token system prompt (rules, tone and examples that piled up over the months) and the top 8 help-center excerpts “to be safe”, and every closed conversation gets a summary written for the support team. Then finance asks a fair question: what does it cost to solve one ticket? Nobody knows. There’s a monthly invoice, and it keeps going up.

Grabbing the first tip you read (switch to a cheaper model!) can backfire, as you’ll see. The fix is boring: price every request from the numbers the API already sends back, then change one thing at a time and measure again. Article 10 turned on prompt caching; this article starts from the version without it, so you can see what caching and five other levers are each worth.

What you’ll build: a cost calculator that reads one hour of the Kitebase bot’s logged requests, prices each one, and replays the log with six levers switched on one after another. It prints the cost per request and per resolved conversation after each lever. It’s plain Python and runs offline: all arithmetic, no API calls.

Step 0: Measure what a request costs

You pay per token, the chunk of text a model reads and writes (roughly 4 characters of English; article 01 covers them). Every response from the Messages API carries a usage object with four counts, and each is billed at its own rate:

usage fieldWhat it countsclaude-opus-5, per million
input_tokensprompt tokens that weren’t cached$5.00
cache_creation_input_tokensprompt tokens written to the cache$6.25 (1.25x input)
cache_read_input_tokensprompt tokens read from the cache$0.50
output_tokenseverything the model wrote$25.00

The two cache fields are zero until you turn caching on (lever 1). Pricing a response is four multiplications:

PRICES = {"claude-opus-5": {"input": 5.00, "output": 25.00, "cache_read": 0.50}, ...}
CACHE_WRITE = 1.25
BATCH = 0.5

def price(usage: dict, model: str, batch: bool = False) -> float:
    p = PRICES[model]
    per_million = (usage["input_tokens"] * p["input"]
                   + usage["cache_creation_input_tokens"] * p["input"] * CACHE_WRITE
                   + usage["cache_read_input_tokens"] * p["cache_read"]
                   + usage["output_tokens"] * p["output"])
    return per_million / 1_000_000 * (BATCH if batch else 1)

To use it you have to log usage with every request, next to what the request was for. One row of the Kitebase log:

{"id": "r05", "at": "06:04:30", "conversation": "c03", "feature": "chat", "turn": 1,
 "text": "Where do I download last month's invoice?", "top_score": 0.64,
 "parts": {"system": 2400, "history": 0, "excerpts": [119, 103, 84, 100, 109, 116, 102, 115], "question": 17},
 "usage": {"input_tokens": 3265, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 0, "output_tokens": 390}}

usage is straight from the response. parts splits the prompt into pieces, which your own code has to log because the API only reports totals; it’s what lets you answer “what if we sent fewer excerpts?” without running anything. top_score is the best excerpt’s search score from article 04. Article 11 covers where logs like this live.

Per request, and per completed task

Cost per request is easy to compute and easy to misread. A customer who needs four messages costs four requests plus a summary, and a conversation that ends with “let me get a human” still cost money. The business pays for solved problems, so the number that matters is cost per completed task: here, total spend divided by conversations the bot resolved without a human.

The sample log is one quiet hour, 06:00 to 07:00, before the working day starts: 30 requests across 10 conversations, 8 of them resolved. It’s small enough to read in full. The first thing main.py prints is where the money goes:

Sample log: 30 requests (20 chat, 10 summaries), 10 conversations, 8 resolved.

Where the money goes (logged usage, everything on claude-opus-5):
  input_tokens    77,420 tokens  $0.3871  45%
  output_tokens   18,743 tokens  $0.4686  55%
  chat requests                 $0.6691  78%
  summary requests              $0.1865  22%

Output is a fifth of the tokens and more than half the bill: on all three Claude models here it costs 5 times the input price. Part of it is text nobody sees. claude-opus-5 thinks before it answers by default, and those thinking tokens are billed as output. The baseline is $0.8557 for the hour, $0.0285 per request, and $0.1070 per resolved conversation. That last number is the one to bring to finance, and the one every lever below gets judged by.

The order to pull the levers

Six levers, cheapest and safest first. Two of them only change how you’re billed: the model gets the same prompt and gives the same kind of answer. The other four change what the model sees or writes, so each one needs an eval run (the labelled test questions from article 08) to confirm answers didn’t get worse. The calculator applies them one on top of another, and each percentage below is against the step before it.

Lever 1: Prompt caching

Every chat request starts with the same 2,400-token system prompt: rules, tone, product facts, examples. The model re-reads it from scratch every time, and you pay full price every time.

Prompt caching lets the provider keep the processed start of a prompt for a while, so the next request that starts with exactly the same text skips re-reading it. On Anthropic’s API a cached prefix lives 5 minutes, and every read restarts the clock. Reads cost a tenth of the input price. The first request, which writes the cache, costs 1.25 times. You mark where the cacheable part ends with cache_control:

params["system"] = [{"type": "text", "text": system, "cache_control": {"type": "ephemeral"}}]

The same request, priced both ways:

r05: "Where do I download last month's invoice?" on claude-opus-5 As logged, no caching RESPONSE.USAGE input_tokens: 3,265 cache_creation_input_tokens: 0 cache_read_input_tokens: 0 output_tokens: 390 PRICE(): TOKENS × RATE PER MILLION 3,265 × $5.00 input $0.0163 390 × $25.00 output $0.0098 COST $0.0261 With the 2,400-token system prompt read from a warm cache RESPONSE.USAGE input_tokens: 865 cache_creation_input_tokens: 0 cache_read_input_tokens: 2,400 output_tokens: 390 PRICE(): TOKENS × RATE PER MILLION 865 × $5.00 input $0.0043 2,400 × $0.50 cache read $0.0012 390 × $25.00 output $0.0098 COST $0.0153 41% less input_tokens is only the uncached part. Add all three fields for the prompt size. 390 output tokens now cost more than the whole 3,265-token prompt. A cache write costs $0.0291, more than no cache at all.
One logged request priced before and after caching. The numbers are what the companion code computes.

A write costs more than not caching, so caching only pays when a read follows within 5 minutes. The calculator replays the log in time order to count them:

if r.cache and prefix >= CACHE_MIN[r.model]:
    key = (r.model, r.feature)  # each model keeps its own cache
    hit = warm or (key in last_used and r.at - last_used[key] <= CACHE_TTL)
    usage["cache_read_input_tokens" if hit else "cache_creation_input_tokens"] = prefix
    usage["input_tokens"] = rest
    last_used[key] = r.at
1. prompt caching           $0.6949     $0.0232      $0.0869    -19%

With caching on, 16 requests read the system prompt from cache and 4 wrote it.

Four times in the hour the bot sat idle for more than 5 minutes, so the next request paid to write the cache again. python main.py --busy (the warm flag above) prices the same requests as if steady traffic kept the cache warm, which is closer to mid-morning: every request reads, and caching saves 25% instead of 19%.

Two things make caching silently do nothing. First, the prefix must match byte for byte: a timestamp or the customer’s name at the top of the system prompt makes every request a write. Second, there’s a minimum length: 512 tokens on claude-opus-5, 1,024 on claude-sonnet-5, 4,096 on claude-haiku-4-5. The summary prompt is 350 tokens, so marking it does nothing, with no error. Article 10 covers both traps in depth.

Lever 2: Send less context

The bot sends the top 8 excerpts with every question, though the answer is almost always in the first one or two. Cutting top_k (how many excerpts search returns) from 8 to 3 removes about 550 tokens per chat request:

2. top_k 8 -> 3             $0.6397     $0.0213      $0.0800     -8%

Only 8%, because after lever 1 output is two thirds of the bill. That’s why you measure after every change. This lever changes what the model sees, so it needs the eval: article 08 showed top_k 1 losing the backup-codes answer for a customer who lost their phone, and 3 passed. Shorter chunks (article 05) and trimming old conversation turns work the same way.

Lever 3: Ask for shorter answers

Output is $25 per million on claude-opus-5, and most support answers don’t need 400 words.

The tempting fix is a small max_tokens. Don’t. It’s a hard stop, not a length request: the reply gets cut off mid-sentence, stop_reason comes back as "max_tokens", and you pay for every token up to the cut. It also counts the thinking, so a tight limit can end a reply before the answer starts. Keep max_tokens generous as a backstop (2,048 in the example) and ask for the length in the system prompt:

SHORT_ANSWER_RULE = "\nAnswer in at most 120 words."

Arithmetic can’t tell you what that saves. You replay a sample with the new prompt and compare output_tokens. The example’s SHORT_ANSWERS = 0.62 stands for “the replay averaged 0.62 times the output tokens”. It’s made up for the demo; it’s the slot for your own measurement:

3. shorter answers          $0.5176     $0.0173      $0.0647    -19%

Run the eval on the same replay: a shorter answer can drop the step a customer needed.

Lever 4: Batch what nobody waits for

The summaries go to the support team’s inbox. Nobody is waiting on them, yet they’re sent like a live chat answer.

The Batch API takes a list of requests, runs them in the background, and bills every token at half price. Most batches finish within an hour; anything not done after 24 hours expires. Each request gets a custom_id so you can match results back:

batch = client.messages.batches.create(requests=[
    {"custom_id": "r02", "params": {"model": "claude-opus-5", "max_tokens": 2048,
                                    "system": SUMMARY_PROMPT, "messages": transcript_c01}},
    # ... one entry per closed conversation
])

# Later, from a scheduled job:
for result in client.messages.batches.results(batch.id):
    if result.result.type == "succeeded":
        save_summary(result.custom_id, result.result.message)
4. batch the summaries      $0.4243     $0.0141      $0.0530    -18%

Same model, same prompt, same answer, so no eval needed. The work is plumbing: a job that submits the batch, checks back until it’s done, and resubmits results that come back "errored" or "expired". Cache hits inside a batch are best-effort, so don’t count on them.

Lever 5: Route simple requests to a cheaper model

“How long does the reset link last?” doesn’t need the most capable model. Routing means picking the model per request, before the call, using only what you know by then. Kitebase’s rule: the first message of a conversation, where search found a clearly matching article.

def is_simple(r: Request) -> bool:
    return r.feature == "chat" and r.turn == 1 and r.top_score >= 0.5

That picks 7 of the 20 chats: reset password, invoice download, two-factor, sign-in email, reset link, backup code, downgrade. The harder ones, like a failing SSO setup or a double charge, stay on claude-opus-5.

The candidates are claude-sonnet-5 ($2 in, $10 out per million) and claude-haiku-4-5 ($1 in, $5 out). Anthropic pitches Haiku 4.5 as its fastest, most cost-effective model for simple tasks, and these questions are narrow and checkable: the answer sits in one excerpt, and the citation check from article 04 catches a made-up source. Sonnet is the safer step down in quality. The calculator prices the 7 simple chats on each:

The 7 simple chats on each model (after levers 1 to 4):
  claude-opus-5     $0.1197  cache: 3 read, 4 written
  claude-sonnet-5   $0.0589  cache: 1 read, 6 written
  claude-haiku-4-5  $0.0280  cache: 0 read, 0 written

Haiku never caches here: the 2,400-token system prompt is under its 4,096-token minimum. It still wins, because in a quiet hour Sonnet’s cache barely works either. Seven requests spread over an hour mostly arrive more than 5 minutes apart, so Sonnet pays for 6 cache writes and reads one:

06:00 06:10 06:20 06:30 06:40 06:50 07:00 Levers 1 to 4: all 20 chats on claude-opus-5 16 read, 4 written idle over 5 minutes: the next one writes Lever 5: claude-opus-5 keeps the 13 harder chats 9 read, 4 written Lever 5: claude-haiku-4-5 takes the 7 simple ones never cached The 2,400-token prompt is under Haiku 4.5's 4,096-token minimum If they went to claude-sonnet-5 instead 1 read, 6 written 5 of the 6 writes are never read: 1.25x for nothing write: 1.25x input price read: 0.1x input price not cached: full price cache warm (5 minutes)
Each model keeps its own cache. Splitting traffic across models spreads it thinner.
5. route simple chats       $0.3879     $0.0129      $0.0485     -9%

With ROUTE_TO = "claude-sonnet-5" the same step saves 1%. Now run python main.py --busy, where every cache stays warm:

  claude-opus-5     $0.0645  cache: 7 read, 0 written
  claude-sonnet-5   $0.0258  cache: 7 read, 0 written
  claude-haiku-4-5  $0.0280  cache: 0 read, 0 written

The answer flips: with a warm cache, Sonnet is 8% cheaper than Haiku. Kitebase has busy days and quiet nights, so the example routes to Haiku, whose worst case is much better: 8% dearer when Sonnet’s cache is fully warm, less than half the price when it isn’t. With a stable prefix over 4,096 tokens Haiku would cache too and win outright. Compute it on a full day of your own log, not on per-token prices.

Two cautions. The calculator reuses Opus’s token counts, but tokenizers (the code that splits text into tokens) differ between model generations, so check a sample with client.messages.count_tokens on the new model. And routing changes who writes the answer, so run the eval on the simple questions with the cheaper model first.

Lever 6: Lower the effort

A summary of a finished conversation doesn’t need deep reasoning, but claude-opus-5 decides for itself how much to think. Output, thinking included, is 79% of what the summaries cost.

output_config={"effort": "low"} tells the model to think less and answer sooner. Levels go from "low" through "medium" to "high" and beyond; leaving it out uses the model’s default. It’s one line in request_params:

if r.effort:
    params["output_config"] = {"effort": r.effort}

Like shorter answers, the saving comes from a replay. The made-up LOW_EFFORT = 0.45 stands for “summaries at low effort averaged 0.45 times the output tokens”:

6. low effort on summaries  $0.3473     $0.0116      $0.0434    -10%

The complex chats keep the default: that’s where thinking earns its cost, and customers read those answers. Set effort per feature and leave it fixed, because changing it between requests in one conversation can invalidate the cache.

Putting it together

Cost per resolved conversation, one lever at a time 0. Baseline all Opus, no cache, top_k 8 $0.1070 1. Prompt caching system prompt read at 0.1x $0.0869 −19% 2. top_k 8 to 3 5 fewer excerpts per chat $0.0800 −8% 3. Shorter answers chat output x0.62 (replay) $0.0647 −19% 4. Batch the summaries 10 summaries at half price $0.0530 −18% 5. Route simple chats 7 of 20 chats on Haiku 4.5 $0.0485 −9% 6. Low effort, summaries summary output x0.45 (replay) $0.0434 −10% Per 1,000 resolved conversations: $106.96 before, $43.42 after all six (−59%). same answers, just cheaper changes what the model sees or writes: run the eval what the lever cut
The companion code's output. Two levers are free; four need an eval before they ship.

All six take a resolved conversation from $0.1070 to $0.0434, 59% less on this log. That stands on a made-up hour and two made-up replay ratios: the method carries over, the percentages don’t.

The credit each lever gets depends on the order. Move routing to the front and it looks like 17%, while caching shrinks to 12%; the end result is the same $0.0434. So “which change saved the most?” has no answer without “in which order?”

The rest of the bill. Embeddings add up when you re-embed everything on every deploy, so re-embed only chunks whose text changed (article 10’s embedding cache). Vector databases bill for storage and queries, and a re-ranker (article 06) is an extra model call per question. Price them the same way: per request, then per completed task.

When not to bother. If the bill is a few hundred dollars a month, your time is worth more. If answers aren’t good enough yet, fix quality first: you need a passing eval before you can tell whether a cut cost you anything.

Try it yourself

The companion example is the calculator: the sample log, the six levers, and request_params(), which turns each request’s settings into real client.messages.create arguments. It runs offline.

Download the runnable example (zip)

cd 12-cost-optimization
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. Set ROUTE_TO = "claude-sonnet-5" and run it. Step 5 drops to 1%. Then run python main.py --busy with the same setting and compare: a warm cache changes which model is cheaper.
  2. Move the ("route simple chats", route_simple) line to the top of LEVERS. Routing now looks like -17% and caching like -12%, and the last line doesn’t move.
  3. Set SHORT_ANSWERS = 1.0, as if the replay showed no change. See how much of the total came from output tokens.

pip install pytest && pytest -q runs the offline tests. They check the pricing, the cache timing and minimums, each lever, and the numbers quoted here.

Common beginner mistakes

  • Reading input_tokens as the prompt size. With caching on it’s only the uncached part. Add all three input fields.
  • Tracking only the monthly total. It can’t tell you what grew. Log usage with the feature and conversation.
  • Using max_tokens to shorten answers. It truncates, and you pay for the truncated reply. Ask for the length in the prompt.
  • Something that changes at the top of the prompt. A timestamp, a request id or the user’s name before the stable part turns every cache read into a write.
  • Switching models on per-token price alone. Cache minimums, cache warmth and token counts change with the model. Price it on your log.
  • Cutting cost without an eval. A cheaper bot that resolves fewer conversations can cost more per resolved conversation.

Questions you will face in production

“What does one solved ticket cost?” Total spend divided by conversations solved without a human, from logged usage. Track it per feature and chart it daily next to the resolution rate, so a saving that hurts quality shows up as a rising cost per solved ticket.

“How do we avoid surprise bills?” Alert when a feature’s cost per request jumps, not just on the monthly total. Keep max_tokens as a backstop, and check cache_read_input_tokens after every prompt change: a broken cache raises no error.

“Is a week of optimization worth it?” Multiply the per-task saving by your monthly volume. $0.06 per conversation at 1,000 conversations a month is $60, not worth a week. At 200,000 a month it’s $12,000.

Check your understanding

You added cache_control to the system prompt, but cache_read_input_tokens is 0 on every request. Name three causes.

Something in the prefix changes per request (a timestamp, a user name), so it never matches byte for byte. The prefix is under the model’s minimum (512 tokens on claude-opus-5, 4,096 on claude-haiku-4-5). Or requests arrive more than 5 minutes apart, so every one writes a fresh cache.

A teammate sets max_tokens=200 to cut output cost. What happens to the bot?

Long answers get cut off mid-sentence with stop_reason "max_tokens", and you still pay for the 200 tokens. On claude-opus-5 thinking counts toward the limit too, so some replies end before the answer starts. Ask for a short answer in the prompt and keep max_tokens as a generous backstop.

Sonnet 5 costs well under half of Opus 5 per token, but routing the simple chats to it saved only 1% in the quiet hour. Why?

Seven requests spread over an hour leave Sonnet’s cache cold, so it pays 1.25x to write the system prompt six times and reads it once. Those simple chats also used to keep Opus’s cache warm for the harder ones. With a warm cache (--busy) the same move saves about 10%. Traffic and caching matter as much as the per-token price.

Which levers can ship without an eval run, and why?

Prompt caching and the Batch API. The model gets exactly the same request and gives the same kind of answer; only the billing changes. Fewer excerpts, shorter answers, a different model and lower effort all change what the model sees or writes.

After routing, cost per request fell but cost per resolved conversation went up. What happened?

The cheaper model resolved fewer conversations: more follow-up messages, more handoffs to a human. Each request got cheaper, but each solved problem got more expensive. That’s why cost per completed task is the number to watch.

What to remember

  • Price every request from its usage: four fields, four rates. Log it next to the feature and conversation.
  • Judge changes by cost per completed task, not per request or per month.
  • Caching and batching don’t change answers: do them first. The other levers need an eval.
  • Output costs 5 times input on these models; ask for shorter answers, don’t truncate them.
  • Caches are per model, have a minimum length and expire after 5 idle minutes. Routing interacts with all three.
  • The saving from each lever depends on your traffic and on the order you apply them. Compute it on your own log.

What to study next

That’s Track 04. The capstone, Build It: A Small RAG Service That Ships, wires ingestion, retrieval with a minimum score, a versioned prompt, citations, evals, caching and cost logging into one small service you can run.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.