Cost Optimization in Production AI
The Kitebase support bot from article 04 has grown up. It answers about 10,000 questions a day. Every question goes to claude-opus-5 with a 2,400-token system prompt (rules, tone and examples that piled up over the months) and the top 8 help-center excerpts “to be safe”, and every closed conversation gets a summary written for the support team. Then finance asks a fair question: what does it cost to solve one ticket? Nobody knows. There’s a monthly invoice, and it keeps going up.
Grabbing the first tip you read (switch to a cheaper model!) can backfire, as you’ll see. The fix is boring: price every request from the numbers the API already sends back, then change one thing at a time and measure again. Article 10 turned on prompt caching; this article starts from the version without it, so you can see what caching and five other levers are each worth.
What you’ll build: a cost calculator that reads one hour of the Kitebase bot’s logged requests, prices each one, and replays the log with six levers switched on one after another. It prints the cost per request and per resolved conversation after each lever. It’s plain Python and runs offline: all arithmetic, no API calls.
Step 0: Measure what a request costs
You pay per token, the chunk of text a model reads and writes (roughly 4 characters of English; article 01 covers them). Every response from the Messages API carries a usage object with four counts, and each is billed at its own rate:
usage field | What it counts | claude-opus-5, per million |
|---|---|---|
input_tokens | prompt tokens that weren’t cached | $5.00 |
cache_creation_input_tokens | prompt tokens written to the cache | $6.25 (1.25x input) |
cache_read_input_tokens | prompt tokens read from the cache | $0.50 |
output_tokens | everything the model wrote | $25.00 |
The two cache fields are zero until you turn caching on (lever 1). Pricing a response is four multiplications:
PRICES = {"claude-opus-5": {"input": 5.00, "output": 25.00, "cache_read": 0.50}, ...}
CACHE_WRITE = 1.25
BATCH = 0.5
def price(usage: dict, model: str, batch: bool = False) -> float:
p = PRICES[model]
per_million = (usage["input_tokens"] * p["input"]
+ usage["cache_creation_input_tokens"] * p["input"] * CACHE_WRITE
+ usage["cache_read_input_tokens"] * p["cache_read"]
+ usage["output_tokens"] * p["output"])
return per_million / 1_000_000 * (BATCH if batch else 1)
To use it you have to log usage with every request, next to what the request was for. One row of the Kitebase log:
{"id": "r05", "at": "06:04:30", "conversation": "c03", "feature": "chat", "turn": 1,
"text": "Where do I download last month's invoice?", "top_score": 0.64,
"parts": {"system": 2400, "history": 0, "excerpts": [119, 103, 84, 100, 109, 116, 102, 115], "question": 17},
"usage": {"input_tokens": 3265, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 0, "output_tokens": 390}}
usage is straight from the response. parts splits the prompt into pieces, which your own code has to log because the API only reports totals; it’s what lets you answer “what if we sent fewer excerpts?” without running anything. top_score is the best excerpt’s search score from article 04. Article 11 covers where logs like this live.
Per request, and per completed task
Cost per request is easy to compute and easy to misread. A customer who needs four messages costs four requests plus a summary, and a conversation that ends with “let me get a human” still cost money. The business pays for solved problems, so the number that matters is cost per completed task: here, total spend divided by conversations the bot resolved without a human.
The sample log is one quiet hour, 06:00 to 07:00, before the working day starts: 30 requests across 10 conversations, 8 of them resolved. It’s small enough to read in full. The first thing main.py prints is where the money goes:
Sample log: 30 requests (20 chat, 10 summaries), 10 conversations, 8 resolved.
Where the money goes (logged usage, everything on claude-opus-5):
input_tokens 77,420 tokens $0.3871 45%
output_tokens 18,743 tokens $0.4686 55%
chat requests $0.6691 78%
summary requests $0.1865 22%
Output is a fifth of the tokens and more than half the bill: on all three Claude models here it costs 5 times the input price. Part of it is text nobody sees. claude-opus-5 thinks before it answers by default, and those thinking tokens are billed as output. The baseline is $0.8557 for the hour, $0.0285 per request, and $0.1070 per resolved conversation. That last number is the one to bring to finance, and the one every lever below gets judged by.
The order to pull the levers
Six levers, cheapest and safest first. Two of them only change how you’re billed: the model gets the same prompt and gives the same kind of answer. The other four change what the model sees or writes, so each one needs an eval run (the labelled test questions from article 08) to confirm answers didn’t get worse. The calculator applies them one on top of another, and each percentage below is against the step before it.
Lever 1: Prompt caching
Every chat request starts with the same 2,400-token system prompt: rules, tone, product facts, examples. The model re-reads it from scratch every time, and you pay full price every time.
Prompt caching lets the provider keep the processed start of a prompt for a while, so the next request that starts with exactly the same text skips re-reading it. On Anthropic’s API a cached prefix lives 5 minutes, and every read restarts the clock. Reads cost a tenth of the input price. The first request, which writes the cache, costs 1.25 times. You mark where the cacheable part ends with cache_control:
params["system"] = [{"type": "text", "text": system, "cache_control": {"type": "ephemeral"}}]
The same request, priced both ways:
A write costs more than not caching, so caching only pays when a read follows within 5 minutes. The calculator replays the log in time order to count them:
if r.cache and prefix >= CACHE_MIN[r.model]:
key = (r.model, r.feature) # each model keeps its own cache
hit = warm or (key in last_used and r.at - last_used[key] <= CACHE_TTL)
usage["cache_read_input_tokens" if hit else "cache_creation_input_tokens"] = prefix
usage["input_tokens"] = rest
last_used[key] = r.at
1. prompt caching $0.6949 $0.0232 $0.0869 -19%
With caching on, 16 requests read the system prompt from cache and 4 wrote it.
Four times in the hour the bot sat idle for more than 5 minutes, so the next request paid to write the cache again. python main.py --busy (the warm flag above) prices the same requests as if steady traffic kept the cache warm, which is closer to mid-morning: every request reads, and caching saves 25% instead of 19%.
Two things make caching silently do nothing. First, the prefix must match byte for byte: a timestamp or the customer’s name at the top of the system prompt makes every request a write. Second, there’s a minimum length: 512 tokens on claude-opus-5, 1,024 on claude-sonnet-5, 4,096 on claude-haiku-4-5. The summary prompt is 350 tokens, so marking it does nothing, with no error. Article 10 covers both traps in depth.
Lever 2: Send less context
The bot sends the top 8 excerpts with every question, though the answer is almost always in the first one or two. Cutting top_k (how many excerpts search returns) from 8 to 3 removes about 550 tokens per chat request:
2. top_k 8 -> 3 $0.6397 $0.0213 $0.0800 -8%
Only 8%, because after lever 1 output is two thirds of the bill. That’s why you measure after every change. This lever changes what the model sees, so it needs the eval: article 08 showed top_k 1 losing the backup-codes answer for a customer who lost their phone, and 3 passed. Shorter chunks (article 05) and trimming old conversation turns work the same way.
Lever 3: Ask for shorter answers
Output is $25 per million on claude-opus-5, and most support answers don’t need 400 words.
The tempting fix is a small max_tokens. Don’t. It’s a hard stop, not a length request: the reply gets cut off mid-sentence, stop_reason comes back as "max_tokens", and you pay for every token up to the cut. It also counts the thinking, so a tight limit can end a reply before the answer starts. Keep max_tokens generous as a backstop (2,048 in the example) and ask for the length in the system prompt:
SHORT_ANSWER_RULE = "\nAnswer in at most 120 words."
Arithmetic can’t tell you what that saves. You replay a sample with the new prompt and compare output_tokens. The example’s SHORT_ANSWERS = 0.62 stands for “the replay averaged 0.62 times the output tokens”. It’s made up for the demo; it’s the slot for your own measurement:
3. shorter answers $0.5176 $0.0173 $0.0647 -19%
Run the eval on the same replay: a shorter answer can drop the step a customer needed.
Lever 4: Batch what nobody waits for
The summaries go to the support team’s inbox. Nobody is waiting on them, yet they’re sent like a live chat answer.
The Batch API takes a list of requests, runs them in the background, and bills every token at half price. Most batches finish within an hour; anything not done after 24 hours expires. Each request gets a custom_id so you can match results back:
batch = client.messages.batches.create(requests=[
{"custom_id": "r02", "params": {"model": "claude-opus-5", "max_tokens": 2048,
"system": SUMMARY_PROMPT, "messages": transcript_c01}},
# ... one entry per closed conversation
])
# Later, from a scheduled job:
for result in client.messages.batches.results(batch.id):
if result.result.type == "succeeded":
save_summary(result.custom_id, result.result.message)
4. batch the summaries $0.4243 $0.0141 $0.0530 -18%
Same model, same prompt, same answer, so no eval needed. The work is plumbing: a job that submits the batch, checks back until it’s done, and resubmits results that come back "errored" or "expired". Cache hits inside a batch are best-effort, so don’t count on them.
Lever 5: Route simple requests to a cheaper model
“How long does the reset link last?” doesn’t need the most capable model. Routing means picking the model per request, before the call, using only what you know by then. Kitebase’s rule: the first message of a conversation, where search found a clearly matching article.
def is_simple(r: Request) -> bool:
return r.feature == "chat" and r.turn == 1 and r.top_score >= 0.5
That picks 7 of the 20 chats: reset password, invoice download, two-factor, sign-in email, reset link, backup code, downgrade. The harder ones, like a failing SSO setup or a double charge, stay on claude-opus-5.
The candidates are claude-sonnet-5 ($2 in, $10 out per million) and claude-haiku-4-5 ($1 in, $5 out). Anthropic pitches Haiku 4.5 as its fastest, most cost-effective model for simple tasks, and these questions are narrow and checkable: the answer sits in one excerpt, and the citation check from article 04 catches a made-up source. Sonnet is the safer step down in quality. The calculator prices the 7 simple chats on each:
The 7 simple chats on each model (after levers 1 to 4):
claude-opus-5 $0.1197 cache: 3 read, 4 written
claude-sonnet-5 $0.0589 cache: 1 read, 6 written
claude-haiku-4-5 $0.0280 cache: 0 read, 0 written
Haiku never caches here: the 2,400-token system prompt is under its 4,096-token minimum. It still wins, because in a quiet hour Sonnet’s cache barely works either. Seven requests spread over an hour mostly arrive more than 5 minutes apart, so Sonnet pays for 6 cache writes and reads one:
5. route simple chats $0.3879 $0.0129 $0.0485 -9%
With ROUTE_TO = "claude-sonnet-5" the same step saves 1%. Now run python main.py --busy, where every cache stays warm:
claude-opus-5 $0.0645 cache: 7 read, 0 written
claude-sonnet-5 $0.0258 cache: 7 read, 0 written
claude-haiku-4-5 $0.0280 cache: 0 read, 0 written
The answer flips: with a warm cache, Sonnet is 8% cheaper than Haiku. Kitebase has busy days and quiet nights, so the example routes to Haiku, whose worst case is much better: 8% dearer when Sonnet’s cache is fully warm, less than half the price when it isn’t. With a stable prefix over 4,096 tokens Haiku would cache too and win outright. Compute it on a full day of your own log, not on per-token prices.
Two cautions. The calculator reuses Opus’s token counts, but tokenizers (the code that splits text into tokens) differ between model generations, so check a sample with client.messages.count_tokens on the new model. And routing changes who writes the answer, so run the eval on the simple questions with the cheaper model first.
Lever 6: Lower the effort
A summary of a finished conversation doesn’t need deep reasoning, but claude-opus-5 decides for itself how much to think. Output, thinking included, is 79% of what the summaries cost.
output_config={"effort": "low"} tells the model to think less and answer sooner. Levels go from "low" through "medium" to "high" and beyond; leaving it out uses the model’s default. It’s one line in request_params:
if r.effort:
params["output_config"] = {"effort": r.effort}
Like shorter answers, the saving comes from a replay. The made-up LOW_EFFORT = 0.45 stands for “summaries at low effort averaged 0.45 times the output tokens”:
6. low effort on summaries $0.3473 $0.0116 $0.0434 -10%
The complex chats keep the default: that’s where thinking earns its cost, and customers read those answers. Set effort per feature and leave it fixed, because changing it between requests in one conversation can invalidate the cache.
Putting it together
All six take a resolved conversation from $0.1070 to $0.0434, 59% less on this log. That stands on a made-up hour and two made-up replay ratios: the method carries over, the percentages don’t.
The credit each lever gets depends on the order. Move routing to the front and it looks like 17%, while caching shrinks to 12%; the end result is the same $0.0434. So “which change saved the most?” has no answer without “in which order?”
The rest of the bill. Embeddings add up when you re-embed everything on every deploy, so re-embed only chunks whose text changed (article 10’s embedding cache). Vector databases bill for storage and queries, and a re-ranker (article 06) is an extra model call per question. Price them the same way: per request, then per completed task.
When not to bother. If the bill is a few hundred dollars a month, your time is worth more. If answers aren’t good enough yet, fix quality first: you need a passing eval before you can tell whether a cut cost you anything.
Try it yourself
The companion example is the calculator: the sample log, the six levers, and request_params(), which turns each request’s settings into real client.messages.create arguments. It runs offline.
Download the runnable example (zip)
cd 12-cost-optimization
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
- Set
ROUTE_TO = "claude-sonnet-5"and run it. Step 5 drops to 1%. Then runpython main.py --busywith the same setting and compare: a warm cache changes which model is cheaper. - Move the
("route simple chats", route_simple)line to the top ofLEVERS. Routing now looks like -17% and caching like -12%, and the last line doesn’t move. - Set
SHORT_ANSWERS = 1.0, as if the replay showed no change. See how much of the total came from output tokens.
pip install pytest && pytest -q runs the offline tests. They check the pricing, the cache timing and minimums, each lever, and the numbers quoted here.
Common beginner mistakes
- Reading
input_tokensas the prompt size. With caching on it’s only the uncached part. Add all three input fields. - Tracking only the monthly total. It can’t tell you what grew. Log
usagewith the feature and conversation. - Using
max_tokensto shorten answers. It truncates, and you pay for the truncated reply. Ask for the length in the prompt. - Something that changes at the top of the prompt. A timestamp, a request id or the user’s name before the stable part turns every cache read into a write.
- Switching models on per-token price alone. Cache minimums, cache warmth and token counts change with the model. Price it on your log.
- Cutting cost without an eval. A cheaper bot that resolves fewer conversations can cost more per resolved conversation.
Questions you will face in production
“What does one solved ticket cost?”
Total spend divided by conversations solved without a human, from logged usage. Track it per feature and chart it daily next to the resolution rate, so a saving that hurts quality shows up as a rising cost per solved ticket.
“How do we avoid surprise bills?”
Alert when a feature’s cost per request jumps, not just on the monthly total. Keep max_tokens as a backstop, and check cache_read_input_tokens after every prompt change: a broken cache raises no error.
“Is a week of optimization worth it?” Multiply the per-task saving by your monthly volume. $0.06 per conversation at 1,000 conversations a month is $60, not worth a week. At 200,000 a month it’s $12,000.
Check your understanding
You added cache_control to the system prompt, but cache_read_input_tokens is 0 on every request. Name three causes.
Something in the prefix changes per request (a timestamp, a user name), so it never matches byte for byte. The prefix is under the model’s minimum (512 tokens on claude-opus-5, 4,096 on claude-haiku-4-5). Or requests arrive more than 5 minutes apart, so every one writes a fresh cache.
A teammate sets max_tokens=200 to cut output cost. What happens to the bot?
Long answers get cut off mid-sentence with stop_reason "max_tokens", and you still pay for the 200 tokens. On claude-opus-5 thinking counts toward the limit too, so some replies end before the answer starts. Ask for a short answer in the prompt and keep max_tokens as a generous backstop.
Sonnet 5 costs well under half of Opus 5 per token, but routing the simple chats to it saved only 1% in the quiet hour. Why?
Seven requests spread over an hour leave Sonnet’s cache cold, so it pays 1.25x to write the system prompt six times and reads it once. Those simple chats also used to keep Opus’s cache warm for the harder ones. With a warm cache (--busy) the same move saves about 10%. Traffic and caching matter as much as the per-token price.
Which levers can ship without an eval run, and why?
Prompt caching and the Batch API. The model gets exactly the same request and gives the same kind of answer; only the billing changes. Fewer excerpts, shorter answers, a different model and lower effort all change what the model sees or writes.
After routing, cost per request fell but cost per resolved conversation went up. What happened?
The cheaper model resolved fewer conversations: more follow-up messages, more handoffs to a human. Each request got cheaper, but each solved problem got more expensive. That’s why cost per completed task is the number to watch.
What to remember
- Price every request from its
usage: four fields, four rates. Log it next to the feature and conversation. - Judge changes by cost per completed task, not per request or per month.
- Caching and batching don’t change answers: do them first. The other levers need an eval.
- Output costs 5 times input on these models; ask for shorter answers, don’t truncate them.
- Caches are per model, have a minimum length and expire after 5 idle minutes. Routing interacts with all three.
- The saving from each lever depends on your traffic and on the order you apply them. Compute it on your own log.
What to study next
That’s Track 04. The capstone, Build It: A Small RAG Service That Ships, wires ingestion, retrieval with a minimum score, a versioned prompt, citations, evals, caching and cost logging into one small service you can run.
Further reading
- Anthropic: Pricing. Per-token prices for every model, including cache writes, cache reads and batch rates.
- Anthropic: Prompt caching. How prefixes match, cache lifetimes, per-model minimums and the
usagefields. - Anthropic: Batch processing. Limits, result types and how to handle errored and expired requests.
- Anthropic: Effort. The effort levels and how they trade depth for tokens.
- Anthropic: Optimizing for cost and intelligence. Anthropic’s own measured cost levers and cost-per-task comparisons between models.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.