Caching for LLM Apps
The Kitebase support bot from article 04 now answers 10,000 questions a day. Two things repeat in the logs. Every request sends the same 859 tokens of rules and help articles, paid in full each time. And “How do I reset my password?” arrives hundreds of times a day, and the model writes a fresh answer to every one.
That’s about $81 a day, mostly for work you’ve already paid for once. Caching stops you paying twice. It’s also where some of the quietest bugs in LLM apps live.
What you’ll build: a response cache for the Kitebase bot that normalises questions, expires answers and forgets them when an article changes; prompt caching on its system prompt, and the timestamp that silently switches it off; and a semantic cache serving a wrong answer. It all runs offline.
Where the money goes
A cache stores the result of expensive work under a key made from its inputs. When the same inputs come back, you return the stored result instead of redoing the work. It’s memoising a function, with an expiry date.
Kitebase’s help center is six short articles, so this version of the bot skips retrieval and puts all six in the system prompt (the instructions sent at the start of every request, from article 02). One question costs:
| Part | Tokens | Price per million (claude-opus-5) | Cost |
|---|---|---|---|
| System prompt: rules + 6 articles | 859 | $5 input | $0.00430 |
| The question | 6 | $5 input | $0.00003 |
| The answer | 150 | $25 output | $0.00375 |
| Total | $0.00808 |
Token counts are estimated at 4 characters per token, and 150 is a typical answer for this bot. At 10,000 questions a day, $80.75.
Two parts of that repeat, and each gets its own cache:
- Whole questions repeat. A response cache stores the final answer, keyed by the question. A hit skips the model entirely.
- The start of every prompt repeats. Prompt caching happens at the provider: Anthropic keeps its work on the 859-token prefix for a few minutes and charges a tenth of the input price to reuse it.
A third kind, the semantic cache, catches the same question in different words. It comes last, because it can hurt you.
Step 1: A response cache with a good key
The simplest response cache is a dict from question to answer. The hard part is the key. Customers type How do I reset my password?, how do i reset my password and How do I reset my PASSWORD. Use the raw string and that’s three misses for one question.
So you normalise first: rewrite the question into one standard form. The rule is strict: only changes that can’t change the meaning.
def normalise(question: str) -> str:
# Only changes that can't change the meaning: case, spacing, end punctuation.
text = unicodedata.normalize("NFKC", question).lower()
return " ".join(text.split()).rstrip("?!. ")
NFKC turns look-alike Unicode characters, like a non-breaking space, into plain ones. Then it lowercases, collapses whitespace and drops the punctuation at the end.
The question isn’t the only input to the answer. A different model, or an edited rule or help article, gives a different answer to the same question. So they go in the key too:
def cache_key(question: str, system: str, model: str = MODEL) -> str:
# Everything that can change the answer goes in the key. The system prompt
# holds the rules and the articles, so editing either gives new keys.
parts = {"model": model, "prompt": fingerprint(system), "q": normalise(question)}
blob = json.dumps(parts, sort_keys=True)
return "kb-answer:" + hashlib.sha256(blob.encode()).hexdigest()[:12]
fingerprint is the first 8 characters of a SHA-256 hash of the system prompt, the trick article 03 uses to tell prompt versions apart. sort_keys=True makes the same inputs always serialise, and so hash, the same way. The companion code prints:
stored "How do I reset my password?" kb-answer:01cc5143cfe0
" how do I reset my PASSWORD" kb-answer:01cc5143cfe0 hit
"How do I reset my password if I never..." kb-answer:8a25c1366512 miss
The messy spelling hits. The locked-out customer from article 04, who needs a different answer, misses.
The gotcha is normalising too much to get more hits. Sort the words and “Move my tasks from Alpha to Beta” shares a key with “Move my tasks from Beta to Alpha”. Drop “on” and “off” as filler and “How do I turn off two-factor?” gets the turn-on instructions. A wrong hit has no model in the loop to notice. When in doubt, don’t normalise: a miss costs a model call, a wrong hit costs a customer.
The same goes for anything personal. If the prompt includes the customer’s plan or account data, that goes in the key, or the request skips the cache. Otherwise one customer’s answer is served to the next.
Step 2: TTLs and invalidation
A cached answer is right only while its inputs stay the same. Two tools keep it honest.
A TTL (time to live) is how long an entry stays valid. After that, the lookup misses and the answer is regenerated:
class ResponseCache:
...
def get(self, key: str) -> str | None:
item = self._store.get(key)
if item is None or self.clock() >= item[0]:
self._store.pop(key, None)
self.misses += 1
return None
self.hits += 1
return item[1]
def set(self, key: str, answer: str) -> None:
self._store[key] = (self.clock() + self.ttl, answer)
clock is passed in, so the demo and tests can move time forward without waiting. In production this is Redis, an in-memory key-value store most teams already run: SET key answer EX 86400 stores a value that Redis deletes after 86,400 seconds.
Invalidation means making stale entries stop being used, and the fingerprint already does it. Edit reset-password.md so the link lasts 60 minutes instead of 30, and every key changes:
same question, 25 hours later kb-answer:01cc5143cfe0 miss
same question, reset-password.md edited kb-answer:23a7138af9f6 miss
The old entries aren’t deleted. Nothing looks them up any more, and the TTL clears them within a day. The key handles changes you can see; the TTL catches the ones you can’t, like a policy that changed before anyone edited the help center. Start with 24 hours for help-center answers. Go shorter when answers depend on things that change during the day, and don’t cache answers built from live account data.
Only cache answers you’d serve again:
def answer(client, cache: ResponseCache, system: str, question: str) -> tuple[str, bool]:
key = cache_key(question, system)
if (cached := cache.get(key)) is not None:
return cached, True
response = ask_claude(client, system, question)
text = "".join(b.text for b in response.content if b.type == "text")
if response.stop_reason == "end_turn": # never cache a refusal or a cut-off answer
cache.set(key, text)
return text, False
A stop_reason of "max_tokens" means the answer was cut off; "refusal" means the model declined. Cache either and everyone who asks gets the broken reply for a day.
The same pattern gives you the embedding cache article 04 mentioned: key each chunk’s vector by the model name plus a hash of the chunk’s text, and re-indexing only embeds changed chunks. It needs no TTL: same text, same model, same vector.
Step 3: Prompt caching at the provider
Every response-cache miss still sends the same 859 tokens of rules and articles at $5 per million.
With prompt caching, the provider keeps its processed version of the start of your prompt for a few minutes. When the next request starts with exactly the same bytes, it reuses that work and bills those tokens at $0.50 per million on claude-opus-5, a tenth of the price.
Why does it have to be the start of the prompt?
The model reads tokens in order, and its work on each token depends on every token before it. So the work for token 500 is reusable only if tokens 1 to 499 are identical too.
Change one character near the start and everything after it is recomputed, even if the rest is the same. A shared start can be reused. A shared middle can’t.
You mark where the reusable part ends with a breakpoint, a cache_control field on a content block. Everything up to and including that block is cached. For the Kitebase bot, that’s the system prompt, the last thing every request shares:
def ask_claude(client, system: str, question: str):
return client.messages.create(
model=MODEL,
max_tokens=1024,
# The breakpoint goes on the last block every request shares, not on the question.
system=[{"type": "text", "text": system, "cache_control": {"type": "ephemeral"}}],
messages=[{"role": "user", "content": question}],
)
The system prompt becomes a list of blocks so one can carry cache_control. A top-level cache_control puts the breakpoint on the last block for you, which suits long conversations. Here the last block is the question, so every request would write an entry nobody reads.
The response reports what happened in usage: cache_creation_input_tokens were written to the cache, cache_read_input_tokens were read from it, and input_tokens is only the rest, after the breakpoint. The example’s offline model of the cache prints what those fields would say for two questions a minute apart, with costs including a 150-token answer:
stable system prompt:
request 1: write 859 read 0 cost $0.00915
request 2: write 0 read 859 cost $0.00421
Uncached, each costs $0.00808. The first costs more, because writing an entry is billed at 1.25 times the input price. After that, each costs about half. python main.py --live sends two real questions and prints the real fields; the counts will differ a little from the estimate.
Three rules trip people up:
- Entries expire after 5 minutes without a read. Each read resets the timer, so steady traffic keeps it alive all day.
{"type": "ephemeral", "ttl": "1h"}lasts an hour, but its writes cost 2 times the input price. - Short prompts don’t cache, and nothing tells you. The minimum is 512 tokens on
claude-opus-5but 4,096 onclaude-haiku-4-5. Move this bot to Haiku and every request quietly shows zero writes and zero reads. - Parallel requests don’t share a fresh entry. An entry becomes readable once the request that wrote it starts streaming its answer. Fire 50 identical requests at once and none of them reads the cache.
OpenAI caches long prefixes automatically and reports them in usage.prompt_tokens_details.cached_tokens. Same prefix rule.
Step 4: The timestamp that turns it off
A month later, someone wants the bot to tell customers whether an invoice is overdue. The model needs today’s date, so they add a line to the top of the system prompt:
def with_timestamp(system: str, now: str) -> str:
return f"Current time: {now}\n\n{system}" # the mistake this demo is about
Every test passes and the answers look fine. The cache hit rate drops to zero, and nothing says so:
timestamp at the top:
request 1: write 868 read 0 cost $0.00920
request 2: write 868 read 0 cost $0.00920
The two stamped prompts share their first 29 of 3,473 characters: 'Current time: 2026-09-27 14:0'
Response cache keys differ too: kb-answer:d1bed0fcd520 vs kb-answer:2cbea15f5528
The prompts differ only in the timestamp, and it comes first. Every request writes a new entry and reads nothing, at $0.00920 a question: worse than no caching. The last line is the second casualty. The timestamp changes the system prompt’s fingerprint, so the response cache never hits either.
The fix: keep the system prompt frozen, and put anything that changes after the breakpoint, here at the start of the user message: f"Today is {date}.\n\n{question}". The response cache key stays date-free, so don’t cache answers that depend on the date, like whether an invoice is overdue.
Anything that changes near the start of the prompt does the same:
- A request ID or
uuid4()in the system prompt. json.dumps(data)withoutsort_keys=True, or looping over aset: same data, different byte order.- The customer’s name or plan in the system prompt.
- A tool list that varies per request. Tools come before the system prompt.
- Switching model. Each model has its own cache.
Requests keep succeeding, so watch the numbers. If cache_read_input_tokens is 0 on repeated requests, diff two rendered prompts: the first difference is the culprit. Better, add a check that the second of two identical requests shows a cache read.
What it all saves
A day of 10,000 questions:
3. A day of 10,000 questions (859-token prefix, 6-token question, 150-token answer)
no caching $ 80.75
prompt caching $ 42.09 (48% less)
+ response cache, 20% repeats $ 33.68 (58% less)
Prompt caching takes input from $43.25 to under $5. Cache writes are left out: steady traffic keeps the entry warm, so a write only happens after a quiet gap of 5 minutes or more, and costs about half a cent more than a read.
Now output is $37.50 of the $42.09, and prompt caching can’t touch it. Only a response-cache hit skips output. The 20% repeat rate is an assumption, not a benchmark: log your cache keys for a week and count the repeats before you trust any number here. Other ways to cut output, like shorter answers and cheaper models, are in article 12: Cost Optimization.
At 500 help articles you’re back to retrieving chunks, as in article 04. Then the fixed rules go before the breakpoint and the chunks and question after it.
Semantic caching, and why to be careful
The response cache misses “How can I turn on two-factor?” when it has “How do I turn on two-factor?” stored.
A semantic cache looks answers up by meaning. It embeds the new question (turns it into a vector, as in article 04), compares it with questions it has already answered, and returns the stored answer if the closest scores above a threshold, a minimum cosine similarity you choose. The example uses the same word-counting stand-in embedder:
def semantic_lookup(entries: list[tuple[list[float], str]], question: str) -> tuple[float, str | None]:
q = embed(question)
best, found = 0.0, None
for vec, cached_answer in entries:
score = sum(a * b for a, b in zip(q, vec)) # both have length 1, so this is the cosine
if score > best:
best, found = score, cached_answer
return best, (found if best >= SEMANTIC_THRESHOLD else None)
With “How do I turn on two-factor?” cached and a threshold of 0.85:
How can I turn on two-factor? 1.00 hit: Go to Settings > Security > Two-factor and scan the QR code.
How do I turn off two-factor? 0.87 hit: Go to Settings > Security > Two-factor and scan the QR code.
The first hit is the point. The second is the risk: asked how to turn two-factor off, the bot explains turning it on. The help center doesn’t cover turning it off, so the right answer was “I don’t know”, and the model never saw the question to say so.
The stand-in exaggerates, since it drops “on” as filler. A real embedding model keeps both words, but it measures how related two texts are, and these two are very related. Only testing tells you whether such a pair clears your threshold, which has to be picked per embedding model.
My default is to leave semantic caching out. Ship the other two caches and read your logs. If they show lots of paraphrased repeats, match against a short list of FAQ answers a person wrote and checked, not whatever the model said last time. Test the threshold on look-alike pairs, and log every semantic hit for review.
Try it yourself
The companion example runs every demo in this article offline.
Download the runnable example (zip)
cd 10-caching-for-llm-apps
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
- Change
with_timestampto returnsystemunchanged, as if the time had moved into the user message. Request 2 reads 859 tokens again. - Set
MIN_CACHEABLE_TOKENS = 4096, the minimum forclaude-haiku-4-5. Both requests showwrite 0 read 0and cost $0.00808: nothing cached, no error. - Raise
SEMANTIC_THRESHOLDinsemantic.pyto0.9. The wrong two-factor answer goes away, and so will some right hits.
pip install pytest && pytest -q runs the offline tests for keys, TTL expiry, invalidation and the timestamp demo. They need no keys and no network.
Common beginner mistakes
- Leaving the version out of the key. Key on the question alone and you serve old answers after fixing the prompt or articles. Add the model and a prompt fingerprint.
- Normalising meaning away. Stripping stopwords or sorting words merges questions that need different answers.
- Caching broken answers. Store only
end_turnresponses: no refusals, cut-offs or errors. - Not measuring. Without a hit rate and
cache_read_input_tokensin your metrics, you won’t notice when a change quietly turns caching off.
Questions you will face in production
“What’s our hit rate?”
Log a cache_hit flag per request for the response cache, and sum the usage cache fields for prompt caching. Keep both on a dashboard: the expensive failures are regressions months after launch.
“Why is the hit rate low when the prompt looks identical?” Diff two rendered requests and find the first difference. Also check the prefix is above the model’s minimum, and that traffic isn’t split across API workspaces, which don’t share prompt caches.
Check your understanding
You edited the reset article to say 60 minutes, and the bot still says 30. What's the likely bug?
The response cache key doesn’t include a fingerprint of the prompt, so the edit changed no keys. Add it; the old entries stop matching and the TTL clears them.
After a deploy, every request shows cache_read_input_tokens of 0 and cache_creation_input_tokens of about 870. What happened?
Something before the breakpoint changes between requests, so each writes a new entry and none reads one. Diff two rendered system prompts and look for a timestamp, a request ID or unsorted JSON.
You move the bot to claude-haiku-4-5 to save money. Prompt caching shows no writes and no reads, and no error. Why?
The prefix is about 859 tokens, and claude-haiku-4-5 only caches prefixes of 4,096 tokens or more. Shorter ones are silently skipped. On claude-opus-5 the minimum is 512.
Someone wants the customer's plan in the system prompt so answers can mention plan limits. What happens to both caches?
Prompt caching is now shared only within a plan, and the response cache key must include the plan or a Free customer gets a Business answer. Put the plan in the user message, after the breakpoint, and the shared prefix survives. The response cache key still needs the plan.
What to remember
- Two things repeat: whole questions and the start of every prompt. A response cache handles the first, prompt caching the second.
- Put everything that can change the answer in the key, normalise only what can’t change the meaning, and set a TTL for what the key can’t see.
- Prompt caching is a byte-for-byte prefix match. Put the breakpoint on the last block every request shares, and anything that changes after it.
- Reads cost 0.1x the input price, writes 1.25x, and prompts under the model’s minimum silently don’t cache. Check
cache_read_input_tokens. - Prompt caching cuts input only. Output is often most of what’s left.
- Semantic caches serve confident wrong answers to look-alike questions. Leave them out until your logs justify one.
What to study next
Most failures in this article, a hit rate that drops to zero or a stale answer, are silent. Seeing them takes traces and metrics on every request: article 11: Production AI Observability. Then article 12: Cost Optimization goes after the output cost caching leaves behind.
Further reading
- Anthropic: Prompt caching. The official docs: breakpoints, TTLs, minimum lengths per model and pricing multipliers.
- OpenAI: Prompt caching. OpenAI’s automatic caching, which prefixes get cached and how to see hits.
- GPTCache: semantic caching library. An open-source semantic cache, useful to read even if you build your own.
- AWS: Bedrock prompt caching. How prompt caching works on Bedrock, for comparison.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.