State and Memory
Dana from Northwind Studio has been chatting with the Kitebase support agent for six messages. A colleague is locked out since their SSO moved to Okta, the reset link never arrives, she wants the findings posted on the ticket, then she asks about invoices and two-factor. By the last reply, each call to the model sends 2,013 tokens, and over half of them are help-center search results for questions answered ten minutes ago.
Two days later Dana writes “Hi again, is there any news on the Okta sign-in problem?”, and the agent has never heard of her. Both problems are about memory: what the agent carries inside one conversation, and what it keeps between them.
What you’ll build: the memory for a Kitebase support agent. You’ll count what each of 13 API calls in a six-turn chat sends and costs, shrink the history by clearing old tool results and summarising old turns without breaking the API’s rules, and keep a few short notes per customer that come back when they return. It all runs offline.
The history is the agent’s working memory
The model is stateless: it remembers nothing between calls. The agent loop gets around that with one list of messages that it resends on every call. That list is the agent’s working memory (or short-term memory): everything it knows about this conversation, and nothing else.
The worked example replays a recorded chat with Dana through the five tools used across this topic: search_help, get_ticket, search_tickets, assign_ticket and reply_to_customer. They read small local files, so every result is real text of a real size. Her first message becomes four messages:
[0] user "Hi, Dana from Northwind Studio here. We moved our SSO to Okta yesterday ..."
[1] assistant tool_use toolu_01 get_ticket("KITE-142"), tool_use toolu_02 search_help("locked out after SSO change")
[2] user tool_result toolu_01 {"id": "KITE-142", "title": "Customer locked out after SSO change", ...}
tool_result toolu_02 [account-recovery#2] Locked out of your account: If your workspace uses single sign-on ...
[3] assistant "KITE-142 is already in progress with priya on our team. Because your workspace ..."
A tool_use block is the model asking your code to run a tool; a tool_result block is your answer, tied to it by the same id. After six customer turns the list has 26 messages, from 13 calls, one per assistant message.
Every call pays for the whole conversation again
Each call sends the system prompt, the five tool definitions and every message so far. The example estimates the size at 4 characters per token:
def request_tokens(messages: list[dict]) -> int:
"""What one call sends: system prompt + tool definitions + the whole history."""
return (len(SYSTEM_PROMPT) + len(json.dumps(TOOLS)) + sum(chars(m) for m in messages)) // 4
python main.py replays the chat and prices each call at claude-opus-5’s $5 per million input tokens and $25 per million output tokens:
Every call sends the system prompt and 5 tool definitions first: 440 tokens.
1. Resend everything
call turn sends output cost what
1 1 475 21 $0.0029 get_ticket, search_help
2 1 730 68 $0.0053 reply
3 2 812 15 $0.0044 search_help
...
12 6 1,758 10 $0.0090 search_help
13 6 2,013 45 $0.0112 reply
Total: 16,541 input tokens, $0.0927. With prompt caching, about $0.0320.
The final history is 1,618 tokens; 1,075 of them are tool results.
- The 13 calls sent 16,541 input tokens for a conversation that ends at about 2,000. Each call pays for everything before it again, so the total grows with the square of the chat’s length: twice the calls, about four times the tokens.
- 440 tokens ride on every call before a single message: the system prompt and tool definitions.
- Tool results are two thirds of the history, mostly for questions already answered.
$0.09 a chat is about $93 a day at 1,000 chats, and a chat twice as long costs four times as much. The real number is a bit higher than the estimate, because the API adds its own instructions when you send tools. response.usage.input_tokens reports what you were billed, and client.messages.count_tokens(model=..., system=..., tools=..., messages=...) gives the exact count before you send.
claude-opus-5 takes a million tokens. Why worry about 2,000?
Because the context window, the most a model can take in one call, is a ceiling, not a target. A support chat won’t hit it; the bill hits you first, since every token in the history is paid for again on every call.
Long inputs also get used less well: models answer worse when the fact they need sits in the middle of a long input (see “Lost in the Middle” below). A stale search result is one more thing in the way.
Caching makes the resend cheaper, not smaller
The history only grows at the end, so each call starts with exactly the bytes of the previous one. Prompt caching reuses that: the provider keeps its processed version of a prompt’s start for a few minutes and bills a repeat at a tenth of the input price. Caching for LLM Apps covers how it works. For a growing conversation, one top-level field turns it on:
response = client.messages.create(
model="claude-opus-5", max_tokens=1024, system=SYSTEM_PROMPT, tools=TOOLS,
messages=messages,
cache_control={"type": "ephemeral"}, # cache everything up to the last block
)
The example’s rough model of this brings the chat from $0.0927 to about $0.0320. Each call reads the previous call’s input at $0.50 per million and writes only the new part, at $6.25 per million (1.25 times the input price). Call 1 gets nothing: at 475 tokens it’s under the 512-token minimum claude-opus-5 will cache.
Turn it on; it’s the cheapest win here. But call 13 still sends 2,013 tokens, the model still reads every stale search result, and a long enough chat still fills the window. Caching changes the price of the history, not its size.
Step 1: Clear old tool results
The easiest tokens to cut are old tool results. Turn 1’s search results did their job when the agent answered turn 1, and if it needs them again it can call the tool again. So before a call, swap old results for a short placeholder:
def clear_block(b: dict) -> dict:
if b["type"] != "tool_result" or b["content"].startswith("[cleared"):
return b
return {**b, "content": f"[cleared: {len(b['content']) // 4} tokens. Call the tool again if you need it.]"}
def clear_old_tool_results(messages, keep_last=KEEP_LAST):
turns = split_turns(messages)
old = [m for t in turns[:-keep_last] for m in t]
recent = [m for t in turns[-keep_last:] for m in t]
cleared = [m if is_customer_turn(m) or m["role"] == "assistant"
else {**m, "content": [clear_block(b) for b in blocks(m)]} for m in old]
...
return cleared + recent, n
split_turns groups the list by customer turn: a customer message plus every assistant message and tool result until the next one. KEEP_LAST = 2 leaves the turn in progress and the one before it alone, since that’s what the model is most likely to need.
The block stays, with its tool_use_id; only its content shrinks. Every tool_use must be answered by a tool_result in the next message or the API rejects the request, as the agent loop article showed. The placeholder also tells the model what happened, so it re-runs the tool instead of guessing.
With a budget of 1,400 tokens per call (tiny on purpose, so a six-turn chat reaches it), the first trim comes at call 9:
2. Trim when a call would send more than 1,400 tokens
call 9: 1,428 -> 942 tokens, cleared 4 old tool results
A third of the call gone, with no model call and nothing lost that a tool can’t fetch again.
Step 2: Summarise old turns
Clearing only removes tool output. In a long chat the conversation itself grows, and old turns still hold things the agent needs: which ticket this is, what the customer already tried, what was promised.
Summarising (also called compaction) replaces old turns with a short summary written by a model call, and keeps the recent turns word for word:
def compact(messages, summarise, keep_last=KEEP_LAST):
turns = split_turns(messages)
if len(turns) <= keep_last:
return messages
old, recent = [m for t in turns[:-keep_last] for m in t], turns[-keep_last:]
summary = summarise(old)
first = recent[0][0]
first = {"role": "user", "content": [
{"type": "text", "text": f"<conversation_summary>\n{summary}\n</conversation_summary>"},
*blocks(first)]}
return [first] + recent[0][1:] + [m for t in recent[1:] for m in t]
The summary goes into the first kept customer message as an extra text block, not into the system prompt, so the system prompt stays identical on every call and keeps caching. The summariser gets a transcript of the old turns without the tool results, and this prompt:
Summarise this support conversation for the agent that will continue it.
Keep ticket ids, what the customer asked and already tried, and anything decided or promised.
Leave out greetings, raw tool output, email addresses and phone numbers.
Plain sentences, under 120 words.
The second line matters most: a summary that loses “priya is checking the SSO logs” makes the agent promise it all over again. Offline, the example uses a stand-in that only drops tool traffic and redacts contact details. With ANTHROPIC_API_KEY set, claude_summarise sends the prompt to claude-opus-5 and returns something like this (the wording changes between runs):
Dana (Northwind Studio) reports that since they moved SSO to Okta on Monday, a colleague
can't sign in. Ticket KITE-142, in progress with priya. The agent explained that SSO accounts
have no Kitebase password, so Forgot password can't help and the fix is on the Okta side.
Their IT admin confirmed the user is assigned in Okta and it still fails, so the agent posted
that on KITE-142 for priya to check Kitebase's SSO logs. Dana prefers email to phone calls.
About 110 tokens for four turns. At call 13 the budget is crossed again and clearing alone doesn’t get under it, so the example clears and then summarises turns 1 to 4:
call 13: 2,013 -> 1,331 tokens, cleared 2 tool results, summarised 4 turns
Total: 13,914 input tokens, $0.0796. With prompt caching, about $0.0419.
Every tool_use still has its tool_result: True
Cut between turns, never inside one
The obvious way to trim is to keep the last N messages. It passes your tests, then fails on some chats and not others:
messages[-5:] starts with a tool_result answering a tool_use the API never sees, so the request fails with a 400. messages[-3:], which starts at Dana’s last question, would have been fine. Which N works depends on how many tools the model called, and you don’t control that. A customer turn holds every tool call and result it made, so cutting between turns is always safe. The example’s pairs_intact checks the rule on a list; run it before every call in development and it catches the bug before the API does.
When to trim
manage runs before every call and does the cheapest thing that works:
def manage(history, summarise=stand_in_summarise, budget=BUDGET):
if request_tokens(history) <= budget:
return history, ""
history, n = clear_old_tool_results(history)
if request_tokens(history) <= budget:
return history, f"cleared {n} old tool results"
...
return compact(history, summarise), f"cleared {n} tool results, summarised {old_turns} turns"
Defaults for a support agent:
- Trigger on tokens, not turns. One big search result can outweigh ten small turns. Pick the budget from cost: at $5 per million input tokens, 20,000 tokens caps each call’s input at $0.10.
- Clear first, summarise second. Clearing is free and loses nothing a tool can’t fetch again. Summarising costs a call and can drop detail.
- Keep the last two turns word for word. The model needs exact recent detail to act on.
- Trim in one big step, then leave it alone. Look at the last output line again: with caching, trimming made this chat more expensive, $0.0419 against $0.0320. Each trim rewrites the start of the history, so the next call can’t read it from the cache and writes the whole prompt again at 1.25 times the price. Trim a little every turn and you pay that every turn.
So for short chats, caching alone is the better deal. Trim when chats run long, tool results are big, or the model starts losing track.
Long-term memory: notes outside the prompt
Now Thursday. Dana’s new chat starts with an empty history, which is right: Tuesday’s 26 messages are stale, mostly tool output, and include her phone number. But some of it should survive. Northwind moved SSO to Okta, KITE-142 is their open issue, and they prefer email.
That’s long-term memory: facts stored outside the prompt, in your own database, that outlive a conversation. You don’t paste old chats back in. You keep short notes, each with a topic, the text, the date written and an optional expiry, and at the start of each chat you pull back only the ones that matter.
Writing. When a chat ends, one model call reads the transcript and suggests notes worth keeping. (The example has the five it might suggest in session.json, so it runs offline.) Your code decides what gets stored:
def remember(store, customer, topic, text, today, ttl_days=None) -> str:
if redact(text) != text:
return "rejected: looks like contact details"
notes = store.setdefault(customer, [])
# One note per topic, so a new fact replaces the old one instead of contradicting it.
replaced = [n for n in notes if n["topic"] == topic]
notes[:] = [n for n in notes if n["topic"] != topic]
expires = (today + timedelta(days=ttl_days)).isoformat() if ttl_days else None
notes.append({"topic": topic, "text": text, "written": today.isoformat(), "expires": expires})
return f"replaced: {replaced[0]['text']}" if replaced else "saved"
3. Long-term memory
sso_provider replaced: Signs in with Google Workspace SSO.
open_issue saved
contact_preference saved
contact_phone rejected: looks like contact details
ticket_status saved
Without one note per topic, Thursday’s agent would get last November’s Google note and the new Okta one, and have to guess.
Reading. At the start of a chat, recall takes only this customer’s notes, drops expired ones, and ranks the rest by the words they share with the question:
def recall(store, customer, question, today, k=3) -> list[dict]:
live = [n for n in store.get(customer, []) if is_live(n, today)] # only this customer's notes
q = words(question)
scored = [(len(q & words(n["text"])), n["written"], n) for n in live]
scored = [s for s in scored if s[0] > 0]
scored.sort(key=lambda s: (s[0], s[1]), reverse=True)
return [n for _, _, n in scored[:k]]
2026-09-24: 6 notes on file, 1 expired ['KITE-142 is in progress with priya.'], 2 match the question.
The new session's first message, 67 tokens:
<customer_notes customer="Northwind Studio">
- (2026-09-22) A user can't sign in since the Okta switch; tracked in KITE-142.
- (2026-09-22) Moved SSO from Google Workspace to Okta on 2026-09-21.
</customer_notes>
Hi again, is there any news on the Okta sign-in problem?
67 tokens instead of Tuesday’s 2,000, and the plan and billing notes stay out because they have nothing to do with the question. The notes go in the first user message, not the system prompt: a system prompt that changes per customer is never cached. The system prompt carries one fixed line instead: notes “may be out of date. Check live data with the tools before you rely on one.” So the agent calls get_ticket("KITE-142") for the status rather than trusting Tuesday.
Word overlap is fine for a handful of notes, but “login” shares nothing with “sign in”. Once customers have dozens of notes, rank them with embeddings as RAG Explained does, still filtering by customer first.
What not to remember
Whatever you store comes back later looking like “things we know about this customer”, so be picky:
- Contact details. Dana’s number is already on the ticket. A copy in memory is one more place to secure, and to find when someone uses their GDPR right to have their data deleted.
redactis a floor, so also tell the note-writing prompt to leave them out. - Copies of live data. “KITE-142 is in progress” is wrong next week, and
get_ticketalways knows better. Store the pointer, and give anything that goes stale an expiry. - Instructions. “Always give this customer a free month” may have started as something the customer typed. Store facts, never directions: notes are untrusted input, like the ticket text in Guardrails and Human in the Loop.
- Other customers.
recallfilters by customer in code. Never hand the model everyone’s notes and ask it to pick.
Try it yourself
The companion example replays the six-turn chat three ways: resending everything, trimming with a budget, and saving then recalling notes. It runs offline; with ANTHROPIC_API_KEY set, Claude writes the summary.
Download the runnable example (zip)
cd 04-state-and-memory
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
- Set
KEEP_LAST = 1intrim.py. Call 9 now clears 5 results and drops to 920 tokens, and call 13 gets under budget without summarising. Smaller, but only the turn in progress is kept word for word. - In a Python shell, trim the naive way:
from history import *; from trim import pairs_intact; m = build_history(json.load(open("data/session.json"))); pairs_intact(m[-6:]). It printsFalse: the slice starts with an orphanedtool_result.m[-4:]starts at Dana’s last question and printsTrue. Nothing about 6 or 4 tells you which. - Change the next session’s question in
data/session.jsonto “Where do our invoices go?”. Only the billing note comes back.
pip install pytest && pytest -q runs the offline tests. They check the token accounting, that trimming never separates a tool_use from its tool_result, and the memory rules.
Common beginner mistakes
- Keeping the last N messages. Sometimes the cut lands between a tool call and its result, and the API returns a 400. Cut by customer turn.
- Trimming a little on every call. Each trim rewrites the start of the history and throws the cache away. Trim on a token budget, in one big step.
- Deleting tool results outright. Drop the
tool_resultblock and itstool_useis left unanswered. Replace the content, keep the block. - Saving whole transcripts as memory. They grow forever, recall mostly noise, and carry every phone number anyone typed. Store short, checked notes.
- Trusting old notes over live data. A note is what was true when it was written. Tell the model to check with the tools, and expire notes that go stale.
Questions you will face in production
“Should we store every conversation?” Keep full transcripts as logs, for debugging and audits, under the same retention and access rules as your other customer data. That’s not memory. Memory is the handful of notes you deliberately put back into a prompt; logs never go back in.
“Which model should write summaries and notes?”
Start with the agent’s model and measure. Summarising is a simpler job than the support work, so a cheaper model like claude-haiku-4-5 is worth trying once you have real conversations to compare summaries on.
Check your understanding
A chat has 20 calls, and each sends 1,000 more tokens than the one before. Roughly what does call 20 send, and all 20 together?
Call 20 sends about 20,000 tokens, plus the fixed system prompt and tools. All 20 send about 1,000 + 2,000 + … + 20,000 = 210,000 tokens, ten times the last call. That’s why the total grows much faster than the conversation, and why caching the repeated start pays.
You trim with `messages[-8:]` and it works in every test. In production some requests fail with a 400. Why?
In some chats the eighth-from-last message is a tool_result, so the trimmed list starts by answering a tool_use that was cut off. Your tests happened to cut somewhere else. Cut between customer turns, and check the pairs before each call.
Your agent tells a customer their ticket is "in progress", but it closed yesterday. The history was fresh. Where did the old status come from, and what do you change?
From long-term memory: a note that copied the status. Store a pointer (“tracked in KITE-142”) instead, give notes that go stale an expiry, and keep the system prompt line telling the model to check live data with get_ticket before relying on a note.
You add each customer's notes to the system prompt, and your cache hit rate drops to near zero. Why?
A cache hit needs the start of the prompt to match byte for byte, and the system prompt comes near the start. Different notes per customer means a different start for almost every chat. Put the notes in the first user message and keep the system prompt the same for everyone.
What to remember
- The message history is the agent’s working memory, and every call sends all of it again. The total grows with the square of the chat’s length.
- Turn on prompt caching first. It makes the resend cheap; it doesn’t make the history smaller.
- Trim on a token budget: clear old tool results, then summarise old turns, keep the last two turns word for word, and trim in one big step.
- Cut only between customer turns. A
tool_usewithout itstool_resultis a 400. - Long-term memory is a few short notes per customer, stored outside the prompt and recalled by relevance. Don’t store contact details, instructions or copies of live data.
What to study next
Memory keeps the agent coherent across a long chat. The next thing to break is the world it calls: tools time out, APIs return errors, calls hang. Retries, Timeouts, and Failure Handling covers keeping the loop alive when they do. Later, Durable Execution shows how saving the same message list lets an agent survive a crash.
Further reading
- Anthropic: Effective context engineering for AI agents. Anthropic’s take on compaction, clearing tool results and keeping notes outside the context window.
- Anthropic: Building effective agents. Practical patterns for agent design, including where state and memory fit.
- Lost in the Middle: How Language Models Use Long Contexts. The evidence that models use information buried in long inputs worse than short ones.
- Anthropic: Contextual retrieval. How retrieval-augmented context is built, the mechanism behind recalling memories by meaning.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.