State and Memory

Dana from Northwind Studio has been chatting with the Kitebase support agent for six messages. A colleague is locked out since their SSO moved to Okta, the reset link never arrives, she wants the findings posted on the ticket, then she asks about invoices and two-factor. By the last reply, each call to the model sends 2,013 tokens, and over half of them are help-center search results for questions answered ten minutes ago.

Two days later Dana writes “Hi again, is there any news on the Okta sign-in problem?”, and the agent has never heard of her. Both problems are about memory: what the agent carries inside one conversation, and what it keeps between them.

What you’ll build: the memory for a Kitebase support agent. You’ll count what each of 13 API calls in a six-turn chat sends and costs, shrink the history by clearing old tool results and summarising old turns without breaking the API’s rules, and keep a few short notes per customer that come back when they return. It all runs offline.

The history is the agent’s working memory

The model is stateless: it remembers nothing between calls. The agent loop gets around that with one list of messages that it resends on every call. That list is the agent’s working memory (or short-term memory): everything it knows about this conversation, and nothing else.

The worked example replays a recorded chat with Dana through the five tools used across this topic: search_help, get_ticket, search_tickets, assign_ticket and reply_to_customer. They read small local files, so every result is real text of a real size. Her first message becomes four messages:

[0] user       "Hi, Dana from Northwind Studio here. We moved our SSO to Okta yesterday ..."
[1] assistant  tool_use toolu_01 get_ticket("KITE-142"), tool_use toolu_02 search_help("locked out after SSO change")
[2] user       tool_result toolu_01 {"id": "KITE-142", "title": "Customer locked out after SSO change", ...}
               tool_result toolu_02 [account-recovery#2] Locked out of your account: If your workspace uses single sign-on ...
[3] assistant  "KITE-142 is already in progress with priya on our team. Because your workspace ..."

A tool_use block is the model asking your code to run a tool; a tool_result block is your answer, tied to it by the same id. After six customer turns the list has 26 messages, from 13 calls, one per assistant message.

Every call pays for the whole conversation again

Each call sends the system prompt, the five tool definitions and every message so far. The example estimates the size at 4 characters per token:

def request_tokens(messages: list[dict]) -> int:
    """What one call sends: system prompt + tool definitions + the whole history."""
    return (len(SYSTEM_PROMPT) + len(json.dumps(TOOLS)) + sum(chars(m) for m in messages)) // 4

python main.py replays the chat and prices each call at claude-opus-5’s $5 per million input tokens and $25 per million output tokens:

Every call sends the system prompt and 5 tool definitions first: 440 tokens.

1. Resend everything
  call  turn   sends  output     cost  what
     1     1     475      21  $0.0029  get_ticket, search_help
     2     1     730      68  $0.0053  reply
     3     2     812      15  $0.0044  search_help
     ...
    12     6   1,758      10  $0.0090  search_help
    13     6   2,013      45  $0.0112  reply
  Total: 16,541 input tokens, $0.0927. With prompt caching, about $0.0320.
  The final history is 1,618 tokens; 1,075 of them are tool results.
  • The 13 calls sent 16,541 input tokens for a conversation that ends at about 2,000. Each call pays for everything before it again, so the total grows with the square of the chat’s length: twice the calls, about four times the tokens.
  • 440 tokens ride on every call before a single message: the system prompt and tool definitions.
  • Tool results are two thirds of the history, mostly for questions already answered.
Input tokens each call sends, 6-turn chat with Northwind Studio Estimated at 4 characters per token. Each call includes the 440-token system prompt and tools. call turn what came back budget 1,400 before → after 1 1 get_ticket, search_help 475 2 1 reply 730 3 2 search_help 812 4 2 search_tickets 1,104 5 2 reply 1,152 6 3 get_ticket 1,242 7 3 reply 1,287 8 4 reply_to_customer 1,357 9 4 reply 1,428 → 942 10 5 search_help 1,468 → 982 11 5 reply 1,715 → 1,228 12 6 search_help 1,758 → 1,272 13 6 reply 2,013 → 1,331 cleared 4 results summarised 4 turns sent with trimming only sent without trimming all 13: 16,541 → 13,914 What call 13 sends without trimming: 2,013 tokens system + 5 tools: 440 customer: 136 agent calls + replies: 360 tool results: 1,075 Tool results are over half of call 13, and the part the agent is least likely to need again.
The worked chat, call by call. The trimming in the top half comes later in this article.

$0.09 a chat is about $93 a day at 1,000 chats, and a chat twice as long costs four times as much. The real number is a bit higher than the estimate, because the API adds its own instructions when you send tools. response.usage.input_tokens reports what you were billed, and client.messages.count_tokens(model=..., system=..., tools=..., messages=...) gives the exact count before you send.

claude-opus-5 takes a million tokens. Why worry about 2,000?

Because the context window, the most a model can take in one call, is a ceiling, not a target. A support chat won’t hit it; the bill hits you first, since every token in the history is paid for again on every call.

Long inputs also get used less well: models answer worse when the fact they need sits in the middle of a long input (see “Lost in the Middle” below). A stale search result is one more thing in the way.

Caching makes the resend cheaper, not smaller

The history only grows at the end, so each call starts with exactly the bytes of the previous one. Prompt caching reuses that: the provider keeps its processed version of a prompt’s start for a few minutes and bills a repeat at a tenth of the input price. Caching for LLM Apps covers how it works. For a growing conversation, one top-level field turns it on:

response = client.messages.create(
    model="claude-opus-5", max_tokens=1024, system=SYSTEM_PROMPT, tools=TOOLS,
    messages=messages,
    cache_control={"type": "ephemeral"},  # cache everything up to the last block
)

The example’s rough model of this brings the chat from $0.0927 to about $0.0320. Each call reads the previous call’s input at $0.50 per million and writes only the new part, at $6.25 per million (1.25 times the input price). Call 1 gets nothing: at 475 tokens it’s under the 512-token minimum claude-opus-5 will cache.

Turn it on; it’s the cheapest win here. But call 13 still sends 2,013 tokens, the model still reads every stale search result, and a long enough chat still fills the window. Caching changes the price of the history, not its size.

Step 1: Clear old tool results

The easiest tokens to cut are old tool results. Turn 1’s search results did their job when the agent answered turn 1, and if it needs them again it can call the tool again. So before a call, swap old results for a short placeholder:

def clear_block(b: dict) -> dict:
    if b["type"] != "tool_result" or b["content"].startswith("[cleared"):
        return b
    return {**b, "content": f"[cleared: {len(b['content']) // 4} tokens. Call the tool again if you need it.]"}


def clear_old_tool_results(messages, keep_last=KEEP_LAST):
    turns = split_turns(messages)
    old = [m for t in turns[:-keep_last] for m in t]
    recent = [m for t in turns[-keep_last:] for m in t]
    cleared = [m if is_customer_turn(m) or m["role"] == "assistant"
               else {**m, "content": [clear_block(b) for b in blocks(m)]} for m in old]
    ...
    return cleared + recent, n

split_turns groups the list by customer turn: a customer message plus every assistant message and tool result until the next one. KEEP_LAST = 2 leaves the turn in progress and the one before it alone, since that’s what the model is most likely to need.

The block stays, with its tool_use_id; only its content shrinks. Every tool_use must be answered by a tool_result in the next message or the API rejects the request, as the agent loop article showed. The placeholder also tells the model what happened, so it re-runs the tool instead of guessing.

With a budget of 1,400 tokens per call (tiny on purpose, so a six-turn chat reaches it), the first trim comes at call 9:

2. Trim when a call would send more than 1,400 tokens
  call  9: 1,428 -> 942 tokens, cleared 4 old tool results

A third of the call gone, with no model call and nothing lost that a tool can’t fetch again.

Step 2: Summarise old turns

Clearing only removes tool output. In a long chat the conversation itself grows, and old turns still hold things the agent needs: which ticket this is, what the customer already tried, what was promised.

Summarising (also called compaction) replaces old turns with a short summary written by a model call, and keeps the recent turns word for word:

def compact(messages, summarise, keep_last=KEEP_LAST):
    turns = split_turns(messages)
    if len(turns) <= keep_last:
        return messages
    old, recent = [m for t in turns[:-keep_last] for m in t], turns[-keep_last:]
    summary = summarise(old)
    first = recent[0][0]
    first = {"role": "user", "content": [
        {"type": "text", "text": f"<conversation_summary>\n{summary}\n</conversation_summary>"},
        *blocks(first)]}
    return [first] + recent[0][1:] + [m for t in recent[1:] for m in t]

The summary goes into the first kept customer message as an extra text block, not into the system prompt, so the system prompt stays identical on every call and keeps caching. The summariser gets a transcript of the old turns without the tool results, and this prompt:

Summarise this support conversation for the agent that will continue it.
Keep ticket ids, what the customer asked and already tried, and anything decided or promised.
Leave out greetings, raw tool output, email addresses and phone numbers.
Plain sentences, under 120 words.

The second line matters most: a summary that loses “priya is checking the SSO logs” makes the agent promise it all over again. Offline, the example uses a stand-in that only drops tool traffic and redacts contact details. With ANTHROPIC_API_KEY set, claude_summarise sends the prompt to claude-opus-5 and returns something like this (the wording changes between runs):

Dana (Northwind Studio) reports that since they moved SSO to Okta on Monday, a colleague
can't sign in. Ticket KITE-142, in progress with priya. The agent explained that SSO accounts
have no Kitebase password, so Forgot password can't help and the fix is on the Okta side.
Their IT admin confirmed the user is assigned in Okta and it still fails, so the agent posted
that on KITE-142 for priya to check Kitebase's SSO logs. Dana prefers email to phone calls.

About 110 tokens for four turns. At call 13 the budget is crossed again and clearing alone doesn’t get under it, so the example clears and then summarises turns 1 to 4:

  call 13: 2,013 -> 1,331 tokens, cleared 2 tool results, summarised 4 turns
  Total: 13,914 input tokens, $0.0796. With prompt caching, about $0.0419.
  Every tool_use still has its tool_result: True

Cut between turns, never inside one

The obvious way to trim is to keep the last N messages. It passes your tests, then fails on some chats and not others:

Before call 13: the last 8 of 25 messages, and two places to cut [0] to [17] turns 1 to 4: 18 messages, old tool results already cleared [18] user "Thanks. Unrelated, but where do I download last month's invoice?" [19] assistant tool_use toolu_07 search_help("download invoice") [20] user tool_result toolu_07 [invoices#1] Invoices and... [21] assistant "Go to Admin > Billing > Invoices and click the invoice date..." [22] user "Last thing: how does someone turn on two-factor?" [23] assistant tool_use toolu_08 search_help("turn on two-factor") [24] user tool_result toolu_08 [two-factor#1] Two-factor... compact() cuts here: between two customer turns messages[-5:] cuts here: between a tool_use and its result COMPACT(): 7 MESSAGES, CALL 13 SENDS 1,331 [0] user: summary of turns 1 to 4 + [18]'s question [1] to [6]: messages [19] to [24], unchanged toolu_07 and toolu_08 each still have their result MESSAGES[-5:] The list now starts with [20], a tool_result for toolu_07. Its tool_use was in [19], which is gone. The API answers 400. Some N work, some don't. Cut by turn.
Keeping the last N messages breaks whenever the cut lands between a tool call and its result.

messages[-5:] starts with a tool_result answering a tool_use the API never sees, so the request fails with a 400. messages[-3:], which starts at Dana’s last question, would have been fine. Which N works depends on how many tools the model called, and you don’t control that. A customer turn holds every tool call and result it made, so cutting between turns is always safe. The example’s pairs_intact checks the rule on a list; run it before every call in development and it catches the bug before the API does.

When to trim

manage runs before every call and does the cheapest thing that works:

def manage(history, summarise=stand_in_summarise, budget=BUDGET):
    if request_tokens(history) <= budget:
        return history, ""
    history, n = clear_old_tool_results(history)
    if request_tokens(history) <= budget:
        return history, f"cleared {n} old tool results"
    ...
    return compact(history, summarise), f"cleared {n} tool results, summarised {old_turns} turns"

Defaults for a support agent:

  • Trigger on tokens, not turns. One big search result can outweigh ten small turns. Pick the budget from cost: at $5 per million input tokens, 20,000 tokens caps each call’s input at $0.10.
  • Clear first, summarise second. Clearing is free and loses nothing a tool can’t fetch again. Summarising costs a call and can drop detail.
  • Keep the last two turns word for word. The model needs exact recent detail to act on.
  • Trim in one big step, then leave it alone. Look at the last output line again: with caching, trimming made this chat more expensive, $0.0419 against $0.0320. Each trim rewrites the start of the history, so the next call can’t read it from the cache and writes the whole prompt again at 1.25 times the price. Trim a little every turn and you pay that every turn.

So for short chats, caching alone is the better deal. Trim when chats run long, tool results are big, or the model starts losing track.

Long-term memory: notes outside the prompt

Now Thursday. Dana’s new chat starts with an empty history, which is right: Tuesday’s 26 messages are stale, mostly tool output, and include her phone number. But some of it should survive. Northwind moved SSO to Okta, KITE-142 is their open issue, and they prefer email.

That’s long-term memory: facts stored outside the prompt, in your own database, that outlive a conversation. You don’t paste old chats back in. You keep short notes, each with a topic, the text, the date written and an optional expiry, and at the start of each chat you pull back only the ones that matter.

Tuesday 2026-09-22, when the chat ends 1. PROPOSE At the end of the chat, a model call reads the transcript and suggests notes. The example has 5. 2. REMEMBER(): CHECK, THEN WRITE sso_provider replaces the Google note open_issue saved, expires in 30 days contact_preference saved: "email, not calls" contact_phone rejected: a phone number ticket_status saved, expires in 1 day 3. NOTES STORE Northwind Studio: 6 Harbor Pine: 1 a JSON file here, a table in production Thursday 2026-09-24, a new chat 4. QUESTION "Hi again, is there any news on the Okta sign-in problem?" 5. RECALL(): FILTER, THEN RANK only Northwind Studio's notes: 6 drop expired: ticket_status ended 09-23 rank by words shared with the question keep the top 3 that share any: 2 notes 6. FIRST MESSAGE <customer_notes> KITE-142 lockout moved SSO to Okta then the question 67 tokens, not the old chat
Write a few checked notes when the chat ends; read back only the relevant ones when the next one starts.

Writing. When a chat ends, one model call reads the transcript and suggests notes worth keeping. (The example has the five it might suggest in session.json, so it runs offline.) Your code decides what gets stored:

def remember(store, customer, topic, text, today, ttl_days=None) -> str:
    if redact(text) != text:
        return "rejected: looks like contact details"
    notes = store.setdefault(customer, [])
    # One note per topic, so a new fact replaces the old one instead of contradicting it.
    replaced = [n for n in notes if n["topic"] == topic]
    notes[:] = [n for n in notes if n["topic"] != topic]
    expires = (today + timedelta(days=ttl_days)).isoformat() if ttl_days else None
    notes.append({"topic": topic, "text": text, "written": today.isoformat(), "expires": expires})
    return f"replaced: {replaced[0]['text']}" if replaced else "saved"
3. Long-term memory
  sso_provider        replaced: Signs in with Google Workspace SSO.
  open_issue          saved
  contact_preference  saved
  contact_phone       rejected: looks like contact details
  ticket_status       saved

Without one note per topic, Thursday’s agent would get last November’s Google note and the new Okta one, and have to guess.

Reading. At the start of a chat, recall takes only this customer’s notes, drops expired ones, and ranks the rest by the words they share with the question:

def recall(store, customer, question, today, k=3) -> list[dict]:
    live = [n for n in store.get(customer, []) if is_live(n, today)]  # only this customer's notes
    q = words(question)
    scored = [(len(q & words(n["text"])), n["written"], n) for n in live]
    scored = [s for s in scored if s[0] > 0]
    scored.sort(key=lambda s: (s[0], s[1]), reverse=True)
    return [n for _, _, n in scored[:k]]
  2026-09-24: 6 notes on file, 1 expired ['KITE-142 is in progress with priya.'], 2 match the question.
  The new session's first message, 67 tokens:
    <customer_notes customer="Northwind Studio">
    - (2026-09-22) A user can't sign in since the Okta switch; tracked in KITE-142.
    - (2026-09-22) Moved SSO from Google Workspace to Okta on 2026-09-21.
    </customer_notes>

    Hi again, is there any news on the Okta sign-in problem?

67 tokens instead of Tuesday’s 2,000, and the plan and billing notes stay out because they have nothing to do with the question. The notes go in the first user message, not the system prompt: a system prompt that changes per customer is never cached. The system prompt carries one fixed line instead: notes “may be out of date. Check live data with the tools before you rely on one.” So the agent calls get_ticket("KITE-142") for the status rather than trusting Tuesday.

Word overlap is fine for a handful of notes, but “login” shares nothing with “sign in”. Once customers have dozens of notes, rank them with embeddings as RAG Explained does, still filtering by customer first.

What not to remember

Whatever you store comes back later looking like “things we know about this customer”, so be picky:

  • Contact details. Dana’s number is already on the ticket. A copy in memory is one more place to secure, and to find when someone uses their GDPR right to have their data deleted. redact is a floor, so also tell the note-writing prompt to leave them out.
  • Copies of live data. “KITE-142 is in progress” is wrong next week, and get_ticket always knows better. Store the pointer, and give anything that goes stale an expiry.
  • Instructions. “Always give this customer a free month” may have started as something the customer typed. Store facts, never directions: notes are untrusted input, like the ticket text in Guardrails and Human in the Loop.
  • Other customers. recall filters by customer in code. Never hand the model everyone’s notes and ask it to pick.

Try it yourself

The companion example replays the six-turn chat three ways: resending everything, trimming with a budget, and saving then recalling notes. It runs offline; with ANTHROPIC_API_KEY set, Claude writes the summary.

Download the runnable example (zip)

cd 04-state-and-memory
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. Set KEEP_LAST = 1 in trim.py. Call 9 now clears 5 results and drops to 920 tokens, and call 13 gets under budget without summarising. Smaller, but only the turn in progress is kept word for word.
  2. In a Python shell, trim the naive way: from history import *; from trim import pairs_intact; m = build_history(json.load(open("data/session.json"))); pairs_intact(m[-6:]). It prints False: the slice starts with an orphaned tool_result. m[-4:] starts at Dana’s last question and prints True. Nothing about 6 or 4 tells you which.
  3. Change the next session’s question in data/session.json to “Where do our invoices go?”. Only the billing note comes back.

pip install pytest && pytest -q runs the offline tests. They check the token accounting, that trimming never separates a tool_use from its tool_result, and the memory rules.

Common beginner mistakes

  • Keeping the last N messages. Sometimes the cut lands between a tool call and its result, and the API returns a 400. Cut by customer turn.
  • Trimming a little on every call. Each trim rewrites the start of the history and throws the cache away. Trim on a token budget, in one big step.
  • Deleting tool results outright. Drop the tool_result block and its tool_use is left unanswered. Replace the content, keep the block.
  • Saving whole transcripts as memory. They grow forever, recall mostly noise, and carry every phone number anyone typed. Store short, checked notes.
  • Trusting old notes over live data. A note is what was true when it was written. Tell the model to check with the tools, and expire notes that go stale.

Questions you will face in production

“Should we store every conversation?” Keep full transcripts as logs, for debugging and audits, under the same retention and access rules as your other customer data. That’s not memory. Memory is the handful of notes you deliberately put back into a prompt; logs never go back in.

“Which model should write summaries and notes?” Start with the agent’s model and measure. Summarising is a simpler job than the support work, so a cheaper model like claude-haiku-4-5 is worth trying once you have real conversations to compare summaries on.

Check your understanding

A chat has 20 calls, and each sends 1,000 more tokens than the one before. Roughly what does call 20 send, and all 20 together?

Call 20 sends about 20,000 tokens, plus the fixed system prompt and tools. All 20 send about 1,000 + 2,000 + … + 20,000 = 210,000 tokens, ten times the last call. That’s why the total grows much faster than the conversation, and why caching the repeated start pays.

You trim with `messages[-8:]` and it works in every test. In production some requests fail with a 400. Why?

In some chats the eighth-from-last message is a tool_result, so the trimmed list starts by answering a tool_use that was cut off. Your tests happened to cut somewhere else. Cut between customer turns, and check the pairs before each call.

Your agent tells a customer their ticket is "in progress", but it closed yesterday. The history was fresh. Where did the old status come from, and what do you change?

From long-term memory: a note that copied the status. Store a pointer (“tracked in KITE-142”) instead, give notes that go stale an expiry, and keep the system prompt line telling the model to check live data with get_ticket before relying on a note.

You add each customer's notes to the system prompt, and your cache hit rate drops to near zero. Why?

A cache hit needs the start of the prompt to match byte for byte, and the system prompt comes near the start. Different notes per customer means a different start for almost every chat. Put the notes in the first user message and keep the system prompt the same for everyone.

What to remember

  • The message history is the agent’s working memory, and every call sends all of it again. The total grows with the square of the chat’s length.
  • Turn on prompt caching first. It makes the resend cheap; it doesn’t make the history smaller.
  • Trim on a token budget: clear old tool results, then summarise old turns, keep the last two turns word for word, and trim in one big step.
  • Cut only between customer turns. A tool_use without its tool_result is a 400.
  • Long-term memory is a few short notes per customer, stored outside the prompt and recalled by relevance. Don’t store contact details, instructions or copies of live data.

What to study next

Memory keeps the agent coherent across a long chat. The next thing to break is the world it calls: tools time out, APIs return errors, calls hang. Retries, Timeouts, and Failure Handling covers keeping the loop alive when they do. Later, Durable Execution shows how saving the same message list lets an agent survive a crash.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.