Context Management: What to Feed, What to Hide

Four tests fail in the Kitebase ticket service, so you tell your coding agent: “The tests are failing. Have a look around the repo and fix it.” It does exactly that. It opens all 14 files, including a 96 KB generated fixture, then runs the suite with pytest -v, which prints one line for each of 607 tests. Before it has changed anything, the session holds about 39,800 tokens. The bug is one line in a 46-line script.

Nothing is broken yet. But everything the agent read stays in the conversation and goes back to the model on every call, and Claude Code’s docs are blunt about the effect: “LLM performance degrades as context fills.” The agent that read everything gets slower, costs more, and has more places to go wrong.

What you’ll build: a script that adds up what one Kitebase session puts in the agent’s context (the rules file, each file read, each test run) and runs the same fix four ways, from about 53,000 estimated tokens down to about 1,000.

The worked example: four bad fixture tickets

Kitebase is the made-up project-tracking app used across this site. Its ticket service is the one from How AI Coding Tools Actually Work, with the --status filter in place and two new files. scripts/make_fixtures.py writes tests/fixtures/tickets.json: 600 tickets converted from an old tracker. tests/test_fixtures.py checks each one’s status is open, in_progress or closed, one test per ticket.

Four of them fail, because of one line in the generator:

# The old tracker's statuses, mapped to Kitebase's STATUSES.
LEGACY_STATUS = {"new": "open", "wip": "in_progress", "started": "In progress", "resolved": "closed"}

"In progress" should be "in_progress". Only four old tickets used started. The fix is one edit and a rerun of the script; what matters here is how much the agent reads on the way.

What fills the context window

The context window is the text the model sees on one call, measured in tokens (chunks of about 4 characters), as in What LLMs Are. A coding agent’s context is the whole session so far:

  • the tool’s own instructions and tool definitions
  • the rules file (CLAUDE.md here, which imports AGENTS.md)
  • your prompt
  • every file it reads
  • every command’s output, like a test run
  • its own replies and edits

Nothing drops out on its own. A file read on turn 2 is still there on turn 30, and the whole thing is sent again on every model call, as article 01 showed.

The companion script estimates the parts you control. Each thing the session adds is an Item, and the estimate is characters divided by 4:

def tokens(text: str) -> int:
    return len(text) // 4  # rough: about 4 characters per token for English and code

def sessions(repo: Path = REPO) -> dict[str, list[Item]]:
    ...
    return {
        "1. Read everything, -v test output": [rules, vague, *(read(repo, f) for f in everything), *verbose],
        "2. Same, with the deny rules": [rules, vague, *(read(repo, f) for f in allowed), *verbose],
        "3. Focused: 3 files, short test output": [rules, focused, *(read(repo, f) for f in FOCUSED_READS), *short],
        "4. Fresh session from a handoff note": [rules, handoff, read(repo, FOCUSED_READS[0]), short[1]],
    }

The test output isn’t made up: it’s saved from real runs of Kitebase’s tests, before and after the fix. Here’s session 1, the “have a look around” run, through to the fix:

1. Read everything, -v test output
  rules file    x1       267   CLAUDE.md + @AGENTS.md
  your prompt   x1        15   prompts/everything.txt
  files read    x14   25,399   biggest: tests/fixtures/tickets.json (23,896)
  test output   x2    27,583   pytest -v: 4 failed, 603 passed / pytest -v: 607 passed
  total               53,264   26.6% of a 200K window

The rules file and your prompt are under 300 tokens. One generated file is 23,896 and two test runs are 27,583. Big files and long output are nearly all of it, and the rest of this article cuts them:

rules, prompt, note files read generated fixture test output 1. Read everything 53,264 tokens tickets.json 23,896 2 x pytest -v 27,583 "Have a look around": 14 files, then the full -v log twice, failing and passing. 2. + deny rules 29,368 tokens 2 x pytest -v 27,583 tickets.json denied. 13 files left: 1,503 The deny rule removed the fixture. The logs are untouched. 3. Focused 1,817 tokens the 3 files the prompt names: 582 2 x pytest -q --tb=short: 868 4. Fresh from handoff 1,038 tokens CLAUDE.md 267, handoff note 146, make_fixtures.py 432, one pytest -q 193 Estimates at about 4 characters per token. Not counted: the tool's system prompt, tool definitions, the model's replies.
The same fix, four ways. The numbers are what the companion script prints.

How big is the window? At the time of writing, Claude Code runs current models on the Anthropic API with a 1 million token window; some setups run 200K. You rarely hit the wall on one task. The cost comes first: money, since each call re-sends everything, and attention, since models recall details less reliably as the context grows. Anthropic calls that context rot. Four FAILED lines are easy to find in 675 tokens and easier to lose in 14,154.

The estimate leaves out the tool’s own instructions, tool definitions and the model’s replies. For the real total in Claude Code, run /context: a grid of your usage by category, which CLAUDE.md files loaded, and suggestions for what to cut. Run it after a big read and you’ll soon know what things cost.

Point the agent at the right files

“Have a look around” is an instruction to read everything. On a real repo that’s dozens of files, and Claude Code’s docs name the failure: an unscoped investigation where Claude “reads hundreds of files, filling the context.”

The fix is to say where to look. Here’s the focused prompt from prompts/focused.txt:

4 of the 600 fixture tests fail: tickets KITE-1000, 1150, 1300 and 1450 have
status 'In progress', which isn't in STATUSES.

tests/fixtures/tickets.json is generated by scripts/make_fixtures.py. Read that
script, tests/test_fixtures.py and tickets/models.py. Don't open tickets.json.
Fix the script, run `python scripts/make_fixtures.py`, then
`python -m pytest -q --tb=short`. Done means 607 passed.

Three files, 582 tokens of reads instead of 25,399. The whole session comes to 1,817.

Three ways to point, in any tool:

  • Name the files in the prompt. In Claude Code, writing @tickets/models.py in a prompt makes Claude read that file before it responds. Other tools have their own syntax, like @ mentions in Cursor.
  • Point at the source, not the output. The generator is 432 tokens and holds the bug. The JSON it writes is 23,896 and only shows the symptom. Same goes for build output, compiled files and anything your code generates.
  • Send exploration somewhere else. When you don’t know where to look, a subagent (a helper agent with its own separate context window) can search and send back a summary. In Claude Code, ask for it: “Use a subagent to find where ticket statuses are validated.” The files it reads stay in its context, and only its answer lands in yours.

The gotcha: point too narrowly and the agent misses the file that matters. Read which files it opened, and add the missing one. And don’t type “don’t open tickets.json” into every prompt: facts true for every task, like which files are generated and which test command to use, belong in the rules file from Setting Up Context. Kitebase’s AGENTS.md now says both.

Trim the test output

Test runs are the other half of session 1. pytest -v prints one line per test, so both runs cost about the same, pass or fail:

RunLinesTokens (estimated)
python -m pytest -v, 4 failed66514,154
python -m pytest -v, all pass61513,429
python -m pytest -q --tb=short, 4 failed36675
python -m pytest -q --tb=short, all pass10193

The passing -v run spends 13,429 tokens to say “607 passed”. The short run says the same in 193, and when tests fail, it still keeps what the agent needs:

______________ test_fixture_ticket_has_a_known_status[KITE-1000] _______________
tests/test_fixtures.py:13: in test_fixture_ticket_has_a_known_status
    assert ticket["status"] in STATUSES, f"{ticket['id']} has status {ticket['status']!r}"
E   AssertionError: KITE-1000 has status 'In progress'
...
4 failed, 603 passed in 0.14s

Most test runners have a quieter mode, like pytest -q or go test without -v. Put the short command in the rules file, with the reason, and the agent will usually use it. Kitebase’s AGENTS.md says: “Don’t use -v: it prints a line for each of the 600 fixture tests.”

The same goes for what you paste. A 3,000-line CI log in your prompt stays in the context all session. Paste the failing test and the error, as in Prompting Coding Agents.

Doesn't Claude Code cut long output on its own?

Yes, by position. At the time of writing, a successful command’s output reaches Claude inline up to about 30,000 characters; past that, Claude Code saves it to a file and hands Claude the path and a 2,000-character preview. A failing command gets a head-and-tail excerpt of about 10,000 characters.

So a failing -v run costs about 2,500 tokens there, not 14,154. But the excerpt doesn’t know which lines matter, and it still costs nearly four times the short run. The bashOutputMaxChars setting changes the limit for successful output. Trimming before the output reaches the tool beats trimming after.

Two stronger options when a command is always noisy. Claude Code’s docs show a hook (a script the tool runs at a fixed point, here before every shell command) that rewrites test commands to keep only the failure lines. And you can hand the whole run to a subagent: “Use a subagent to run the test suite and report only the failing tests.” The log stays in its context; you get the summary.

Keep secrets and generated files out

Pointing at files is a habit. For some files you want a rule instead: huge generated files never worth reading, and secrets, like a .env with an API token, that must never reach the model. Each tool has its own mechanism, and none is airtight.

In Claude Code, the mechanism is a Read deny rule in .claude/settings.json, the settings file from Setting Up Context. Kitebase’s:

{
  "permissions": {
    "allow": ["Bash(python -m pytest *)", "Bash(python scripts/make_fixtures.py)"],
    "deny": ["Read(./.env)", "Read(./.env.*)", "Read(./tests/fixtures/**)"]
  }
}

The patterns use .gitignore syntax. Claude Code’s settings reference says matching files are excluded “from file discovery and search results”, reads are denied, and the Edit and Write tools are blocked on them too. Run the estimator with these rules and session 2 drops to 29,368 tokens:

The deny rules kept out: tests/fixtures/tickets.json (about 23,896 tokens)

Blocking edits is a bonus: a hand edit to a generated file would be overwritten anyway. The agent regenerates it with python scripts/make_fixtures.py, and the deny rule doesn’t stop that script from writing the file.

That’s also the gotcha. A deny rule covers Claude’s file tools and shell commands that name the file, like cat .env. It doesn’t cover grep -r TOKEN ., which reads every file without naming one, or a script that opens the file itself. Fine for the generator; not fine for secrets. For that, Claude Code’s docs point to its sandbox, which blocks the path for every process at the operating-system level. Better still, don’t keep live secrets in the working copy an agent runs in.

Other tools, at the time of writing:

ToolHow you keep files outWhat it doesn’t cover
Claude CodeRead(...) deny rules in .claude/settings.jsoncommands that don’t name the file, subprocesses
Cursor.cursorignore, on top of .gitignore and a built-in list (lock files, .env, binaries, dependency folders)the agent’s terminal and MCP tools; Cursor says protection isn’t guaranteed
GitHub Copilotcontent exclusion, set in the repo or org settings on GitHubagent mode in Copilot Chat in the IDE

Claude Code doesn’t use an ignore file: its old ignorePatterns setting is deprecated in favour of deny rules. Its respectGitignore setting only hides gitignored files from the @ file picker. Whether the agent’s own searches skip them depends on which search runs, so don’t count on .gitignore.

Clear, compact or start fresh

Say you ran session 1 anyway. The agent has found the bug and the context holds 39,835 estimated tokens, nearly all of it a fixture and a log you no longer need. You have three choices.

Keep going. Fine if the history is still useful, like when you’re deep in one hard problem. Claude Code also auto-compacts: when the context nears its limit, it replaces the conversation with a summary on its own.

/compact summarizes the conversation now. Give it a focus so the summary keeps what you need: /compact keep the 4 failing ticket ids and the fix plan. Claude Code’s docs list what survives:

  • The system prompt still applies, and the project-root CLAUDE.md is re-read from disk.
  • Up to five recently used files are re-read. A file over 5,000 tokens comes back as a path only.
  • Nested CLAUDE.md files and path-scoped rules are summarized away, and reload only when the agent next reads a file they cover. A rule that must survive belongs in the root file.
  • Everything else, including the exact test output, is only what the summary kept.

You can steer compaction from the rules file too. Kitebase’s CLAUDE.md:

@AGENTS.md

## Compact instructions
When compacting, keep the files changed so far, the test command, and the names of any tests still failing.

/clear empties the context and starts a new conversation. The old one isn’t deleted; /resume brings it back. Use it between unrelated tasks, and after two failed corrections on the same problem, the rule from Prompting Coding Agents. /compact is itself one big request, since the model has to read everything to summarize it. /clear costs nothing.

SESSION 1 SO FAR: 39,835 CLAUDE.md + AGENTS.md 267 "have a look around" 15 14 files, with tickets.json 25,399 pytest -v: 4 failed 14,154 + the model's replies The bug is found. Most of this is noise. /compact /clear SAME SESSION, COMPACTED system prompt, CLAUDE.md, AGENTS.md: re-read from disk a summary the model writes of everything else up to 5 recent files re-read; any over 5,000 tokens as a path the exact test output: gone, only what the summary kept Keeps the thread. The model decides what the summary holds. NEW SESSION FROM A HANDOFF NOTE: 1,038 CLAUDE.md + AGENTS.md 267 @HANDOFF.md: 15 lines you read before clearing 146 reads scripts/make_fixtures.py, fixes line 21 432 pytest -q --tb=short: 607 passed 193 You decide what carries over. The dead ends stay behind.
Two ways out of a heavy session. Compaction keeps the thread; clearing keeps only what you carry over.

Two smaller tools: /btw asks a side question whose answer never enters the conversation, and /rewind (or Esc twice) can summarize from a chosen message onwards, leaving the earlier part word for word.

The default: /clear when the task changes; /compact with a focus when you’re mid-task and the history still matters.

Hand off with a short summary

After /clear, the new conversation knows nothing the old one learned. A handoff note carries that part over: a short file the old session writes and the new one reads. Before clearing, ask for it. From prompts/handoff.txt:

Before I clear this session, write a handoff note to HANDOFF.md for a fresh
session: the goal, what you found (file and line), what's done, what's left,
and the exact commands to check it. Under 15 lines. No pasted code or logs.

The companion’s session/handoff.md is an example of a good one, written by hand for this article:

# Handoff: 4 failing fixture tests

Goal: `python -m pytest -q --tb=short` passes (607 tests).

Found: tests/fixtures/tickets.json is generated by scripts/make_fixtures.py.
LEGACY_STATUS (line 21) maps "started" to "In progress". STATUSES has "in_progress".
Only 4 tickets use "started": KITE-1000, 1150, 1300, 1450.

Done: nothing edited yet.
...
Don't open or hand-edit tickets.json. It's regenerated, and it's about 24,000 tokens.

It’s 146 tokens. The new session reads it, reads the generator, fixes line 21 and runs the short tests: 1,038 tokens in all, against 53,264 for session 1 carried through to the end.

Two habits make handoffs work:

  • Read the note before you clear. It’s the model’s summary of its own work, and a wrong conclusion in it (“the bug is in search.py”) becomes the new session’s starting fact. It takes 30 seconds to check.
  • Start the new session from it. /clear, then a first prompt like “Continue from @HANDOFF.md.” Delete the file when the task is done, or add HANDOFF.md to .gitignore so it never gets committed.

A plain file also works across days and across tools. To get the whole conversation back instead, claude --continue reopens the last session in the folder and claude --resume lets you pick one, junk included.

Try it yourself

The companion example is the Kitebase repo with its four failing tests, the prompts from this article, the captured test output, and the estimator. It runs offline and needs no API key.

Download the runnable example (zip)

cd 05-context-management
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. echo "KITEBASE_API_TOKEN=test-only" > kitebase/.env and run python main.py again. Session 1 now reads .env; session 2 doesn’t, because Read(./.env) denies it. Delete the file afterwards.
  2. Remove "Read(./tests/fixtures/**)" from kitebase/.claude/settings.json and rerun. Session 2 grows back to nearly the size of session 1.
  3. Open Claude Code in kitebase/ and paste prompts/everything.txt. When it’s done, run /context. Then /clear, paste prompts/focused.txt, and compare.

pytest -q runs the offline tests, including one that applies the fix on a temporary copy and checks all 607 tests pass.

Common beginner mistakes

  • “Have a look around” as a prompt. It tells the agent to read everything. Name the files, or send the search to a subagent.
  • Letting it read generated files. Fixtures, lockfiles, build output and vendored code are big and rarely the source of the bug. Deny or ignore them, and point at what generates them.
  • Verbose output by default. -v on a big suite costs the same when everything passes. Put the quiet command in the rules file.
  • Trusting an ignore rule with secrets. Deny rules and .cursorignore stop the file tools, not every command. Keep live secrets out of the working copy, or use the sandbox.
  • One session for everything, or clearing with no handoff. Stale context from the last task is noise; a bare /clear makes the next session rediscover what this one knew. Clear between tasks, with a short note you’ve read.

Questions you will face in production

“Does context size matter if the window is 1 million tokens?” Yes. The window is how much fits, not how well the model uses it. A bloated session costs more on every call and recalls detail less reliably. Treat 1M as headroom, not a target.

“Should the team commit deny rules?” Yes: put .env files, secrets folders and big generated paths in the project’s .claude/settings.json (or .cursorignore) so every teammate gets them. They’re a sensible default, not a security boundary. Team Workflows and Guardrails covers the rest of the shared setup.

“Auto-compact kicked in and the agent forgot a rule. What happened?” Likely it was in a nested CLAUDE.md or a path-scoped rule. Those are summarized away and come back only when the agent reads a matching file again. Move rules that must always apply into the root CLAUDE.md, and add a compact instructions section for what the summary must keep.

Check your understanding

A session feels slow and the agent has started repeating a mistake you already corrected. What do you check, and what do you do?

Run /context to see what’s filling it; usually big file reads or long command output. If the task has changed or you’ve corrected the same thing twice, ask for a short handoff note, read it, /clear, and start from the note. If you’re mid-task and the history matters, /compact with a focus instead.

Your repo has a 40,000-token generated OpenAPI client that agents keep opening. What's the fix?

Keep it out and point at its source. In Claude Code, add a Read deny rule for its folder in .claude/settings.json; in Cursor, add it to .cursorignore. Then say in the rules file which spec file or script generates it and the command to regenerate. The agent reads the small source instead of the big output.

You added Read(./.env) to the deny list. Can the agent still see your token?

It can’t read .env with its file tools or with cat .env. It could still reach it through a command that doesn’t name the file, like grep -r TOKEN ., or a script that opens the file itself. For a hard guarantee, use the sandbox, or keep the real token out of that working copy.

Why is pytest -q --tb=short better than letting the tool truncate a -v log?

Quiet mode drops the lines you don’t need (one per passing test) and keeps the ones you do: each failure, the assertion, the summary. A tool’s cut is by position, so it keeps the start and end of the log whatever is in them. Here the short run is 675 tokens and has all four failures.

What to remember

  • Everything the agent reads or runs stays in its context for the rest of the session and is re-sent on every call. Big files and long output are nearly all of it.
  • Name the files, point at sources instead of generated output, and send open-ended searches to a subagent.
  • Use quiet test commands and paste only the failing lines. Put both in the rules file.
  • Keep huge and secret files out with deny rules or ignore files, and know what they don’t cover.
  • /clear between tasks, /compact with a focus mid-task, and hand off through a short note you’ve read.

What to study next

Good context makes the agent’s first try much more likely to be right, but not certain. Reviewing and Trusting AI-Written Code is about what to do with the diff: the mistakes agents make, and how to check them.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from; the Claude Code details were checked against version 2.1.281 and its docs. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.