How AI Coding Tools Actually Work

You ask a coding agent to “add a --status filter to search_tickets”. A minute later it hands you a tidy diff across four files. It read the code first, ran the tests, saw three of them fail, fixed its own bug, added two tests, and reported “All 6 tests pass.” Then you try it:

$ kitebase search sso --status archived
No tickets found.

There’s no status called archived. The project’s rules file says an unknown status must raise an error, and the tool had that rule in front of it the whole time. The same run that caught one bug shipped another.

Both halves come from the same machinery. This article takes it apart: what the tool sends to the model, the loop that reads, edits and runs code, and why that loop is good at some jobs and bad at others.

What you’ll build: a toy coding agent, about 270 lines of Python, that works on a tiny Kitebase ticket service. It gathers context, searches the code, reads five files, edits two, runs the tests, fixes the failure and prints the diff. It runs offline with a scripted model, but the files and the tests are real.

The worked example: a tiny ticket service

Kitebase is the made-up project-tracking app used across this site. Its ticket service is ten files:

kitebase/
├── AGENTS.md            # the rules file
├── pyproject.toml
├── tickets/
│   ├── models.py        # Ticket dataclass, STATUSES = ("open", "in_progress", "closed")
│   ├── data.py          # 8 sample tickets, KITE-138 to KITE-145
│   ├── search.py        # search_tickets()
│   ├── cli.py           # the `kitebase search <query>` command
│   └── __init__.py, __main__.py
└── tests/
    ├── test_search.py   # 3 tests
    └── test_cli.py      # 1 test

The function the task is about:

def search_tickets(tickets: list[Ticket], query: str) -> list[Ticket]:
    """Tickets whose title or customer contains every word of the query."""
    words = query.lower().split()
    return [
        t for t in tickets
        if all(w in f"{t.title} {t.customer}".lower() for w in words)
    ]

kitebase search sso (run as python -m tickets search sso in the example) prints three tickets: KITE-139 (closed), KITE-142 (in progress) and KITE-144 (open). The task is to let you add --status open and get only KITE-144.

AGENTS.md is a rules file: a Markdown file in the repo that tells coding tools how this project works. Its rules section, which matters later:

## Rules
- Run the tests with `python -m pytest -q` before you say you're done.
- Every new behavior gets a test.
- A status must be one of `STATUSES`. Raise `ValueError` for anything else.
- Keep search functions pure: a list of tickets in, a list of tickets out.

Three ways to use the same model

Every coding tool is built on the same kind of model from What LLMs Are: text in, text out, one token (a chunk of about 4 characters) at a time. What differs is how much the tool sends, how many times it calls the model, and whether anything checks the result. There are three shapes.

what goes to the model model calls what you get back AUTOCOMPLETE as you type code before and after your cursor: search.add_argument("--st| + a snippet from an open tab about 190 tokens 1 call no tools grey text, Tab to accept: atus", choices=STATUSES) nothing runs it, so nobody sees STATUSES was never imported CHAT you ask your question: "How do I add a --status filter?" + the files you paste or mention + the chat so far 1 call per message an answer with a code block; you paste it in, run the tests, and come back with the error you are the loop AGENT you hand over a task the task: "Add a --status filter to search_tickets." + AGENTS.md, the file tree, 4 tools about 545 tokens on turn 1 7 calls in a loop, 13 tool calls a diff across 4 files, after it ran the tests itself: 3 failed, then 6 passed you review the diff Going down: more context, more calls, more time, and more of the checking done for you.
The same task in the three shapes. The numbers come from the companion example.

Autocomplete suggests the rest of the line, or the next few lines, as you type. It’s one small model call per suggestion. Say you’ve typed search.add_argument("--st in cli.py. The tool sends the code before your cursor, the code after it, and a snippet from another file you have open. This is fill-in-the-middle: the model is trained to write the missing piece between a prefix and a suffix. The companion’s autocomplete.py builds that request, trimmed here:

<file_sep>tickets/models.py
STATUSES = ("open", "in_progress", "closed")
<file_sep>tickets/cli.py
<fim_prefix>...
    search.add_argument("query")
    search.add_argument("--st<fim_suffix>
    args = parser.parse_args(argv)
    ...<fim_middle>

The <fim_...> markers are the format some open code models use; each product has its own and doesn’t publish it. The request is about 190 tokens, and a plausible reply is atus", choices=STATUSES). It’s also a NameError waiting to happen, because cli.py never imports STATUSES. The model saw it in the open tab and assumed. Nothing runs the suggestion, so you find out when you do.

Chat is a conversation in a side panel. You ask how to add the filter, mention the files it needs, and get an answer with a code block: one call per message. You paste the code in, run it, and bring back the error. You are the loop.

Agent mode is where the tool runs that loop for you. You hand over a task, and the model calls tools (search, read a file, edit a file, run a command) until it decides it’s done. This is the agent from What an AI Agent Actually Is: a model in a loop with tools. A coding agent is that loop with file and shell tools. On the Kitebase task it makes 7 model calls and 13 tool calls.

At the time of writing, most products offer more than one shape. GitHub Copilot and Cursor do autocomplete, chat and agent work inside your editor. Claude Code runs in the terminal as an agent and has no autocomplete. Choosing Your Tool compares them. The rest of this article is about agent mode, since that’s where most of the machinery is.

Where the context comes from

The model knows nothing about your repo. It sees only its context window: the text sent with this one call, up to a limit measured in tokens. A real codebase won’t fit, so before the first call the tool has to choose what goes in, and that choice decides most of what comes out.

The toy agent builds a system prompt (the instructions sent at the top of every call) from three things: its own instructions, the rules file, and the file tree. Then it adds the four tool definitions: the name, description and input schema the model sees for each tool. python main.py prints the sizes:

Context before the first call:
  instructions             about   62 tokens
  rules file (AGENTS.md)   about  164 tokens
  file tree                about   46 tokens
  tool definitions         about  252 tokens
  file contents            0 of 10 files (the model has to ask)

Look at the last line. The model knows tickets/search.py exists, but not a single line of it. Everything else it learns, it has to ask for with a tool, and every answer is added to the conversation and sent again on the next call:

TURN 1: ABOUT 545 TOKENS instructions 62 AGENTS.md, the rules file 164 file tree: 10 paths, no contents 46 4 tool definitions 252 the task 10 built once, sent on every call TURN 7: ABOUT 2,959 TOKENS everything from turn 1 545 search results, 7 lines 136 5 files it read 551 5 edits: old and new text 592 2 test runs 365 the model's own notes 128 tool requests, ids, wrapping 642 grows with every tool result NEVER IN THE CONTEXT the 4 files it didn't open, like tickets/data.py why the team wants statuses validated (a ticket, a chat) git history a teammate's branch that also changes search.py how the CLI runs in production The model knows what's in the first two boxes and nothing else. For the third, it guesses.
What the model can see on the first call and the last one. The third box is everything it has to guess.

Real tools gather context from the same places, plus a few more:

  • The rules file, loaded at the start of every session. Claude Code reads CLAUDE.md (and, at the time of writing, AGENTS.md when there’s no CLAUDE.md). Cursor reads .cursor/rules/ and AGENTS.md. GitHub Copilot reads .github/copilot-instructions.md and AGENTS.md. Same idea, different file names.
  • What you point at: files you mention by name, the file open in your editor, the lines you selected.
  • What it finds: searching the code as it works. Claude Code, for one, searches by running grep and find through its shell tool on macOS and Linux.

When the answer depends on something in the third box, like why the team wants statuses validated, the model fills the gap with whatever looks most likely. That’s the hallucination problem from What LLMs Are, in code.

Why not just load the whole repo into the context?

Two reasons: cost and attention.

Every call re-sends the whole conversation. A 200,000-token repo sent on each of this example’s 7 calls is 1.4 million input tokens, $7 at claude-opus-5’s $5 per million, for one small change.

And models track details less well in very long prompts, so 20 relevant files beat 2,000 mostly irrelevant ones. Context Management covers how to feed the right slice.

The loop: read, edit, run the tests

Once the context is built, the agent runs the loop from Plan, Act, Observe: The Agent Loop in Code: call the model with the tools and the whole conversation, run the tool calls it asks for, send the results back, repeat until it stops asking. The toy’s run_agent is that loop with four tools:

ToolWhat it doesClaude Code’s version, at the time of writing
search(text)Every line containing the text, as path:line: textgrep and find via Bash
read_file(path)One file, with line numbersRead
edit_file(path, old, new)Replace exact text in one fileEdit
run_tests()python -m pytest -q, last 25 lines of outputBash, which runs any command

Here is the Kitebase run, turn by turn:

The model picks every step: what to read, what to change, when it's done. 7 calls, one per turn. READ turns 1 and 2 search 7 matching lines read_file x5 551 tokens of code EDIT turn 3 search.py new status param cli.py new --status flag RUN TESTS turn 4 3 failed 1 passed AttributeError: None.lower() EDIT turn 5 reads the error, only filters when a status is given + 2 new tests RUN TESTS turn 6 6 passed exit code 0 DONE turn 7 end_turn a summary and a diff the test output goes back to the model IT STOPS WHEN THE TESTS PASS, NOT WHEN THE CODE IS RIGHT $ kitebase search sso --status archived No tickets found. AGENTS.md says an unknown status must raise ValueError. The rule was in the context on all 7 turns. No test checks it.
The worked example as the companion prints it. Seven turns, and the last one is a judgment call by the model.

Read before you edit

Turn 1 is a search, because the model can’t read a file it can’t find. Turn 2 asks for five files at once (the definition, its caller, their tests and the model class), since none of those reads depends on another:

Turn 2: sent about 764 tokens. stop_reason = 'tool_use'
  model: "It's defined in tickets/search.py and called from tickets/cli.py. I'll read both, their ..."
  read_file(path='tickets/search.py')
      -> 10 lines, 93 tokens
  read_file(path='tickets/cli.py')
      -> 19 lines, 176 tokens
  ...three more

stop_reason = 'tool_use' means “run these and come back”. The model: line is the note the model writes before its tool calls, and the first thing to read when an agent does something odd.

Edit by exact replacement

Turn 3 edits two files. The edit tool doesn’t take a line number or a whole new file. It takes a piece of the old text and the text to replace it with:

def edit_file(repo: Path, path: str, old: str, new: str) -> str:
    target = _inside(repo, path)
    text = target.read_text()
    count = text.count(old)
    if count == 0:
        raise ValueError(f"old text not found in {path}. Read the file again and copy it exactly.")
    if count > 1:
        raise ValueError(f"old text appears {count} times in {path}. Include more lines so it's unique.")
    target.write_text(text.replace(old, new))
    return f"Edited {path}: replaced {old.count(chr(10)) + 1} line(s) with {new.count(chr(10)) + 1}."

Claude Code’s Edit tool works the same way: an old_string, a new_string, no fuzzy matching. If the model misremembers one space, the edit fails with a clear error instead of landing in the wrong place. That’s also why coding agents read so much: they need the exact text to edit it.

Run the tests, read the failure

Turn 3’s edit to search.py has a bug, the kind models write all the time:

        and t.status == status.lower()

When you run kitebase search sso without the flag, status is None, and None has no .lower(). On turn 4 the agent runs the tests, as AGENTS.md asks, and gets this back:

Turn 4: sent about 2,045 tokens. stop_reason = 'tool_use'
  run_tests()
      -> exit code 1: 3 failed, 1 passed in 0.01s
         AttributeError: 'NoneType' object has no attribute 'lower'

That output goes into the next call like any tool result. On turn 5 the model reads it, changes the line to and (status is None or t.status == status.lower()), and adds two tests. Turn 6: 6 passed. Turn 7 has no tool call, so the loop ends.

This is the step that makes an agent more than a chat window with file access: its first try was wrong, and the loop turned that into an error it could act on. Autocomplete had no such step. The toy runs only the tests, but a real agent’s shell runs anything, so the linter, the type checker and the build are all more chances to catch its own mistakes.

It stops when it thinks it’s done

There’s no check in run_agent for “is the task done?” The loop stops when the model replies without asking for a tool, so when the model decides it’s finished. Here it decided the moment the tests passed:

Final answer:
Added an optional `status` argument to `search_tickets` and a `--status` flag to
`kitebase search`, with a test for each. All 6 tests pass.

Every word is true, and --status archived still returns “No tickets found.” instead of the ValueError that AGENTS.md asks for. The rule was in the system prompt on all 7 calls, and the model even read models.py, where STATUSES lives. No test asked, so nothing made it act.

So rules files are context, not enforcement. Claude Code’s own docs say the same about CLAUDE.md: the model reads it and tries to follow it, with no guarantee. A rule the tests don’t need will sometimes be skipped.

The fix is to turn the rule into a check. Add one test to tests/test_search.py:

def test_unknown_status_raises():
    with pytest.raises(ValueError):
        search_tickets(TICKETS, "sso", status="archived")

Now the rule is part of “done”. On the offline run the second test run comes back 1 failed, 6 passed with DID NOT RAISE ValueError. The scripted model announces “All 6 tests pass” anyway, because it doesn’t read results: a reminder to read the tool output, not the summary. A real model would see the failure and, most of the time, add the check.

Why they’re fast at some things and wrong at others

You can predict how an agent will do on a task from three questions:

  • Is what it needs in the context, or findable? Adding a flag to an argparse command needs three files, all found by one search. “Make search behave the way the product team wants” needs a conversation that isn’t in the repo.
  • Is this a common pattern? Optional filter arguments appear in countless public codebases the model learned from. An odd in-house convention appears in yours alone, so it has to be written down.
  • Is there a check? With tests, a wrong first try becomes a failure message and a second try. Without them, the first plausible answer is the final one.

Tasks with three yeses, like adding a flag, testing a pure function or a mechanical rename, go fast and mostly right. A single no, like a change that depends on why the code is the way it is or a bug that only shows in production, goes wrong in a confident, well-formatted way.

Speed is also less clear-cut than it feels. In a 2025 study by METR, 16 experienced open-source developers worked on 246 real tasks in large codebases they knew well, with and without early-2025 AI tools. With the tools they took 19% longer, yet afterwards they believed the tools had sped them up by 20%. It’s one study, with tools that have changed since, but it’s a good reason to measure instead of going by feel.

Size matters too. The Kitebase run sent about 13,250 input tokens over its 7 calls, roughly 7 cents at claude-opus-5’s $5 per million. A task that reads 30 real files sends far more, and the longer the conversation, the more early detail the model loses. Small, checkable tasks are cheaper and more accurate. Prompting Coding Agents covers how to scope them.

Try it yourself

The companion example is the toy coding agent plus the autocomplete request builder. It needs no API key; --claude runs the same loop against claude-opus-5 if you have one.

Download the runnable example (zip)

cd 01-how-ai-coding-tools-work
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Every run edits a temporary copy of kitebase/, so the original never changes. Then try these:

  1. python main.py --run "search sso --status archived". You’ll see “No tickets found.” after the diff: the broken rule from this article.
  2. Add test_unknown_status_raises (and import pytest) to kitebase/tests/test_search.py and run again. The second test run fails, and the scripted model still claims success. Remove the test afterwards.
  3. python main.py --show-context prints the whole turn-1 system prompt. Delete kitebase/AGENTS.md and run it again: nothing left in the prompt tells the model to run the tests or add any.

pytest -q runs the offline tests. They need no key and no network.

Common beginner mistakes

  • Expecting agent results from autocomplete. Autocomplete is one call with no checks. It’s great for the next line and blind to the rest of the repo.
  • Assuming the tool has seen your code. It has seen the file tree and whatever it chose to read. Name the files that matter in your request.
  • Writing rules and not checks. A rule in AGENTS.md is a strong hint. A test fails loudly. Put the rules you can’t afford to break into tests, linters or hooks (scripts some tools run before or after each edit or command).
  • Trusting the final summary. “All tests pass” is the model’s description of the tool output. Scroll up and read the actual output, then read the diff.
  • Handing over tasks with no way to check them. If the repo has no tests for the area, the agent’s first plausible answer is its last. Ask it to write the test first.

Questions you will face in production

“Is it safe to let an agent edit and run things on my machine?” It runs with your permissions, so treat it that way. At the time of writing, Claude Code’s manual permission mode reads files in your project without asking, but asks before editing a file or running a shell command other than a short list of read-only ones. Other tools have similar settings. Work on a branch, keep auto-approval for low-risk commands like the test runner, and review the diff like a pull request from a fast new teammate.

“Can it use our internal tools, like the ticket tracker or the staging database?” Yes, by giving it more tools. The standard way to plug your own systems into a coding tool is MCP, and MCP and Custom Tools in Your Editor shows how.

Check your understanding

Autocomplete suggests a line that uses a function your file never imports. Why didn't it notice?

Autocomplete is one model call on the code around your cursor and a few snippets from open files. It probably saw the function in another tab. Nothing runs the suggestion or checks the imports, so the mistake reaches you, not the model.

An agent edits the wrong file: it changes a search helper nobody uses instead of the one the CLI calls. What was probably missing from its context?

The link between the CLI and the real function. It likely found the unused helper first and never read cli.py to see which one is called. Name the right file in your request, or ask it to search for the callers before editing. The model can only work from what it read.

Your rules file says "never change the public API of search.py", and the agent changes it anyway. What do you do?

Accept that a rules file is context, not enforcement, and add a check. For example, a test that calls search_tickets with the old signature, or a hook that blocks edits to that file. Then breaking the rule makes a check fail, and the loop has to deal with it.

Why does the edit tool take a piece of the old text instead of a line number?

Line numbers shift as soon as an earlier edit adds a line, so “replace line 9” can hit the wrong line without any error. Exact old text either matches once or fails loudly, and the failure tells the model to read the file again. Claude Code’s Edit tool has the same two checks: the old text must match exactly and appear only once.

What to remember

  • Autocomplete is one call on the code around your cursor. Chat is one call per message, with you as the loop. An agent runs the loop itself, with tools to search, read, edit and run code.
  • The model only knows what’s in its context: instructions, the rules file, the file tree, and whatever it reads. Everything else, it guesses.
  • The read, edit, run-the-tests loop is what lets an agent fix its own mistakes. Fast tests make it much better.
  • The loop stops when the model decides it’s done, not when the code is right. Read the tool output and the diff.
  • Rules files are hints. Turn the rules that matter into tests or other checks.

What to study next

Now that you know the machinery, Choosing Your Tool: Cursor, Claude Code, Devin Desktop compares the products built on it. After that, Setting Up Context: Rules Files and Project Config is about writing the file that fills the top of every context window, and Prompting Coding Agents about the request that starts the loop.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.