Prompting Coding Agents
You open a coding agent in the Kitebase ticket service and type add a --status filter to search_tickets. Thirty seconds later it says “Done!” The diff (the set of line changes it made) adds a module you didn’t ask for, changes the output format, and lets --status done (Kitebase calls it closed) quietly print nothing. Three tests fail. The agent never ran them, and you never asked it to.
The model isn’t bad at coding. You gave it one sentence and it guessed the rest. In a chat, a vague question gets a vague answer that you then act on. A coding agent (an LLM in a loop with tools to read files, edit them and run commands, as in article 01) turns your words straight into edits. Nobody asks what you meant. The guesses land on disk.
What you’ll build: the same change prompted two ways, vague and good, plus a check script that decides whether a change is really done: the tests pass, and only the expected files changed.
The worked example: a tiny ticket service
Kitebase is a made-up project-tracking app. Its ticket service is two files and a test folder:
ticket-service/
kitebase/tickets.py Ticket, 5 sample TICKETS, search_tickets()
kitebase/cli.py python -m kitebase.cli search <query> [--assignee NAME]
tests/test_search.py 4 tests, passing
tests/test_status_filter.py 3 tests, failing on purpose
Here’s the function the change is about:
STATUSES = ("open", "in_progress", "closed")
def search_tickets(
tickets: list[Ticket], query: str, assignee: str | None = None
) -> list[Ticket]:
"""Tickets whose title contains `query` (case-insensitive), optionally for one assignee."""
q = query.lower()
results = [t for t in tickets if q in t.title.lower()]
if assignee is not None:
results = [t for t in results if t.assignee == assignee]
return results
And the CLI in use:
$ python -m kitebase.cli search export
KB-101 open CSV export drops the last row
KB-102 closed Export button hidden on mobile
KB-104 open CSV export ignores the date filter
You want --status open to keep only the open ones. Before touching the agent, write down what “working” means as tests. tests/test_status_filter.py:
def test_search_filters_by_status():
assert [t.id for t in search_tickets(TICKETS, "export", status="open")] == ["KB-101", "KB-104"]
def test_cli_status_flag(capsys):
main(["search", "export", "--status", "closed"])
assert capsys.readouterr().out == "KB-102 closed Export button hidden on mobile\n"
def test_cli_rejects_unknown_status(capsys):
with pytest.raises(SystemExit):
main(["search", "export", "--status", "done"])
assert "invalid choice: 'done'" in capsys.readouterr().err
The first pins the new parameter, the second the output format. The third says a typo must fail loudly: argparse (Python’s built-in argument parser) prints invalid choice when a flag only accepts a fixed list of values. All three fail today. That file is the finish line you’ll hand to the agent.
The vague prompt, and what it produces
Here’s the whole prompt, from prompts/01-vague.txt:
add a --status filter to search_tickets
Real agent output varies from run to run, so the transcript below is illustrative: written by hand to show a typical run, not captured. The files it describes are in the companion code’s sample_runs/vague/, and the check on them is real.
ILLUSTRATIVE, not a captured session
> add a --status filter to search_tickets
Read kitebase/tickets.py
Read kitebase/cli.py
I'll put the filter in a small reusable module so other commands can use it.
Write kitebase/filters.py
Edit kitebase/tickets.py search_tickets takes status, calls filter_by_status
Edit kitebase/cli.py adds --status; output now shows the assignee too
Done! `kitebase search` now accepts --status (e.g. --status open). I also
made the output easier to scan by adding the assignee and separating the
columns with "|".
It sounds finished. Now run the check script from the companion code on what it left behind:
$ python check_done.py --sample vague
Tests: 3 failed, 4 passed
failing: tests/test_search.py::test_cli_prints_one_line_per_ticket
failing: tests/test_status_filter.py::test_cli_status_flag
failing: tests/test_status_filter.py::test_cli_rejects_unknown_status
Changed files:
kitebase/cli.py modified expected
kitebase/filters.py new NOT EXPECTED
kitebase/tickets.py modified expected
NOT DONE:
- tests don't pass: 3 failed, 4 passed
- kitebase/filters.py was added, but it isn't one of the expected files
Each problem traces back to a question the prompt left open:
- Where does the filter go? It guessed “a new module”. Now filtering lives in two places.
- Which values are valid? It guessed “any string”, so
--status donefinds nothing and exits happily. - Can the output change? It guessed “yes, make it nicer”, breaking an existing test and any script that parses the output.
- When is it done? It guessed “when the code looks right” and never ran the tests.
None of those guesses is crazy. A new teammate given one sentence and no chance to ask might make the same ones. The fix is to answer the questions up front.
A good prompt answers four questions
A good agent prompt reads like a small ticket for a fast engineer who has never seen your codebase. It covers four things:
- The goal: the behaviour you want, with a concrete example.
- Where to look: the files and functions involved, and an existing pattern to copy.
- Constraints: what must not change.
- How to verify: the command that proves it’s done, and the result you expect.
Here’s the full prompt from prompts/02-good.txt:
Add a --status filter to `kitebase search`, so
`python -m kitebase.cli search export --status open` shows only open tickets
whose title contains "export".
Where to look: search_tickets in kitebase/tickets.py and build_parser in
kitebase/cli.py. Follow how --assignee already works, from the flag down to
the filter. Valid statuses are the ones in STATUSES in kitebase/tickets.py.
tests/test_status_filter.py describes the behaviour and fails right now:
FAILED tests/test_status_filter.py::test_search_filters_by_status
TypeError: search_tickets() got an unexpected keyword argument 'status'
FAILED tests/test_status_filter.py::test_cli_status_flag - SystemExit: 2
FAILED tests/test_status_filter.py::test_cli_rejects_unknown_status
3 failed, 4 passed
Constraints:
- Only change kitebase/tickets.py and kitebase/cli.py.
- Don't edit any tests. Don't change the output format. No new dependencies.
- Keep status optional, so existing calls to search_tickets still work.
Done means `python -m pytest -q` passes with all 7 tests. Run it and show me
the output before you say you're finished.
It’s about 140 words, and every line answers a question the vague run had to guess at. Three details do most of the work.
“Follow how —assignee already works.” Pointing at an existing pattern beats describing one. The agent reads the real code and copies its shape: an optional keyword argument, a list comprehension next to the assignee one, a flag passed straight through. In Claude Code you can also write @kitebase/tickets.py to pull a file into the prompt, and the tool reads it before it starts.
The constraints. “Only change these two files” turns an unexpected filters.py from a style debate into a clear violation. “Don’t edit any tests” closes the most tempting shortcut, covered below.
“Run it and show me the output.” The verify step, covered next.
The illustrative run for this prompt (hand-written again; the files are in sample_runs/good/):
ILLUSTRATIVE, not a captured session
Read kitebase/tickets.py
Read kitebase/cli.py
Read tests/test_status_filter.py
Edit kitebase/tickets.py status=None parameter, filtered next to assignee
Edit kitebase/cli.py --status flag, passed through to search_tickets
Bash python -m pytest -q
FAILED tests/test_status_filter.py::test_cli_rejects_unknown_status
1 failed, 6 passed
The flag accepts any string. The test wants argparse to reject unknown
values, so I'll restrict it to STATUSES.
Edit kitebase/cli.py choices=STATUSES
Bash python -m pytest -q
7 passed
Done. I changed kitebase/tickets.py and kitebase/cli.py only. Output of
python -m pytest -q: 7 passed.
The line in cli.py that matters:
search.add_argument("--status", choices=STATUSES, help="only tickets with this status")
And the check agrees:
$ python check_done.py --sample good
Tests: 7 passed
Changed files:
kitebase/cli.py modified expected
kitebase/tickets.py modified expected
DONE: tests pass and only the expected files changed.
Isn't a 140-word prompt slower than writing the code myself?
For a two-line change you already know how to write, yes. Write it.
The prompt pays off when the agent saves you the reading: finding where things live, matching the pattern, running the tests and fixing what fails. And the stable parts (the test command, “don’t edit tests”) belong in your rules file from article 03, so you stop typing them.
Give it a check it can run
Of the four parts, the verify line matters most. Claude Code’s docs put the problem in one line: “Claude stops when the work looks done.” An agent loops: act, look at the result, decide whether to keep going. Without a command to run, “look at the result” means rereading its own diff, and code you just wrote always looks right.
A test run is a signal that doesn’t depend on the agent’s judgment: pass or fail, with the failing test’s name. In the good run, the first pass left --status accepting any string. Rereading didn’t catch it. pytest did.
A check is anything that prints a result the agent can read: tests, a type checker, a linter, a build, a script that diffs output against a known-good file. Put the expected result in the prompt (all 7 pass), so “almost” doesn’t count.
Two habits make it stick:
- Ask for evidence, not a claim. “Show me the output” means the last thing you read is
7 passed, not “I’ve verified everything works”. - Check the check yourself. An agent told to make tests pass has one shortcut: change the tests. The good prompt forbids it, and
check_done.pycatches it anyway. Here’s the part that looks at file changes:
for name, kind in changes.items():
what = {"new": "added", "modified": "modified", "deleted": "deleted"}[kind]
if name.startswith("tests/"):
# An agent that edits the test it was told to pass has moved the finish line.
found.append(f"a test file was {what}: {name}")
elif name not in expected:
found.append(f"{name} was {what}, but it isn't one of the expected files")
changes comes from comparing a SHA-256 fingerprint of every file with baseline.json, saved before the agent started. In a real repo, git diff --name-only gives you the same list. It’s the same instinct as validating an LLM’s JSON with a schema in Prompting as Code: don’t trust the output’s description of itself.
The gotcha: a check only proves what it tests. If there were no test_cli_rejects_unknown_status, the first pass would have gone green with the typo bug in it. When you write the test first, write one for the failure you care about, not just the happy path.
Paste the failing test or the error
The good prompt includes the failing output itself, not “the status tests are failing”:
FAILED tests/test_status_filter.py::test_search_filters_by_status
TypeError: search_tickets() got an unexpected keyword argument 'status'
Those two lines name the test and the exact mismatch, so the agent starts at the right function instead of reproducing the problem. Same for bugs: “the export crashes” sends the agent looking; the traceback’s last ten lines send it to the file and line.
In Claude Code you can paste it, reference a file with @, or pipe it in: cat error.log | claude. In any tool, include the command you ran, so the agent can rerun it to check its fix.
Trim it. The relevant 5 to 30 lines help. A 3,000-line CI log mostly fills the agent’s context window (the text a model can see at once, from article 01 of AI Engineering) with noise, and the one line that matters has to compete with everything else.
Ask for a plan first on bigger changes
The --status filter is small. You could describe the diff in one sentence: “add an optional parameter and a flag, like --assignee.” For changes like that, go straight to the prompt; a plan is overhead. Claude Code’s docs draw the line in the same place: when the whole diff fits in one sentence, planning costs more than it saves.
The next Kitebase feature is different: tickets get a closed date, and search gets --closed-since 2026-09-01. That touches the model, the sample data, search, the CLI and the tests, and it hides decisions: the date’s type, what happens to open tickets, what a bad date does. Get the first one wrong and everything after it is built on it.
So ask for a plan first, and say you don’t want code yet. From prompts/03-plan-first.txt:
Don't change any files yet. I want a plan first.
Next feature: tickets get a closed date, and `kitebase search` gets a
--closed-since flag, e.g. `--closed-since 2026-09-01` shows only tickets
closed on or after that day.
Read kitebase/tickets.py, kitebase/cli.py and the tests, then tell me:
1. Which files you'd change, and what changes in each.
2. How the closed date is stored on Ticket, and its type.
3. What --closed-since does with tickets that aren't closed.
4. What happens with a badly formatted date.
5. Which tests you'd add, by name.
Keep it short. Wait for me to approve before you write code.
In Claude Code you can enforce “don’t change files” instead of just asking: plan mode is a permission mode where the agent reads files and can run commands to explore, but doesn’t edit your source. Press Shift+Tab until the status bar shows ⏸ plan mode on, or start the session with claude --permission-mode plan. Ctrl+G opens the plan in your editor to change it directly. Other tools have their own read-only or “ask” modes, and in any tool the prompt’s first line does most of the job.
A plan might come back like this (illustrative again):
ILLUSTRATIVE, not a captured session
1. tickets.py: add closed_at to Ticket, set it on KB-102.
search_tickets gets closed_since. cli.py: add --closed-since.
2. closed_at: str | None, e.g. "2026-09-14". Compare as strings.
3. Tickets with no closed_at are left out.
4. Parse with python-dateutil, so "Sept 1" works too.
5. test_search_filters_by_closed_since, test_cli_closed_since_flag,
test_cli_rejects_bad_date.
Two lines deserve pushback. Comparing dates as strings works for 2026-09-14 and breaks the day someone stores 2026-9-4. Point 4 adds a dependency for formats nobody asked for. Your reply is one sentence: “Store closed_at as a datetime.date, parse the flag with date.fromisoformat, no new dependency.” Fixing that in the plan costs a sentence. Fixing it after the code exists costs a review, a revert and a new prompt.
Work in small steps
A plan also gives you the steps. Don’t send the approved plan back as one “now build all of it”. Split it so each prompt ends at a passing check:
- Add
closed_at: date | NonetoTicketand set it on the closed sample ticket. Existing 7 tests still pass. - Add
closed_sincetosearch_tickets, with its test. 8 pass. - Add the
--closed-sinceflag and the bad-date test. 10 pass.
Each step gives you a diff you can read in a couple of minutes, and a point to return to. Commit after each green step; if step 3 goes sideways, git checkout . (plus git clean -fd for new files) takes you back to the end of step 2. One big run hides its early mistakes under its later work, and you end up approving a 300-line diff because reading it properly would take an hour. Article 07 goes further on changes too big for one sitting.
Small steps also mean watching while it works. If the agent heads the wrong way (“I’ll add a filters.py so…”), stop it. In Claude Code, Esc stops the agent mid-action and keeps the conversation, so your next message redirects it: “no new module; filter inside search_tickets.” A correction before it builds on the mistake is one sentence; after, it’s a rewrite.
Know when to stop and restart
Sometimes correcting doesn’t converge. You told it to keep the old output; it did, but filters.py stayed. You told it no new module; it moved the filter and now accepts any status again. Every attempt and correction is still in the session, and the agent rereads all of it, wrong turns included.
Claude Code’s docs give a concrete rule that works for any tool: after two failed corrections on the same issue, clear the session and start again with a better prompt that includes what you learned.
Restarting takes three moves:
- Undo the code. In a repo,
git checkout .reverts edited files andgit clean -fdremoves new ones likefilters.py. In Claude Code,Esctwice (or/rewind) can restore files to an earlier checkpoint (a snapshot it takes before each prompt’s changes), but checkpoints only track its own file edits, not changes made by shell commands. Git is the safer undo. - Clear the conversation.
/clearin Claude Code, or a new chat in other tools. - Send a better prompt. Start from the good one and add what the failed session taught you.
prompts/04-restart.txtadds three lines:
An earlier attempt went wrong, so to be explicit:
- Filter inside search_tickets, next to the assignee filter. No new module.
- Use argparse `choices=STATUSES`, so `--status done` fails with argparse's
normal "invalid choice" error.
- Keep the output line exactly as it is: f"{t.id} {t.status:<12} {t.title}".
What you learned travels as three lines; the failed attempts stay behind. Restarting feels like losing work, but the work was wrong, and the useful part is now in the prompt.
Try it yourself
The companion example has the ticket service, the four prompts, the two illustrative runs and check_done.py. It runs offline, with no API calls.
Download the runnable example (zip)
cd 04-prompting-coding-agents
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python check_done.py # untouched project: 3 failed, NOT DONE
python check_done.py --sample vague # the vague run: NOT DONE
python check_done.py --sample good # the good run: DONE
Then point a real agent at it:
- Copy
ticket-service/somewhere,git initand commit it, and start your agent inside it (claudefor Claude Code). Pasteprompts/01-vague.txt, then runpython check_done.py --project path/to/your/ticket-service. Did it run the tests unasked? git checkout .andgit clean -fd, clear the session, and pasteprompts/02-good.txt. Run the check again.- Paste
prompts/03-plan-first.txtin plan mode. Find one decision in the plan you’d change, and change it before any code exists.
pytest -q runs the check script’s own offline tests, including one that edits a test file and confirms the check still says NOT DONE.
Common beginner mistakes
- Prompting like a chat. One sentence, sent to something that edits files. The agent won’t ask what you meant; it’ll guess, and you’ll review the guesses.
- No command that proves it’s done. Without “run
pytest -q, all 7 must pass”, the agent stops when the diff looks finished to it. - Describing a pattern instead of pointing at one. “Use our usual style” gets a style. “Follow how
--assigneeworks” gets yours. - Summarising the error. “The tests fail” makes the agent rediscover what you know. Paste the failing lines and the command.
- One big prompt for a multi-file feature. Plan first, then build in steps that each end green.
- Correcting forever. Two failed corrections means the session is working against you. Undo,
/clear, write the better prompt.
Questions you will face in production
“Long detailed prompts or short ones?” Precise, not long. Spend words on the goal, the files, the constraints and the done check; cut adjectives like “clean” and “production-ready” that change nothing the agent does. If a prompt grows past a screen, the task is probably two tasks.
“The agent keeps ignoring our conventions. Do I repeat them in every prompt?” No. Anything you’d type into every prompt (the test command, “don’t edit tests”, naming rules) goes in the rules file from article 03. The prompt carries only what’s specific to this task.
“It changed files I didn’t expect. Is that always bad?”
Not always, but it needs a reason. Run git diff --name-only before you review, so the surprise shows up in one line instead of halfway through a long diff. Article 06 covers reviewing the rest.
Check your understanding
The agent says "Done, all tests pass." What do you do before reviewing the diff?
Look for the evidence: the command it ran and its output, ideally a line like 7 passed. If it isn’t there, ask for it or run the tests yourself. Then check which files changed (git diff --name-only), and especially whether any test file did.
Your teammate's prompt is "fix the flaky login test, make sure it's solid". What's missing, and what would you add?
Where to look (the test file and the code it covers), the failure itself (the pasted output from a failing run), a constraint (“fix the cause; don’t add retries or sleeps to the test”) and a check with a clear bar, like “run the test 20 times in a row, all must pass”. “Solid” tells the agent nothing it can check.
Which of these should get a plan first: renaming a variable in one file, or adding a closed date that changes the Ticket model, search and the CLI?
The closed date. It touches several files and hides decisions (the date’s type, what happens to open tickets, bad input) that are cheap to fix in a plan and expensive to fix in code. The rename can be described in one sentence, so skip the plan.
You've corrected the agent twice and it's still wrong. What exactly do you do next?
Undo its edits (git checkout . and git clean -fd, or rewind to a checkpoint), clear the session with /clear, and send the original prompt plus a few lines saying what went wrong and what to do instead. The failed attempts stay behind; the lesson comes along as text.
What to remember
- An agent turns your prompt straight into edits. Anything you leave out, it guesses.
- A good prompt answers four questions: the goal, where to look, the constraints, and the command that proves it’s done.
- Point at existing code to copy, and paste the real failing output instead of describing it.
- Ask for evidence, and check the result yourself: tests pass, only the expected files changed, no test was edited.
- Plan first when you can’t describe the diff in one sentence, then build in small steps that each end green.
- After two failed corrections, undo, clear and restart with a better prompt.
What to study next
A good prompt only works if the agent can see the right things while it works. Next is Context Management: what fills the context window during a session, what to keep out of it, and how to scope a large repo so the agent works from facts instead of guesses. The prompting habits here are a special case of the ones in Prompting as Code, which applies them to the prompts inside your own product.
Further reading
- Claude Code docs: Best practices. Verification, plan mode, specific prompts, course-correcting and the two-corrections rule, from the people who build the tool.
- Claude Code docs: Permission modes. What plan mode and the other modes allow, and how to switch between them.
- Cursor documentation. How to reference files, set rules and stop an agent mid-task in Cursor.
- Anthropic: Building effective agents. Why a clear goal and a checkable signal matter to anything running an agent loop.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.