Large Multi-File Changes and Refactors

You ask a coding agent to move Kitebase’s status handling into its own module. It’s a good job for an agent: four files each know a piece of the status rules, and you want them in one place. A minute later it’s done. Seven files changed, 46 lines added, 18 removed, and the tests say 12 passed. You merge it.

The next morning a teammate’s script that reads kitebase summary breaks, because the summary now lists closed tickets first. Somewhere in those 64 lines the agent rewrote a loop with Counter, and the order changed. No test checked the order, so nothing noticed.

The agent didn’t do anything unusual. One big change gives you one big diff, and a big diff hides small mistakes from you and from the tests. This article runs the same change as five small ones, each checked and committed before the next.

What you’ll build: the Kitebase refactor as a written plan of five steps, and a script that applies them one at a time to a real git repo. It runs the tests after every step, commits each green one and rolls back the step that breaks something. It also tries one step two ways in two git worktrees. Everything runs offline.

The worked example: statuses in four places

Kitebase is the made-up project-tracking app used across this site. Its ticket service is the one from How AI Coding Tools Actually Work, grown a little: search can filter by status, and there’s a kitebase summary command that counts tickets by status.

$ python -m tickets summary
Open         3
In progress  2
Closed       3

Here’s where the status rules live, found with one grep:

$ grep -rn "STATUSES\|LABELS" tickets
tickets/models.py:3:STATUSES = ("open", "in_progress", "closed")
tickets/search.py:1:from .models import STATUSES, Ticket
tickets/search.py:6:    if status is not None and status not in STATUSES:
tickets/cli.py:4:from .models import STATUSES
tickets/cli.py:8:LABELS = {"open": "Open", "in_progress": "In progress", "closed": "Closed"}
tickets/cli.py:16:    search.add_argument("--status", choices=STATUSES)
tickets/cli.py:22:            print(f"{LABELS[status]:<12} {count}")
tickets/report.py:1:from .models import STATUSES, Ticket
tickets/report.py:5:    """How many tickets have each status, in the order of STATUSES."""
tickets/report.py:6:    counts = {status: 0 for status in STATUSES}

Adding a fourth status, say blocked, means finding all four files. The goal is a new tickets/status.py that holds all of it.

This is a refactor: changing how code is organised without changing what it does. The second half is the hard part. “Without changing what it does” is a claim, and with an agent writing the code, you need a way to check it. Today, 9 tests pass.

Plan first, and write the plan down

Prompting Coding Agents covered asking for a plan before any code when you can’t describe the diff in one sentence. In Claude Code that’s plan mode, where the agent reads and explores but doesn’t edit your files: press Shift+Tab until the status bar shows ⏸ plan mode on, or start with claude --permission-mode plan. For a refactor, ask for a particular shape of plan. From the companion’s prompts/01-plan.txt, trimmed:

Don't change any code yet. I want a plan first.

Move all status handling into a new module, tickets/status.py, without
changing behaviour: same CLI output, same errors.

Give me a numbered plan where:
- each step changes one or two files,
- `python -m pytest -q` passes after every step,
- the old code stays until nothing uses it, and is deleted in the last step.

For each step, name the files and the test count you expect. Check the
tests cover everything the refactor touches; if something isn't covered,
make step 1 a test that pins its current behaviour.

When I approve the plan, save it to PLAN.md before you touch any code.

Each constraint makes the plan easier to check. “One or two files per step” keeps every diff readable. “Tests pass after every step” makes each step verifiable on its own. “Name the test count” gives you a number to compare: if step 2 should end at 13 passed and ends at 12, a test went missing.

The approved plan, from PLAN.md in the companion code:

- [ ] 1. Pin the summary output with a test. Nothing tests what `kitebase summary` prints.
      Files: tests/test_cli.py. Expect 10 passed.
- [ ] 2. Add tickets/status.py (`STATUSES`, `LABELS`, `check_status`, `label`) and its tests.
      Nothing uses it yet. Files: tickets/status.py, tests/test_status.py. Expect 13 passed.
- [ ] 3. search.py and report.py import from tickets/status.py. `models.STATUSES` stays.
      Files: tickets/search.py, tickets/report.py. Expect 13 passed.
- [ ] 4. cli.py imports from tickets/status.py. The `LABELS` dict goes. Keep argparse `choices`.
      Files: tickets/cli.py. Expect 13 passed.
- [ ] 5. Delete `STATUSES` from models.py and update AGENTS.md.
      Files: tickets/models.py, AGENTS.md. Expect 13 passed.

The file also holds the goal, rules for every step (one step per prompt, only the named files, don’t change existing tests) and what “done” means. Read it for the mistakes that are cheap to fix now: a missing file (did it find report.py?), a step that deletes something early, a step with no test count. Each fix is one sentence in a reply.

Why save the plan to a file instead of leaving it in the chat?

Because the chat doesn’t last and the file does.

A five-step refactor outlives one conversation. You’ll /clear between steps, come back tomorrow, or hand step 4 to a teammate. A plan in the chat is gone after /clear, and gets squeezed when a long session is compacted (summarised to free up space). PLAN.md is there for every new session: “Read PLAN.md and do step 3” is a complete prompt. Tick steps off as they land, and delete the file when you’re done.

Pin the behaviour before you move it

Look at step 1 again. Before anything moves, the plan adds a test. Nothing in Kitebase checked what kitebase summary prints, and the refactor is about to touch the code that prints it.

A test like this has a name: a characterization test (or pinning test), one that records what the code does today, right or wrong, so you notice if it changes. Here it is:

def test_summary_prints_every_status_in_order(capsys):
    assert main(["summary"]) == 0
    assert capsys.readouterr().out == "Open         3\nIn progress  2\nClosed       3\n"

That’s the test that would have caught the one-shot bug from the opening. The companion’s tests prove it: apply step 1, then the one-shot patch on top, and 12 passed becomes a failure with a name:

1 failed, 12 passed
FAILED tests/test_cli.py::test_summary_prints_every_status_in_order

The bug was in report.py, where the agent “simplified” the counting:

def count_by_status(tickets: list[Ticket]) -> dict[str, int]:
    """How many tickets have each status."""
    return dict(Counter(t.status for t in tickets))

It looks equivalent, and the existing report test agrees, because comparing two dicts ignores their order. But Counter orders by first appearance, and the first sample ticket is closed. It also drops any status with zero tickets.

To find what needs pinning, ask the agent while it plans which touched behaviour has no test, as 01-plan.txt does. The gotcha: a pinning test pins bugs too. That’s fine. A refactor shouldn’t change behaviour, not even to fix a bug. Fix bugs in their own commit, so a changed test always means a change you meant.

Add the new, move the callers, delete the old

The order of steps 2 to 5 is the core of a safe multi-file change. It’s usually called parallel change, or expand and contract:

  1. Expand: add the new thing next to the old one. Nothing uses it, so nothing can break.
  2. Migrate: move the callers over, one or two at a time. Old and new both work, so every step stays green.
  3. Contract: delete the old thing, once nothing uses it.

Step 2 is pure expansion. It adds tickets/status.py and its three tests, and touches nothing else:

STATUSES = ("open", "in_progress", "closed")
LABELS = {"open": "Open", "in_progress": "In progress", "closed": "Closed"}


def check_status(status: str) -> str:
    """Return the status unchanged, or raise ValueError if Kitebase doesn't know it."""
    if status not in STATUSES:
        raise ValueError(f"unknown status {status!r}")
    return status


def label(status: str) -> str:
    """The name people see, e.g. 'In progress'."""
    return LABELS[check_status(status)]

Then each migration step is small. Step 3’s change to search.py:

-from .models import STATUSES, Ticket
+from .models import Ticket
+from .status import check_status
 ...
-    if status is not None and status not in STATUSES:
-        raise ValueError(f"unknown status {status!r}")
+    if status is not None:
+        check_status(status)

In this diff, lines starting with - were removed and + were added. You can check it in ten seconds.

BEFORE AFTER STEP 2 AFTER STEP 3 AFTER STEP 4 AFTER STEP 5 status.py not yet defines all of it defines all of it defines all of it defines all of it models.py STATUSES STATUSES STATUSES STATUSES Ticket only search.py from models from models from status from status from status report.py from models from models from status from status from status cli.py from models from models from models from status from status Old and new both exist for three commits. Every caller works whichever one it uses, so every step stays green. Drift: delete models.STATUSES during step 3 and cli.py, still "from models", fails with ImportError.
Where each file gets its statuses, commit by commit.

The rule that keeps contraction safe: delete only when a search says nothing uses the old thing. Before step 5, grep -rn "models import STATUSES" tickets must print nothing. Deleting early is an easy mistake for an agent: from the files it just edited, the old thing looks unused.

Commit after every green step

When a step’s tests pass, commit it before the next step starts. Each commit is a snapshot you can return to with one command, so a bad step can only cost you that step.

Start on a branch, so main stays untouched until the whole refactor is reviewed:

git switch -c refactor-status

The companion’s run_steps.py does exactly what you’d do by hand. Each step in steps/ is a patch, standing in for what an agent would do with that step’s prompt. The loop, trimmed:

for n, patch in enumerate(steps, start=first):
    title, files = apply(repo, patch)
    tests = run_tests(repo)
    if not tests.ok:
        print(f"  tests:   {tests.summary}  <- this step broke them")
        git(repo, "reset", "-q", "--hard")  # back to the last green commit
        git(repo, "clean", "-fdq")          # and remove any untracked files the step left
        return False
    git(repo, "commit", "-q", "-m", f"Step {n}: {title}")

Apply, test, commit if green, roll back and stop if red. python run_steps.py runs the plan:

Baseline: 9 passed, committed 592446c

Step 1: Pin the summary output with a test
  changed: tests/test_cli.py
  tests:   10 passed  -> committed 0db0c1f
Step 2: Add tickets/status.py next to the old code
  changed: tests/test_status.py, tickets/status.py
  tests:   13 passed  -> committed fc1af99
Step 3: Move search.py and report.py to tickets/status.py
  changed: tickets/report.py, tickets/search.py
  tests:   13 passed  -> committed a3d6b47
Step 4: Move cli.py to tickets/status.py
  changed: tickets/cli.py
  tests:   13 passed  -> committed 6d6aa42
Step 5: Delete STATUSES from models.py
  changed: AGENTS.md, tickets/models.py
  tests:   13 passed  -> committed 51af83c

Every test count matches the plan, and every step changed only the files the plan named. (The script fixes the commit author and date, so your hashes match these.)

ONE PROMPT, ONE DIFF "Move all status handling into status.py" ONE DIFF 7 files, +46 -18 PYTEST -Q 12 passed KITEBASE SUMMARY before: Open, In progress, Closed after: Closed, Open, In progress Every check it ran is green. No test pinned the order, so the change ships silently, somewhere in 64 changed lines. FIVE STEPS, FIVE COMMITS, TESTS AFTER EACH 1. PIN tests/ test_cli.py 10 passed 0db0c1f 2. ADD NEW status.py test_status.py 13 passed fc1af99 3. MOVE search.py report.py 13 passed a3d6b47 4. MOVE cli.py 13 passed 6d6aa42 5. DELETE OLD models.py AGENTS.md 13 passed 51af83c Step 1's test fails on the one-shot bug. Pin what you are about to move. One or two files per step. Each commit is a green point you can go back to.
The same refactor both ways. The numbers come from the companion code.

When a step goes wrong, the rollback is git reset --hard and then git clean -fd. That’s safe only because everything good is already committed: reset --hard throws away every uncommitted change, not just the bad ones.

Why do you need both git reset --hard and git clean -fd?

They undo different things.

git reset --hard puts every tracked file (one git already knows about) back the way the last commit had it. A brand-new file the agent created, like a stray tickets/filters.py, isn’t tracked, so reset leaves it where it is. git clean -fd deletes untracked files and folders. Run git clean -nd first to see what it would delete: files you haven’t added to git yet, including ones you wrote yourself, go too.

Claude Code has its own undo. It takes a checkpoint (a snapshot of the files it edits) before each prompt you send, and Esc twice or /rewind restores the code, the conversation or both. Its docs say plainly that this isn’t a replacement for git: it only tracks edits made by its file-editing tools, not files changed by shell commands like mv or a formatter. Use checkpoints for the last few minutes, commits for anything you’d be sorry to lose.

One step per prompt

With PLAN.md in the repo, each step’s prompt is short. From prompts/02-step.txt:

Read PLAN.md and do step 3 only.

Touch only the files step 3 lists. If it needs another file, stop and tell
me why instead of editing it.

When you're done, run `python -m pytest -q` and `git status --short`, show
me both outputs, and stop. Don't start step 4.

“Do step 3 only” and “don’t start step 4” both matter, because agents like to finish the job. Given the whole plan and a green test run, one can easily carry on into the next step, or tidy something from a later step while it’s in the file. That’s drift: work the step didn’t ask for. It’s usually well meant, and it breaks the one property that makes steps useful: each is small and checked.

python run_steps.py --drift swaps in a step 3 that also deletes STATUSES from models.py. The agent had just moved the two callers it was working on, so the old constant looked unused:

Step 3: Move search.py and report.py to tickets/status.py
  changed: tickets/models.py, tickets/report.py, tickets/search.py
  tests:   1 error  <- this step broke them
           ERROR tests/test_cli.py
           ImportError: cannot import name 'STATUSES' from 'tickets.models'
  rolled back to fc1af99 Step 2: Add tickets/status.py next to the old code

Two signals, before any code review: a file the plan doesn’t name for step 3, and a test error naming the exact problem, since cli.py still imports the constant until step 4. The rollback costs one step. Steps 1 and 2 are still committed, and the retry is the same prompt plus one line: “Don’t delete anything from models.py; that’s step 5.”

In a live session you can often stop drift before it lands. The agent writes a note before it edits, something like “models.STATUSES is no longer used, so I’ll remove it”. Press Esc to stop it mid-action and redirect. Short prompts in fresh sessions also keep the context small; Context Management covers why that matters.

Try two approaches in parallel with worktrees

Sometimes a step has two reasonable designs. In step 4 the CLI could keep argparse’s choices=STATUSES, or validate with type=check_status, which sounds more consistent since all validation then goes through the new module.

You could try one, roll back, try the other. Or run two agents at once, but not in one folder, where they’d overwrite each other’s edits. A git worktree is a second working folder for the same repo, on its own branch, sharing one history. Edits in one folder don’t touch the others.

git worktree add -b attempt-a ../kitebase-attempt-a
git worktree add -b attempt-b ../kitebase-attempt-b

Each makes a new folder at your current commit, on a new branch. Start an agent in each with the step prompt plus the design to try. python run_steps.py --attempts does this after step 3:

Attempt a in ../kitebase-attempt-a (branch attempt-a): Move cli.py to tickets/status.py, validating --status with check_status
  tests:   1 failed, 12 passed
           FAILED tests/test_cli.py::test_search_rejects_unknown_status
Attempt b in ../kitebase-attempt-b (branch attempt-b): Move cli.py to tickets/status.py
  tests:   13 passed
Kept attempt b: fast-forwarded main to 6d6aa42 Step 4: Move cli.py to tickets/status.py
Removed both worktrees.

Attempt A fails for a real reason. With type=check_status, a typo like --status archived gets argparse’s generic invalid check_status value: 'archived' instead of invalid choice: 'archived' plus the valid values. That’s a user-visible change, and an existing test catches it. Attempt B is committed on its branch, and git merge --ff-only attempt-b moves main to that commit. A fast-forward merge just moves the branch pointer, which works because nothing else changed main meanwhile.

MAIN CHECKOUT kitebase/ main at a3d6b47 (step 3 done) git worktree add, once per attempt WORKTREE A ../kitebase-attempt-a branch attempt-a --status type=check_status 1 FAILED, 12 PASSED argparse now says "invalid check_status value", not "invalid choice". A user-visible change. WORKTREE B ../kitebase-attempt-b branch attempt-b --status choices=STATUSES 13 PASSED: KEEP IT commit on attempt-b, then on main: git merge --ff-only attempt-b main is now 6d6aa42, step 4 Three folders, one .git: shared history, separate files, so the attempts can't touch each other. Afterwards: git worktree remove on both. With Claude Code, claude --worktree attempt-a makes .claude/worktrees/attempt-a on a new branch, worktree-attempt-a.
Two attempts at step 4, isolated from each other. The tests pick the winner.

Claude Code can create worktrees for you. At the time of writing, claude --worktree attempt-a (or -w) makes .claude/worktrees/attempt-a/ on a new branch, worktree-attempt-a, and starts the session inside it. On exit it usually removes a worktree with no changes and asks about one that has some. Three things to know:

  • It branches from your default branch, not from where you are. The new worktree starts from main as your remote has it, without your step 1 to 3 commits. Mid-refactor, set "worktree": {"baseRef": "head"} in .claude/settings.json, or create the worktree with git worktree add and run claude inside it.
  • It’s a fresh checkout. Untracked files like .env and your virtualenv aren’t there. Reinstall, or list gitignored files to copy in a .worktreeinclude file.
  • Add .claude/worktrees/ to .gitignore, so the folders don’t show up as untracked files.

Clean up with git worktree remove <folder> for each attempt and git branch -D for the losing branch. Parallel attempts cost two agent runs and two reviews, so save them for steps where you honestly don’t know which design wins.

Verify each step, not just the tests

A step can pass the tests and still be wrong. Before each commit, check four things:

  1. The test count matches the plan. A lower number can mean a test was deleted or skipped.
  2. The changed files match the plan. git status --short lists only the step’s files. The drift showed up here first.
  3. You read the diff. If it’s too long to read, the step was too big. Roll it back and split it.
  4. No callers are left before a delete. The grep before step 5 prints nothing.

At the end, run the real commands by hand (python -m tickets summary should print the same three lines as before) and review the branch commit by commit: five short diffs you already understand. Reviewing and Trusting AI-Written Code covers what to look for.

Try it yourself

The companion example is the Kitebase codebase with status spread over four files, PLAN.md, the two prompts, the five step patches and run_steps.py. It needs git and no API key.

Download the runnable example (zip)

cd 07-large-changes-and-refactors
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python run_steps.py              # the plan: 5 green steps, 5 commits
python run_steps.py --one-shot   # one big change: tests pass, the summary changes
python run_steps.py --drift      # step 3 breaks the tests and is rolled back
python run_steps.py --attempts   # step 4 tried two ways in two git worktrees

Every run works on a copy in a temporary folder, so kitebase/ never changes. Then try these:

  1. python run_steps.py --drift --keep /tmp/kb-drift, then run git log --oneline and git status in /tmp/kb-drift/kitebase. Steps 1 and 2 are committed and the tree is clean.
  2. Run with --keep again and add a fourth status, "blocked", to STATUSES and LABELS in tickets/status.py. The summary grows a Blocked 0 line and --status blocked works: a one-file change, the payoff of the refactor. Three tests fail because they pinned the old list; this time you changed the behaviour on purpose, so update them.
  3. Copy kitebase/ into a fresh git repo, start a real agent in plan mode, and paste prompts/01-plan.txt. Did its plan find all four files and add a pinning step?

pytest -q runs the offline tests: every step stays green and touches at most two files, the pinned test catches the one-shot bug, the drift is rolled back, and the worktree attempts clean up after themselves.

Common beginner mistakes

  • One prompt for the whole refactor. You get one diff too big to read, and a failure that could be anywhere in it.
  • Refactoring code no test covers. If nothing checks the behaviour you’re moving, green tests prove nothing. Pin it first.
  • Deleting the old code in the middle. Contract last, and only after a search shows no callers are left.
  • Not committing between steps. Then rollback means git reset --hard on work you wanted to keep, or untangling a half-good tree by hand.
  • Letting the agent carry on to the next step. “Stop after step 3” is part of the prompt, and the changed-file list is how you check it did.
  • Trusting checkpoints like commits. They don’t see files changed by shell commands, and they live in the tool’s session, not in your repo.

Questions you will face in production

“The change touches 200 files. Do I really do them one or two at a time?” No. Group identical mechanical changes into batches, one directory or module at a time, with tests and a commit after each. The other rules still apply: old code stays until the end, one batch per prompt. Try the prompt on two or three files first and fix it before running it on the rest. At the time of writing, Claude Code’s /batch command splits a large change across subagents (separate agent sessions it starts and coordinates), each working in its own worktree, and its docs show a shell loop over claude -p for scripted fan-out.

“Should I let the agent commit for me?” Letting it run git commit is fine, after you’ve read the step’s diff and seen the test output. Don’t give up the pause between steps: an agent that commits, starts the next step and pushes has removed every point where you’d have caught drift.

“One pull request for the whole refactor, or one per step?” For a change this size, one pull request with a commit per step, which reviewers read commit by commit. For a bigger one, land the expand step first (it changes no behaviour, so it’s easy to approve), then the migrations, then the delete. Every step leaves the tests green, so any of them can ship alone.

Check your understanding

An agent's refactor passes all the tests, but a CLI command's output changed. What was missing, and what would you add first next time?

A test that pins that output. The tests only prove the behaviour they check, and nothing checked the order of the summary lines. Next time, make step 1 of the plan a characterization test for every behaviour the refactor touches, and ask the agent which touched behaviour has no test.

Step 3 of your plan should change two files. The agent's step changed three, and the tests pass. What do you do?

Look at the third file before committing. It might be a harmless import fix, or it might be work from a later step, like deleting the old constant early. If it isn’t in the plan, roll the step back and rerun it with a line saying what not to touch. Passing tests don’t make drift safe; they just mean the drift hasn’t broken anything yet.

Why does the plan add tickets/status.py in its own step, before anything uses it?

So there’s a green commit where both the old and the new code exist. From there, each caller can move on its own, and each of those steps stays green because nothing has been taken away yet. If you added the module and moved every caller in one step, a mistake in any of them would fail the whole step.

You start two Claude Code sessions with claude --worktree to try two designs for step 4, and both are missing your step 3 changes. Why?

By default, claude --worktree branches from the repository’s default branch, not from your current commit. Your refactor branch’s commits aren’t on main yet. Set worktree.baseRef to "head" in the project’s settings, or create the worktrees yourself with git worktree add from your branch and start claude inside each one.

What to remember

  • Plan first, in plan mode, and save the plan to a file. Each step names its files and its expected test count.
  • Pin the behaviour you’re about to move with a test before you move it. Green tests only prove what they check.
  • Add the new, move the callers, delete the old. Every step keeps the tests green.
  • Commit after every green step. Rolling back a bad step is then git reset --hard plus git clean -fd, and it costs one step.
  • One step per prompt, and check the changed files against the plan. Drift shows up there first.
  • Use git worktrees to try two designs side by side, and let the tests pick.

What to study next

Big changes often need more than the code: the ticket that describes the change, the staging database, the CI status. MCP and Custom Tools in Your Editor shows how to give the agent tools for those systems. When several people run agents on the same repo, Team Workflows and Guardrails covers the shared rules that keep changes like this one reviewable.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.