Choosing Your Tool: Cursor, Claude Code, Devin Desktop

You want a coding agent to do a small job on Kitebase’s ticket service: add a --status filter to search_tickets. First you have to pick a tool, and every thread says something different. The names don’t even hold still: Devin Desktop, one of the three in this article’s title, was called Windsurf until June 2, 2026, according to its own FAQ. Product comparisons go stale fast, this one included.

What doesn’t go stale is how you work. Do you live in an editor or a terminal? Do you want the next line finished, or the whole task done? Is your company fine with code leaving your laptop? Will a team share the setup? Answer those, then test the tools that survive on a real task.

What you’ll build: a small evaluation kit: the Kitebase ticket service, the task as a prompt you paste into each tool, hidden acceptance tests, and a scorecard script. You run the same task in two tools, score both the same way, and pick from your own numbers.

Why a feature table doesn’t help

A grid with a row per feature and a column per product fails for two reasons. First, the ticks converge. At the time of writing, every product here ships more than one surface (a place you use the tool from: an editor, a terminal, a plugin, a web app). Cursor is an editor that also has a terminal agent (Cursor CLI). Claude Code started in the terminal and also has VS Code and JetBrains extensions and a desktop app (Claude Code platforms). GitHub Copilot started as editor autocomplete and now has a terminal agent too (Copilot CLI). A grid says they’re the same. They aren’t.

And the grid ages. Every product fact in this article comes from the product’s own docs or, for Claude Code, the installed CLI (claude --help, version 2.1.281), checked in September 2026. Treat each one as true at the time of writing and check again before you rely on it.

So this article asks four questions about you instead:

  1. Where do you type? Editor-first or terminal-first.
  2. How much do you hand over? Autocomplete or agent.
  3. Where do your code and data go? Privacy and company policy.
  4. How will a team share it? What transfers between tools.

Question 1: where do you type?

Editor-first means the AI lives inside your editor: a side panel for chat and agent tasks, suggestions as you type, changes shown as inline diffs you accept hunk by hunk (a hunk is one block of changed lines). Terminal-first means you start the agent in a shell in your project directory, describe the task, and review the finished diff afterwards, in whatever editor you like.

Here’s the Kitebase task done both ways:

Editor-first you drive, the AI assists 1. YOU OPEN kitebase/tickets.py and start typing the new parameter def search_tickets(tickets, query, status= 2. AUTOCOMPLETE SUGGESTS None): ... if status is not None: ... grey text ahead of your cursor, Tab to accept 3. YOU ASK THE SIDE PANEL "now add --status to the CLI" an inline diff appears in __main__.py 4. YOU ACCEPT THE HUNKS, RUN TESTS python -m pytest -q 4 passed Terminal-first you hand over, then review 1. YOU START THE AGENT ~/kitebase-claude-code $ claude and paste the prompt from TASK.md 2. IT READS THE CODE tickets.py __main__.py test_tickets.py plus AGENTS.md, for the project's rules 3. IT EDITS AND RUNS THE TESTS 3 files changed, +18 -5 lines python -m pytest -q 4 passed 4. YOU REVIEW THE WHOLE DIFF in your own editor, or with git diff You see every line as it's written. Best when you already know the change. You see the result, not each step. Best when the task is clear and tested.
Same task, same result. What differs is how often you're in the loop.

Neither is better. The editor shows every line as it’s written, which suits changes you could mostly write yourself. The terminal shows the result, which suits clear tasks with tests to check them.

Since most products offer both, ask which surface is the product’s home, usually the most complete one. At the time of writing:

ProductHome surfaceAlso available as
CursorIts own editor, which imports your VS Code settings and keybindingsA terminal agent (agent), JetBrains IDEs via ACP
Claude CodeThe terminal (claude)VS Code and JetBrains extensions, a desktop app, web, mobile
Devin Desktop (was Windsurf)Its own editor, which can import VS Code or Cursor settingsDevin CLI, JetBrains via ACP, plugins for other editors
GitHub CopilotExtensions for VS Code, Visual Studio, JetBrains IDEs, Xcode and EclipseA terminal agent (copilot)

ACP, the Agent Client Protocol, is a standard way for an editor to host an agent someone else built (Cursor in JetBrains, Devin in JetBrains).

The gotcha: the other surfaces are real but thinner. Claude Code’s docs say scripting and the Agent SDK are CLI-only. Cursor in JetBrains needs a paid plan and JetBrains’ AI Assistant plugin. The old Windsurf JetBrains plugin is in maintenance mode. And Cursor uses the Open VSX extension registry, not the VS Code Marketplace, so check your must-have extensions first (Cursor: migrate from VS Code).

Default: pick a tool whose home is where you already work. If your editor setup took years to tune, a terminal agent or a plugin adds AI without touching it. If you’d happily switch editors, an AI-first editor gives the tightest integration.

Question 2: how much do you hand over?

Article 01: How AI Coding Tools Actually Work described the two modes. Autocomplete makes one model call per suggestion: here’s the code around my cursor, guess what comes next. Agent mode runs the full loop: read files, edit, run commands, check the output, repeat until the model decides it’s done.

On the Kitebase task they help at different moments. You type status: str | None = None, autocomplete offers the filter lines as grey text, you press Tab. It can’t “also add the CLI flag, and a test”. The agent takes the whole task and comes back with this, plus the flag and the test:

def search_tickets(
    tickets: list[Ticket], query: str, status: str | None = None
) -> list[Ticket]:
    q = query.lower()
    matches = [t for t in tickets if q in t.id.lower() or q in t.title.lower()]
    if status is not None:
        matches = [t for t in matches if t.status == status]
    return matches

That’s from demo/demo-a in the companion kit, a run I wrote by hand to show a good result: 3 files, 18 lines added, 5 removed.

Where the products sit, at the time of writing:

  • Cursor has both: Tab, its autocomplete, which can edit several lines at once (Cursor Tab), and its agent.
  • Devin Desktop has both: Tab and Autocomplete, plus a local agent. That was Cascade; the FAQ says its replacement is called Devin Local.
  • GitHub Copilot has both: inline suggestions as you type, plus agent mode and an agent that prepares pull requests for review.
  • Claude Code is an agent on every surface. Its docs describe one engine behind the CLI, desktop app and IDE extensions, and don’t list as-you-type completions. If you want both, pair it with your editor’s autocomplete.

Also ask whether the tool must run headless: with nobody watching, in CI or a script. Claude Code, the Cursor CLI and the Copilot CLI all have a mode that prints a result and exits:

claude -p "Explain what search_tickets does" --output-format json   # Claude Code
agent -p "Explain what search_tickets does"                         # Cursor CLI
copilot -p "Explain what search_tickets does"                       # Copilot CLI

The flags come from claude --help, Cursor’s headless docs and Copilot’s CLI docs. Each has its own rules for what it may do unasked in this mode; read them before putting one in a pipeline.

The gotcha: each mode lets bad code in its own way. With autocomplete, it’s pressing Tab on grey text you didn’t read. With an agent, it’s a one-line request that edits ten files. Neither replaces reading the diff, which article 06: Reviewing and Trusting AI-Written Code covers.

Question 3: where do your code and data go?

This one can rule a tool out however much you like it. Unless you run a model on your own hardware, your code goes to someone else’s servers. Here’s what goes out for the Kitebase task:

Your machine 1. YOUR REPO kitebase/tickets.py kitebase/__main__.py tests/test_tickets.py AGENTS.md .env DATABASE_URL=… in .gitignore and ignore files 2. THE REQUEST your prompt (TASK.md) AGENTS.md tickets.py, __main__.py command output rebuilt on every step THE GAP: COMMANDS THE AGENT RUNS grep -rn DATABASE . prints the .env line Ignore files and deny rules don't cover it. The output joins the request. 3. TOOL VENDOR the company that makes the tool sometimes skipped 4. MODEL PROVIDER runs the model or your own cloud account Ask of each: is it stored, for how long, and is it used for training?
What leaves your laptop for one task, and the gap most people don't know about.

The request holds the files the agent read, your rules file and every command’s output, sent again on each step of the loop. Cursor’s docs say prompts and code context go to model providers such as OpenAI, Anthropic and Google (Cursor: privacy and data governance). Claude Code’s docs say it sends all prompts and model outputs over the network, encrypted with TLS 1.2 or later (Claude Code: data usage).

For each candidate, find four answers in its docs:

  1. Who receives it? The tool vendor, the model provider, or both. Some tools can use your own cloud account instead: Claude Code can run through Amazon Bedrock or Google Cloud.
  2. Is it used for training? Often that depends on the plan. Claude Code’s docs say Anthropic doesn’t train on code sent under commercial terms (Team, Enterprise, API); on consumer plans (Free, Pro, Max) it’s your choice in settings. Cursor says code is never used for training with Privacy Mode on, which is the default for Enterprise teams.
  3. How long is it kept? Look for zero data retention (ZDR): the provider processes the request and doesn’t store it. Check the exceptions: Cursor lists a few models outside its ZDR agreements. Copies can sit on your own disk too: Claude Code keeps session transcripts in plain text under ~/.claude/projects/ for 30 days by default.
  4. Does any feature store your repository? Cloud agents need your code on the vendor’s machines. Cursor says Cloud Agents are its only feature that stores code, and they’re optional.

Take the answers to whoever owns security at your company before you point a tool at company code. The classic mistake is a personal account on a work laptop, on a plan whose training default nobody checked; Cursor even has a device policy, Allowed Team IDs, to block exactly that.

The gap: ignore files are not a wall

Every tool has a way to hide files: .cursorignore, .devinignore, Copilot’s content exclusion, Claude Code’s deny rules. List .env there and you’d expect the password to stay out of the request. It stays away from the tool’s file-reading tools. The agent can still run a command that prints it. Each tool’s docs say so:

  • Cursor: terminal commands and MCP tools run outside its file access controls, so they may still read ignored files (Cursor: ignore files).
  • GitHub Copilot: agent mode in Copilot Chat in IDEs doesn’t support content exclusion (Copilot: excluding content).
  • Claude Code: deny rules cover its file tools and commands it recognises, like cat, but not grep -r pattern . or a script that opens files itself (Claude Code: permissions).

So an agent looking for the database config runs grep -rn DATABASE ., the output includes the .env line, and it’s in the next request. The fix is to not have the secret there. Keep production credentials off the machine the agent runs on, use a local database with a throwaway password, and treat ignore files as a way to cut noise, not as access control. For a real boundary, use OS-level sandboxing; Claude Code and the Devin CLI both document one.

Question 4: how will a team share it?

Once three people use AI tools, their setups drift: one agent knows the project runs pytest, another writes unittest classes. Split the setup into what transfers between tools and what doesn’t.

What transfers. A rules file is plain text the tool pastes into the context window at the start of each session; article 03 covers them properly. The cross-tool one is AGENTS.md, and at the time of writing all four products read it: Cursor, Claude Code (alongside or instead of CLAUDE.md), Copilot and Devin Desktop. The Kitebase project’s, trimmed:

# Kitebase tickets

- Run the tests with `python -m pytest -q`. They must pass before you finish.
- Don't change the assertions in existing tests. Add new tests instead.
- Keep changes to the files the task needs.

Your tests and your review bar transfer too. They don’t care which tool wrote the code.

What doesn’t. Permissions, allowed models, MCP servers and privacy settings live in each tool’s own config. Claude Code’s docs say to commit .claude/settings.json so everyone who clones the repo gets the same permissions (Claude Code: settings). Cursor lets an admin force Privacy Mode on for the team. Copilot lets organisation owners choose features and excluded files.

Switching is cheaper than it used to be: claude import copies config from Codex, Gemini or Cursor, and Devin Desktop can import VS Code or Cursor settings.

Default for a team: one approved tool per company (question 3 usually settles it), free choice of surface within it, and conventions in AGENTS.md so nobody is locked in. Article 09: Team Workflows and Guardrails goes further.

Decide with your own results

The four questions usually leave two candidates, and reviews won’t separate them. What matters is how each does with you, on your kind of code, so give both the same task and score them the same way.

PROJECT/ untouched Kitebase scores 2/10 setup COPY A ../kitebase-tool-a no eval/ folder COPY B ../kitebase-tool-b no eval/ folder TOOL A same TASK.md 11 min, 0 nudges TOOL B same TASK.md 6 min, 1 nudge MAIN.PY SCORE 5 hidden acceptance tests the copy's own tests added a test? left old ones alone? only touched files the task needs? your minutes and nudges A 10/10 B 7/10 B missed validation, edited a test Copies sit outside the kit, so the agent can't read the tests it's scored on. Faster isn't better if you have to fix what it skipped.
The kit's loop. The A and B numbers are the hand-made demo runs, not results from real tools.
  1. python main.py setup <tool> copies the untouched project to ../kitebase-<tool>, outside the kit, so the agent can’t find the acceptance tests and code to them.
  2. Open that folder in the tool, paste the prompt from TASK.md word for word, and time it. Each follow-up you send (“you forgot the test”) is a nudge.
  3. python main.py score <tool> --minutes 11 --nudges 0 runs five acceptance tests from eval/, runs the copy’s own tests, and diffs the copy against the untouched project.

The core of the scoring is a few checks on that diff:

checks = {
    "its own tests pass": own_exit == 0,
    "added a test for the filter": diff["tests_added"] > 0,
    "left existing tests alone": diff["test_lines_removed"] == 0,
    "only touched files the task needs": not diff["out_of_scope"],
    f"needed at most 1 follow-up prompt ({nudges})": nudges <= 1,
}

Each catches a failure you’d otherwise miss: an agent can turn its own tests green by loosening one, or do work nobody asked for. python main.py with no arguments scores the two demo runs:

Scorecard: demo-b
  [4/5] acceptance tests pass
  [ok]  its own tests pass
  [ok]  added a test for the filter
  [--]  left existing tests alone
  [--]  only touched files the task needs
  [ok]  needed at most 1 follow-up prompt (1)
        outside the task: README.md
  Score 7/10. 4 files changed, +18 -4 lines, 6 min.

                                            demo-a        demo-b
acceptance tests                               5/5           4/5
left existing tests alone                       ok            --
only touched files the task needs               ok            --
minutes                                         11             6
score                                        10/10          7/10

(Comparison trimmed.) demo-b was nearly twice as fast and its own tests are green. But it accepts --status done instead of rejecting it, loosened an assertion in an existing test, and edited the README. Whether five saved minutes is worth that is your call; the scorecard makes it visible, where “tests pass” wouldn’t.

The gotcha is the floor. An untouched copy passes 2 of the 5 acceptance tests, because two check that old behaviour still works. So the other checks only count once a copy passes more than that: a tidy diff that doesn’t do the task is worth nothing. It’s the same reason LLM evaluation pipelines compare every score with a baseline: a number means little until you know what it’s measured against.

And one run proves little. Agents don’t produce the same output twice, so run the task twice per tool in fresh copies, then once more on a real ticket from your codebase where you know the right diff. Give both tools the same AGENTS.md, so you compare tools, not setups.

Try it yourself

The kit runs offline and calls no model. The tools you compare do that, with your own accounts.

Download the runnable example (zip)

cd 02-choosing-your-tool
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py                       # the demo scorecards
python main.py setup cursor          # then open ../kitebase-cursor in Cursor
python main.py setup claude-code     # and ../kitebase-claude-code in Claude Code

Then try these:

  1. python main.py setup baseline, then python main.py score baseline --minutes 0 --nudges 0. You get 2/10 and a note that the checks don’t count. That’s the floor.
  2. Run the task in two tools, score each, then python main.py compare. Read both diffs yourself before you believe the numbers.
  3. In one run, drop the prompt’s last line (“Run the tests before you say you’re done”). See whether that tool runs them anyway.

pytest -q runs the kit’s offline tests. CHECKLIST.md turns the four questions into a list to fill in per tool.

Common beginner mistakes

  • Choosing from a feature grid. Every product ticks most boxes. The grid can’t tell you how a tool fits your day.
  • Trying tools on different tasks. That compares tasks, not tools. Same prompt, same project, same rules file.
  • Trusting “tests pass”. The agent can edit its own tests. Check against tests it never saw.
  • Using a personal account on company code. Its training and retention defaults may not be what your company agreed to. Ask first.
  • Treating ignore files as security. They hide files from the tool’s file reads, not from commands. Keep secrets out of the working directory.

Questions you will face in production

“Which one is the best?” It depends on where you type, how much you hand over and what your company allows, so answer the four questions and run the kit on the tools that pass. If two score the same, pick the one whose home surface you’d rather spend the day in.

“Security hasn’t approved any AI coding tool. Can I try one anyway?” On code that isn’t your company’s, yes: that’s what the Kitebase kit is for. Bring the four data answers to the security review, with links to each vendor’s docs. It turns “is AI safe?” into a list they can check.

“Should the whole team use the same tool?” Standardise what has to match: the approved vendor, privacy settings, AGENTS.md, the tests and the review bar. Let people choose the surface. Diffs from two surfaces of one tool, following the same rules, look the same in review.

Check your understanding

You live in PyCharm and don't want to change editors. Which question narrows your options most, and what do you check?

Where you type. Check each candidate’s JetBrains support and whether it’s the home surface or a thinner one. At the time of writing, Copilot and Claude Code have JetBrains integrations, and Cursor and Devin reach JetBrains through ACP. A terminal agent also works next to PyCharm without touching it.

Your repo's .env is listed in .cursorignore. Can its contents still reach the model provider?

Yes. The ignore file blocks the agent’s file-reading tools, but Cursor’s docs say terminal commands run outside those controls. If the agent runs a command whose output includes the file, like grep -rn DATABASE ., that output goes into the next request. Move the secret out of the working directory.

Tool A scores 10/10 in 11 minutes, tool B 9/10 in 5 minutes, one run each. What next?

Nothing’s decided yet. Read both diffs to see what B missed, then run each again in fresh copies: one run can’t separate a one-point gap from run-to-run variation. Then try both on a real ticket from your codebase.

What to remember

  • Every major product ships several surfaces now. Pick by how you work, not by which boxes it ticks.
  • Four questions: where you type, how much you hand over, where your code goes, and how a team shares it.
  • The request holds your prompt, the files read and command output, on every step. Know who receives it, whether it’s trained on and how long it’s kept, on your plan.
  • Ignore files and deny rules aren’t a security boundary. Keep secrets out of the agent’s working directory.
  • Put conventions in AGENTS.md, which all four products read, so switching stays cheap.
  • Choose from your own results: same task, fresh copies, tests the agent never saw, and a baseline.

What to study next

Whichever tool you pick, what decides output quality next is what you tell it about your project. Article 03: Setting Up Context covers rules files like AGENTS.md and CLAUDE.md properly: what to put in, what to leave out, and the ignore and indexing settings that shape what the agent sees. It works the same in every tool here.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.