What Is MCP? A Beginner’s Guide to the Model Context Protocol
Kitebase’s support lead wants to ask Claude one thing: “What’s going on with KITE-142, and is there a help article I can send the customer?” Claude can’t answer: it has never seen Kitebase’s tickets or help center, and has no way to look.
So you write glue for your support bot: a ticket-lookup function in the format Claude’s API expects. Then the team wants it in Claude Desktop, which plugs things in its own way. Then an engineer wants it in Cursor. Same lookup, three integrations, and every new app or system multiplies the work.
MCP, the Model Context Protocol, fixes that: you write the Kitebase side once, as a small program called an MCP server, and any app that speaks MCP can use it.
What you’ll build: an MCP server for Kitebase with two tools (search_help and get_ticket), plus a 100-line stand-in for Claude Desktop that launches it, shows what the model is given, and calls the tools for the question above. It runs offline; an API key lets Claude pick the calls itself.
Three words first: LLM, context, tool calling
If you’ve read the AI Engineering articles, skim this section.
An LLM (large language model) is a program that takes text in and writes text back. Claude is one. You call it over an HTTP API: send the conversation, get a reply (Your First LLM Integration shows the call).
The model’s context is everything it can see in one request: your instructions, the conversation, anything you pasted in. That’s all it knows about your situation, and it can’t go and fetch more. (What LLMs Are covers the context window, the limit on how much fits.)
Tool calling is how you get around that. With your request you send a list of functions the model may ask for, each with a name, a description and a JSON Schema for its arguments (a standard JSON format for “these fields, these types, these are required”). Instead of answering, the model can reply with a request like this:
{"type": "tool_use", "id": "toolu_01", "name": "get_ticket", "input": {"ticket_id": "KITE-142"}}
Your code runs get_ticket, sends the result back in the next request, and the model carries on with the real ticket in its context. The model never runs anything. It asks; your code decides. Giving an Agent Tools goes deeper on writing good tools.
The problem: every app wires every tool its own way
Tool calling solves “the model can’t see Kitebase” for one app. Your support bot defines get_ticket in Claude’s format and runs the loop, and that code lives inside the support bot.
Claude Desktop can’t use it. Neither can Cursor, or a teammate’s bot on another model provider, whose API calls the schema field parameters where Claude’s calls it input_schema. Each app needs its own Kitebase glue, and each other system you want to reach (GitHub, your Postgres) needs glue in each app too.
This is the N×M problem: N apps times M systems. Three by three is nine integrations to write and maintain, and a fourth app adds three more. With a shared protocol it’s N+M: each app speaks it once, each system implements it once. Six pieces, and a fourth app adds one.
Isn't 6 versus 9 a small win?
At three by three, yes. The gap grows fast: 10 apps and 20 systems is 200 integrations against 30.
The bigger win is who does the work. Without a standard, Kitebase waits for each AI app to build an integration, or builds one per app. With one, Kitebase ships one server and every MCP app can use it on day one.
What MCP is
MCP is an open protocol: a published set of rules for how an AI app and an outside program talk, not a product you buy. Anthropic released it as an open standard on November 25, 2024. The spec, official SDKs and docs live at modelcontextprotocol.io.
The spec names its inspiration: the Language Server Protocol, which lets any editor get Go autocompletion from one language server instead of every editor writing its own. MCP does the same for AI apps.
Messages are JSON-RPC 2.0, a small convention for sending requests as JSON: each request has a method name, params, and an id, and the response carries the same id back. Nothing in it is AI-specific, so any language with a JSON library can speak MCP.
The spec is versioned by date. The current revision is 2026-07-28, which version 2 of the official mcp Python SDK targets. The SDK puts that version string in every request for you.
Hosts, clients and servers
MCP has three roles. In the Kitebase setup:
- The host is the AI app you use: Claude Desktop, Claude Code, Cursor, or your own support bot. It holds the conversation, talks to the model, and decides which servers to connect.
- A client is the connection the host keeps to one server. Each client talks to exactly one server, so a host connected to two servers runs two clients. They’re built into the host; you rarely see them.
- A server is the program that offers capabilities: here,
server.py, which looks up Kitebase tickets and searches its help center. Servers are what you’ll mostly write. Hosts already exist.
Two things in that picture matter later. The model never talks to a server: the host sits in the middle of every call, which is where it can ask you before a tool runs. And a server only sees the calls sent to it. The spec’s design principles keep servers from reading the whole conversation or seeing into other servers, so the Kitebase server never learns what you asked GitHub.
How the bytes travel is the transport. The spec defines two:
- stdio: the host launches the server as a child process and writes one JSON message per line to its standard input; the server answers on standard output. This is how
server.pyruns. - Streamable HTTP: each message is an HTTP POST to one endpoint on a remote server, for when Kitebase hosts one server for all its customers.
The gotcha with stdio: standard output is the wire. The spec says a stdio server must not write anything to stdout that isn’t an MCP message. Add print("looking up", ticket_id) to a tool and that text lands where the host expects JSON. The Python SDK’s client logs Failed to parse JSONRPC message from server and skips the line; other hosts may not be so forgiving. Log to standard error instead: print(..., file=sys.stderr), or the logging module, which writes to stderr by default. MCP Architecture covers transports and the message flow in depth.
What a server offers: tools, resources, prompts
A server can offer three kinds of things. The spec’s split is about who decides to use each one:
| Primitive | Who decides to use it | Kitebase example |
|---|---|---|
| Tools | The model | get_ticket(ticket_id), search_help(query) |
| Resources | The app | kitebase://help/{slug}, the full text of a help article |
| Prompts | The user | triage_ticket, a ready-made instruction picked from a menu |
A tool is a function the model can call, the same idea as tool calling above. A resource is data the host can read and put in front of the model, addressed by a URI like a file path, such as kitebase://help/sso-login. A prompt is a template the user picks from the host’s menu, often shown as a slash command.
Here’s the heart of server.py, using the official mcp Python SDK:
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
mcp = MCPServer("kitebase")
@mcp.tool()
def get_ticket(ticket_id: str) -> dict:
"""Look up a Kitebase ticket by id, like KITE-142. Returns its title, status and assignee."""
ticket = TICKETS.get(ticket_id.strip().upper())
if ticket is None:
raise ToolError(f"No ticket {ticket_id!r}. Ticket ids look like KITE-142.")
return ticket
@mcp.resource("kitebase://help/{slug}")
def help_article(slug: str) -> str: ...
@mcp.prompt()
def triage_ticket(ticket_id: str) -> str: ...
if __name__ == "__main__":
mcp.run() # stdio by default
TICKETS is a dictionary standing in for Kitebase’s API, and search_help is written the same way. The decorator does the protocol work: the function name becomes the tool name, the docstring the description, and the type hints the JSON Schema. The model reads that description to decide when to call the tool, so write it for the model: what it does, what the input looks like, what comes back.
Default to tools. Most servers are mostly tools, and every MCP host supports them. Add resources when the app should attach data itself, and prompts only when your users would pick them from a menu. Tools, Resources, Prompts takes each one properly.
One question, end to end
Here’s the whole trip for the support lead’s question, with the real messages from the companion code:
Before the first question, the host asks the server what it offers with a tools/list request. The reply is each tool’s name, description and input schema. The host turns those into the model’s tool format. For Claude that’s three fields, which is most of what main.py does as a host:
def to_claude_tools(mcp_tools) -> list[dict]:
return [
{"name": t.name, "description": t.description or "", "input_schema": t.input_schema}
for t in mcp_tools
]
For get_ticket it prints:
{
"name": "get_ticket",
"description": "Look up a Kitebase ticket by id, like KITE-142. Returns its title, status and assignee.",
"input_schema": {
"type": "object",
"properties": {"ticket_id": {"title": "Ticket Id", "type": "string"}},
"required": ["ticket_id"],
"title": "get_ticketArguments"
}
}
That translation is why one server works everywhere: Kitebase describes get_ticket once, and each host converts it to whatever its model’s API wants.
When Claude replies with tool_use for get_ticket, the host sends a tools/call request down the stdio pipe. Here’s the real line, captured from the SDK’s client, with the _meta block (protocol version and client details, sent on every request) trimmed:
{"jsonrpc": "2.0", "id": 3, "method": "tools/call",
"params": {"name": "get_ticket", "arguments": {"ticket_id": "KITE-142"}, "_meta": {...}}}
And the server’s reply, the same id with a result:
{"jsonrpc": "2.0", "id": 3, "result": {"content": [{"type": "text",
"text": "{\"id\": \"KITE-142\", \"title\": \"Customer locked out after SSO change\", \"status\": \"in_progress\", \"assignee\": \"priya\", ...}"}],
"isError": false, "resultType": "complete"}}
The host passes that text to Claude as a tool_result. Claude asks for search_help next, the host makes one more tools/call, and Claude writes its answer from two real results: the ticket is in progress with Priya, and sso-login is the article to send. Nothing in the server knows which model asked, or which app it’s running in.
When a tool fails
Ask for a ticket that doesn’t exist and get_ticket raises ToolError. The SDK turns it into a normal result with isError: true:
-> tools/call get_ticket {"ticket_id": "KITE-999"}
is_error: True
Error executing tool get_ticket: No ticket 'KITE-999'. Ticket ids look like KITE-142.
That’s a tool execution error in the spec’s terms: the call worked, the tool didn’t, and the message goes back to the model so it can fix its input and try again. It’s different from a protocol error, when the request itself is broken, which comes back as a JSON-RPC error instead of a result.
The gotcha: raise any other exception, say a KeyError, and this SDK treats it as a crash. The model gets only Error executing tool get_ticket, with no hint of what to change, and the details stay in the server’s log. Raise ToolError with a message that says what a good input looks like whenever the model could do better on a second try.
Is MCP just tool calling with extra steps?
They’re different halves. Tool calling is how a model asks for a function; it’s part of each provider’s API. MCP is how a host finds and runs functions that live in someone else’s program, the same way for every server.
The two meet in the host, which turns MCP tool definitions into tool-calling definitions, and tool-calling requests into tools/call messages. MCP also covers what tool calling doesn’t: resources, prompts and transports.
Try it yourself
The companion example is server.py, the Kitebase server, and main.py, which plays the host: it launches the server over stdio just as Claude Desktop would, lists what it offers, and calls both tools for the worked question.
Download the runnable example (zip)
cd 01-what-is-mcp
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
It prints the server’s tools, resources and prompts, the get_ticket definition shown above, and the two tool results: the KITE-142 ticket, then sso-login and account-recovery.
Then try these:
- In
call_like_claudeinmain.py, changeKITE-142toKITE-999. You get theToolErrormessage back as a result, not a crash. - Add
print("looking up", ticket_id)toget_ticketinserver.pyand runmain.pyagain. You’ll seeFailed to parse JSONRPC message from server: your print went down the wire. Change it toprint("looking up", ticket_id, file=sys.stderr)(andimport sys) and it shows up as a harmless log line. - Set
ANTHROPIC_API_KEYand run it again. Now Claude gets the tool list and picks the calls itself, andmain.pyprints each one before it runs. A run costs a few cents at most.
To use the server from a real host, point the host at the Python that has mcp installed. Claude Desktop’s claude_desktop_config.json and Claude Code’s project .mcp.json both use this shape:
{
"mcpServers": {
"kitebase": {
"command": "/full/path/to/01-what-is-mcp/.venv/bin/python",
"args": ["/full/path/to/01-what-is-mcp/server.py"]
}
}
}
Restart the host and ask it the KITE-142 question. Connecting to Claude Desktop, Cursor, Etc. covers each host’s setup.
pip install pytest && pytest -q runs the offline tests. They call the tools as plain functions, connect to the server through the SDK’s in-process client, and launch it over stdio once. No key, no network.
Common beginner mistakes
- Printing to stdout in a stdio server. Your debug line lands in the JSON stream. Log to stderr.
- Relative paths in the host config. The host doesn’t launch your server from your project folder, so
python server.pycan’t find the file, or finds a Python withoutmcpinstalled. Use absolute paths to both. - Copying 1.x examples. Many tutorials start with
from mcp.server.fastmcp import FastMCP. That’s the 1.x SDK; in 2.x the class isMCPServerand the old import fails with an error pointing at the migration guide. The companion pinsmcp>=2.2,<3. - Vague tool descriptions. “Gets a ticket” leaves the model guessing when to call it and what the id looks like. Say what it does, what the input looks like, and what comes back.
- Plain exceptions for expected failures. The model only sees
Error executing tool get_ticket. RaiseToolErrorwith a message it can act on.
Questions you will face in production
“Is it safe to install someone else’s MCP server?” Treat it like any program you run: a stdio server runs with your permissions and can do anything its code does. The spec calls tools arbitrary code execution and says hosts should keep a human in the loop who can deny a call. Install servers from sources you trust and prefer hosts that ask before running a tool. Tool descriptions are text the model reads, so a hostile server can hide instructions in them.
“Local stdio server or remote HTTP server?” Start with stdio: nothing to deploy. Move to Streamable HTTP when many users need the same server and you don’t want each of them installing it, like Kitebase offering one server to all its customers. That brings authentication with it; Connecting to Claude Desktop, Cursor, Etc. shows how hosts connect to remote servers.
“Should I write an MCP server or just use tool calling in my app?” If the only consumer is your own app, plain tool calling is less machinery. Write a server when more than one host should use the same capability, or when you want people to use your product from the AI apps they already have.
Check your understanding
Your company uses Claude Desktop, Cursor and an in-house support bot, and wants all three to read from Jira and your Postgres. How many pieces do you write with MCP, and how many without?
Without a shared protocol, 3 × 2 = 6 integrations. With MCP, the apps already have clients (your bot gets one from the SDK), so you write or install 2 servers: Jira and Postgres. A fourth app costs nothing on the server side.
Claude calls get_ticket with "142" and gets back only "Error executing tool get_ticket". It tries "142" again. What should the server do differently?
The tool raised a plain exception, which the SDK treats as a crash and hides. Raise ToolError("No ticket '142'. Ticket ids look like KITE-142.") instead: the message reaches Claude as an isError result, and it can retry with KITE-142.
The Kitebase server is connected. Can it read what the user asked the GitHub server earlier in the chat?
No. Each server has its own client and only sees the calls the host sends it, like tools/call get_ticket. The conversation stays in the host, and servers can’t see into each other.
You want Claude to be able to look up a customer's plan and seat count on its own. Tool, resource or prompt?
A tool, say get_customer(customer_id), because the model decides when it needs the data. A resource needs the app to attach it; a prompt is a template the user picks.
What to remember
- MCP is an open protocol that lets any AI app use any MCP server, turning N×M integrations into N+M.
- The host is the AI app and holds the conversation, with one client per server. The server is what you write.
- Servers offer tools (the model decides), resources (the app decides) and prompts (the user decides). Start with tools.
- Messages are JSON-RPC 2.0.
tools/listfinds the tools;tools/callruns one. Stdio for local servers, Streamable HTTP for remote ones. - The model never talks to a server. The host translates tool definitions and calls between the model’s API and MCP.
- In a stdio server, stdout is the wire. Log to stderr, and raise
ToolErrorwith messages the model can act on.
What to study next
Next, MCP Architecture: Clients, Servers, Transports opens up the pipe between client and server: the JSON-RPC messages and the two transports in detail. Then Setting Up Your First MCP Server walks through building and registering a server step by step.
Further reading
- Model Context Protocol specification. The source for the roles, the three primitives and the transports described here. The overview and architecture pages are short.
- Anthropic: Introducing the Model Context Protocol. The original announcement and the problem it set out to solve.
- MCP Python SDK documentation. The SDK the companion code uses, including the migration guide for older examples that use
FastMCP. - Server features in the spec. The table of who controls tools, resources and prompts, with links to each.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.