Wrapping an API as an MCP Server
So far the Kitebase server has used sample data bundled with it. The real tickets live behind Kitebase’s REST API: a set of URLs (endpoints) you call over HTTP, like GET /v1/tickets/KITE-142, each returning JSON. Kitebase’s API has nine of them.
The quickest way to put that API in front of Claude is one tool per endpoint, maybe generated from the API docs. Then the support lead asks: “Northwind Studio says their SSO login is broken again. Is there a ticket for it, and who has it? If nobody does, give it to priya.” Claude calls list_customers to learn that Northwind is cus_4Qx1, list_tickets with that id, list_users to learn that priya is usr_7Hq2, and copies both ids into update_ticket, reading about 4,000 characters of JSON on the way. And the day the key expires, the API’s error says Invalid API key provided: kb_test_****, and Claude repeats it to the user.
None of that is the model’s fault. It’s how the server was shaped.
What you’ll build: an MCP server with three tools, find_tickets, get_ticket and assign_ticket, over a fake Kitebase REST API that runs inside your Python process. The fake has what a real API has: an API key, fat JSON, pages, error codes and a rate limit. The server answers the question in two tool calls, keeps the key from the model, and turns every failure into a sentence the model can act on. It runs offline.
If you haven’t built an MCP server yet, Setting Up Your First MCP Server covers MCPServer, @mcp.tool() and ToolError. This article uses all three.
The API you’re wrapping
Here’s the fake Kitebase API in the companion code’s fake_kitebase.py, shaped like most SaaS APIs:
GET /v1/tickets ?customer_id= &status= &q= &limit= &cursor=
GET /v1/tickets/{id}
PATCH /v1/tickets/{id} {"assignee_id": ...}
GET /v1/tickets/{id}/comments
POST /v1/tickets/{id}/comments
GET /v1/customers ?q=
GET /v1/customers/{id}
GET /v1/users
GET /v1/users/{id}
Every request needs the header Authorization: Bearer kb_test_4f9c2e7a1d. That’s a bearer token: whoever holds the string gets in, so treat it like a password. And every ticket comes back with 24 fields, most of which no one asked for:
{"id": "KITE-157", "object": "ticket", "number": 157,
"title": "SSO login loops back to the sign-in page",
"status": "open", "priority": "high", "assignee_id": null, "customer_id": "cus_4Qx1",
"reporter_id": "usr_9Lm4", "project_id": "prj_support", "channel": "email",
"sla": {"policy_id": "sla_business", "breach_at": "2026-09-26T20:12:44Z", "paused": false},
"custom_fields": {"cf_1029": "enterprise", "cf_2044": null, "cf_3187": ["eu"], "cf_4410": 3},
"internal": {"shard": "eu-3", "etag": "W/\"9dc2\"", "version": 7, "search_rank": null},
...}
That’s about 1,160 characters for one ticket, and it says cus_4Qx1, not “Northwind Studio”. The API is built for a programmer who joins ids up in code.
The fake runs in-process: httpx (the HTTP client used here) builds real requests and responses, but its MockTransport hands them to a Python function instead of the network. Point the same server at a real URL and nothing else changes.
Your server sits between the model and this API. Four things decide whether it works: which tools exist, where the key lives, how much comes back, and what a failure looks like.
Design tools around tasks, not endpoints
A programmer reads the docs, chains calls and keeps state between them. A model picks a tool from a list and fills in the arguments from the conversation. Every extra call is another round trip through the model, and every id it copies between calls is a chance to copy the wrong one.
So start from the jobs. For a support team they sound like “what’s open for Northwind about SSO?”, “what’s going on with KITE-142?” and “give KITE-157 to priya”. Three jobs, three tools:
Three rules make a tool task-shaped:
- Take the values a person would say.
find_ticketstakescustomer="Northwind", notcustomer_id="cus_4Qx1", andassign_tickettakesassignee="priya". The server looks up the ids. - Fan out inside the tool. One
find_ticketscall makes three HTTP requests (customers, tickets, users). - Hide what the task doesn’t need. Kitebase has four ticket statuses. Someone asking “what’s open?” means three of them, so the tool has
include_closed: bool = Falseinstead of a status list.
For the worked question, that’s two calls:
-> find_tickets {"customer": "Northwind", "query": "SSO"}
KITE-157 [open, high, unassigned] SSO login loops back to the sign-in page (updated 2026-09-26)
KITE-142 [in_progress, high, priya] Customer locked out after SSO change (updated 2026-09-24)
-> assign_ticket {"ticket_id": "KITE-157", "assignee": "priya"}
KITE-157 (SSO login loops back to the sign-in page) is now assigned to priya (Priya Shah).
Here’s how find_tickets turns a customer name into the id the API wants:
if customer:
matches = (await kb.request("GET", "/v1/customers", params={"q": customer}))["data"]
if not matches:
raise ToolError(f"No Kitebase customer matches {customer!r}. Try a shorter part of the name.")
if len(matches) > 1:
names = ", ".join(c["name"] for c in matches)
raise ToolError(f"{customer!r} matches several customers: {names}. Which one?")
params["customer_id"] = matches[0]["id"]
Resolving names is where the fiddly cases live: no match, or two. Say so, and the model can ask the user “Which one?” instead of guessing.
Default to three to six tools for an API wrapper. The opposite failure is one kitebase(action, params) tool that does everything: the model has to learn the whole API from one description, and you’ve just moved the mirror inside a function.
Why not generate the tools from the API's OpenAPI spec?
OpenAPI is a standard file format describing every endpoint of an API, and tools exist that turn one into an MCP server automatically. That gets you a mirror: one tool per endpoint, internal ids in every argument, full JSON in every result.
Fine for a prototype. For something people use every day, write the few task tools by hand. Every tool definition also sits in the model’s context on every request, used or not.
Names and descriptions the model chooses from
The model picks a tool from its name, its description and its argument schema, and nothing else. It never sees your code. The SDK builds all three from the function: the docstring becomes the description, and Annotated[..., Field(...)] from Pydantic adds a description or limits to an argument:
@mcp.tool(annotations=ToolAnnotations(read_only_hint=True))
async def find_tickets(
customer: Annotated[str, Field(description='Customer name or part of it, e.g. "Northwind".')] = "",
query: Annotated[str, Field(description='Words to match in the title or description, e.g. "SSO".')] = "",
include_closed: bool = False,
limit: Annotated[int, Field(ge=1, le=20)] = 5,
) -> str:
"""Find Kitebase support tickets by customer and keywords, most recently updated first.
Returns one line per ticket with its id, status, priority and assignee.
Use get_ticket for one ticket's description and comments.
"""
The server lists that as JSON, and the host hands it to Claude (trimmed here):
{"name": "find_tickets",
"description": "Find Kitebase support tickets by customer and keywords, most recently updated first. ...",
"input_schema": {"properties": {
"customer": {"type": "string", "default": "", "description": "Customer name or part of it, e.g. \"Northwind\"."},
"limit": {"type": "integer", "default": 5, "minimum": 1, "maximum": 20}, ...}}}
What makes that work:
- Name the product and the object. “Find Kitebase support tickets” can’t be confused with a GitHub issues tool the user also has installed. A tool called
searchcan. - Say what comes back. “One line per ticket with its id, status, priority and assignee” tells the model it can answer “who has it?” without another call.
- Point to the neighbour. “Use get_ticket for one ticket’s description” stops it calling
find_ticketsagain, hoping for more. - Show an example value and cap numbers.
e.g. "Northwind"says a partial name works.le=20meanslimit=500is rejected before your code runs.
read_only_hint=True is a tool annotation, a hint that the tool changes nothing. A host may use it to decide whether to ask the user before a call, so assign_ticket sets it to False. It’s a hint, not a guard. MCP Server Best Practices goes further on descriptions and on guarding tools that change things.
The gotcha: two tools whose descriptions overlap. Add a search_tickets next to find_tickets and the model has to guess, differently on different days. If you can’t say in one sentence when to use each, merge them.
Keep the API key out of the model’s view
Anything in the model’s context can end up in its answer or a transcript. So the key never goes there: not as a tool argument, not in a result, not in an error.
The server reads the key from an environment variable (a setting the process gets from whoever starts it), once, at startup:
def connect() -> Kitebase:
url = os.environ.get("KITEBASE_API_URL")
... # with no URL set, the server uses the in-process fake instead
key = os.environ.get("KITEBASE_API_KEY")
if not key:
sys.exit("KITEBASE_API_KEY is not set. The server needs it to call Kitebase.")
return Kitebase(http_client(url, key)) # sets "Authorization: Bearer <key>" on every request
A missing key stops the server at startup with a clear message, instead of a confusing failure on the first call. No tool has a parameter for it (the companion tests check that no schema mentions a key, token or auth). And tools never touch httpx directly: they call kb.request(...), the one place that handles HTTP, the key and errors.
Where does the variable come from? For a local server, from the host that launches it: host configs usually have a place for environment variables next to the command (check your host’s docs for the exact field). For a hosted server, from the platform’s secret manager. The code doesn’t change.
The gotcha is the error path. When the key is wrong, the fake API does what many real APIs do and echoes part of it back:
{"error": {"type": "authentication_error", "message": "Invalid API key provided: kb_test_****"}}
Pass that through and part of the key is now in the chat. So the server never forwards a 401’s message: it logs the detail for you and tells the model a fixed sentence, shown in the error section below.
Ask for one small page, and say when there’s more
Northwind has 23 tickets, 14 of them not closed. A bigger customer could have 4,000. APIs don’t return lists like that in one go: they paginate, returning one page at a time. Kitebase uses cursor pagination: each page comes back with "has_more": true and a next_cursor you send to get the next page.
The tempting version loops until has_more is false and hands the model everything. That’s rarely what the question needs, and one call can fill the model’s context window (the limit on how much text it can see at once). Instead, ask for as many rows as you’ll show, and use has_more to tell the model there’s more:
page = await kb.request("GET", "/v1/tickets", params=params) # params include limit=5
names = {u["id"]: name for name, u in (await kb.users()).items()}
lines = [ticket_line(t, names) for t in page["data"]]
if page["has_more"]:
# Say more exist instead of fetching every page into the model's context.
lines.append(f"(Showing the {limit} most recently updated. More match: narrow the search.)")
find_tickets(customer="Northwind") returns:
KITE-157 [open, high, unassigned] SSO login loops back to the sign-in page (updated 2026-09-26)
KITE-142 [in_progress, high, priya] Customer locked out after SSO change (updated 2026-09-24)
KITE-128 [waiting, low, unassigned] Dark mode chart labels unreadable (updated 2026-09-23)
KITE-131 [open, low, ana] Gantt view cuts off long task names (updated 2026-09-21)
KITE-101 [open, low, priya] Invite emails going to spam (updated 2026-09-18)
(Showing the 5 most recently updated. More match: narrow the search.)
The last line does real work. Without it, the model might tell the user “Northwind has five open tickets”. With it, the model knows the list is partial and narrows it with query="SSO", as in the worked example. The description says “most recently updated first”, so it also knows these five are the freshest.
Follow the cursor only when the task needs the whole set, like “how many open tickets does Northwind have?”. Then cap the number of pages and return a count or summary, not every row.
Trim the response to what the task needs
Even one page of five tickets is 5,677 characters of JSON. The model needs about a tenth of it. So the server picks the fields the task needs and writes them as one line per ticket:
def ticket_line(t: dict, names: dict[str, str]) -> str:
who = names.get(t["assignee_id"], "unassigned")
return f"{t['id']} [{t['status']}, {t['priority']}, {who}] {t['title']} (updated {t['updated_at'][:10]})"
main.py measures the difference for the same customer, estimating tokens (the unit models read and bill in) at about 4 characters each:
--- Trimming: the same kind of question, raw API page vs tool result ---
GET /v1/tickets?customer_id=cus_4Qx1 (23 tickets): 26,001 characters, about 6,500 tokens
find_tickets(customer='Northwind') (5 lines): 505 characters, about 126 tokens
Pagination and trimming together: about 50 times less text, with the answer to “who has it?” easier to find. get_ticket does the same for one ticket (status, assignee, customer, the first 400 characters of the description, the last three comments) in about 430 characters.
Lines rather than compact JSON is a style choice: lines don’t repeat the keys on every row, JSON is better if a program reads the result. Dropping the 18 fields nobody needs matters far more than the format.
The gotcha: trimming away something the next call needs. Every line keeps the ticket id because assign_ticket takes it. Before you cut a field, check whether another tool needs it as an argument.
Turn HTTP errors into sentences the model can act on
Every HTTP response has a status code: 2xx means it worked, 4xx means the request was wrong, 5xx means the API itself broke. The model can only react to what you tell it about each.
As in Setting Up Your First MCP Server, raise ToolError for failures you expect: the SDK turns it into a result with is_error: true and your message. Any other exception reaches the model as a bare Error executing tool get_ticket. So tool_error maps each status to one sentence saying what happened and what to do:
def tool_error(response: httpx.Response) -> ToolError:
status = response.status_code
error = response.json()["error"] # the full version also handles bodies that aren't JSON
log.warning("%s %s -> %s %s", response.request.method, response.request.url.path, status, error)
if status in (401, 403):
# Never pass the API's message on here: it can include part of the key.
return ToolError("Kitebase rejected this server's API key. That's a setup problem, not "
"something a retry fixes: tell the user to check KITEBASE_API_KEY.")
if status == 404:
return ToolError(f"{error['message']}. Ticket ids look like KITE-142; "
"use find_tickets if you don't know the id.")
if status in (400, 422):
return ToolError(f"Kitebase rejected the request: {error['message']}.")
...
Here’s what the model reads for a ticket that doesn’t exist:
-> get_ticket {"ticket_id": "KITE-999"} is_error: True
Error executing tool get_ticket: No such ticket: 'KITE-999'. Ticket ids look like KITE-142; use find_tickets if you don't know the id.
The SDK adds the Error executing tool get_ticket: prefix. The rest is yours, and it’s a next step instead of a dead end.
The split that matters is who can fix it. A 404 or 422 is about the arguments, so the API’s own message helps the model fix them. A 401 is the server’s setup, so the model should stop and tell the user. A 5xx is Kitebase’s problem: one retry, then tell the user.
Two more habits. Log the detail to stderr: log.warning records the status, the API’s error and its request_id for when you debug. (On a stdio server, stdout carries the MCP messages, so a stray print() breaks the connection.) And check what you can before calling the API. assign_ticket looks sam up in the cached user list and says who does exist, without sending a request:
-> assign_ticket {"ticket_id": "KITE-157", "assignee": "sam"} is_error: True
Error executing tool assign_ticket: No Kitebase user 'sam'. Usernames: ana, marco, priya.
Rate limits: wait a little, or hand it back
A rate limit caps how many requests one key can make in a time window, say 60 a minute. Go over and the API answers 429 Too Many Requests, usually with a Retry-After header saying how many seconds to wait. The model doesn’t know the quota exists, so the server handles it. Its request method waits out short limits and hands long ones back:
for attempt in range(3):
response = await self.http.request(method, path, **kwargs)
if response.status_code == 429 and attempt < 2:
wait = float(response.headers.get("Retry-After", 1))
if wait <= MAX_RETRY_WAIT: # 10 seconds
log.info("429 from Kitebase, waiting %ss and retrying", wait)
await self.sleep(wait)
continue
if response.is_success:
return response.json()
raise tool_error(response)
With a limit of 2 requests every 2 seconds, get_ticket (4 requests) hits it on the third, waits, and succeeds. The model gets the normal result and never knows:
kitebase: 429 from Kitebase, waiting 2.0s and retrying
-> get_ticket {"ticket_id": "KITE-142"}
KITE-142: Customer locked out after SSO change
...
HTTP behind it: GET /v1/tickets/KITE-142 200
GET /v1/tickets/KITE-142/comments 200
GET /v1/customers/cus_4Qx1 429
GET /v1/customers/cus_4Qx1 200
GET /v1/users 200
If Retry-After says 30, sitting silently inside a tool call is worse than saying so. The server raises “Kitebase is rate limiting this server. Wait 30 seconds before trying again.” and the model can tell the user.
The cheapest request is the one you don’t make. The user list barely changes, so kb.users() caches it for 5 minutes.
The gotcha is retrying writes. Retrying the PATCH after a 429 is safe, because a 429 means the API did nothing. After a timeout it isn’t always safe: the change may have gone through before the connection dropped. Setting an assignee is idempotent (doing it twice ends the same as once), but adding a comment isn’t, and a blind retry posts it twice. That’s why the server only retries 429s. Retries, Timeouts, and Failure Handling covers backoff and idempotency in depth.
Try it yourself
The companion example is the server, the fake Kitebase API, and a main.py that plays the host: it runs the worked example, prints the HTTP requests behind each call, measures the trimming and triggers each kind of failure.
Download the runnable example (zip)
cd 05-wrapping-an-api
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
The first thing it prints after the tool list is the worked example, with the requests behind it:
-> find_tickets {"customer": "Northwind", "query": "SSO"}
KITE-157 [open, high, unassigned] SSO login loops back to the sign-in page (updated 2026-09-26)
KITE-142 [in_progress, high, priya] Customer locked out after SSO change (updated 2026-09-24)
HTTP behind it: GET /v1/customers?q=Northwind 200
GET /v1/tickets?status=open,in_progress,waiting&limit=5&customer_id=cus_4Qx1&q=SSO 200
GET /v1/users 200
Set ANTHROPIC_API_KEY and Claude picks the calls itself. python server.py runs the server over stdio for a host like Claude Desktop, against the fake API unless KITEBASE_API_URL is set.
Then try these:
- Make
ticket_linereturnjson.dumps(t)(addimport json). The trimming line jumps from about 126 tokens to about 1,514 for the same 5 tickets. - In
tool_error, make the 401 branch returnToolError(error["message"]), the API’s own message. The wrong-key case now shows the modelInvalid API key provided: kb_test_****. - In
main.py, change the rate-limit demo toFakeKitebase(rate_limit=2, window=30, ...). The wait is now longer thanMAX_RETRY_WAIT, so the model is told to wait 30 seconds instead.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- One tool per endpoint. The model chains calls and copies internal ids between them. Write one tool per job, and let it take the names people say.
- The API key as a tool argument. It puts the key in the model’s context and in every transcript. Read it from the environment at startup.
- Forwarding every API error message. Some include part of the key. Pass them on for 400s and 404s, which are about the arguments; write your own for auth and server errors.
- Returning every page of raw JSON. One call can fill the context window. Ask for the rows you’ll show, say when more exist, and keep only the fields the task needs.
- Letting exceptions escape. A
KeyErrorreaches the model asError executing tool get_ticketand nothing else. RaiseToolErrorfor the failures you can predict.
Questions you will face in production
“How many tools should an API wrapper have?” As few as cover the jobs people ask for, usually three to six. Write down ten real requests from users, then group them.
“The API renamed a field and the tool broke. How do I catch that?”
It shows up as a KeyError in your trimming code, and the model sees a bare “Error executing tool”. Keep recorded responses from the real API in your tests, like the fake here, and rerun them before moving to a new API version. If the API lets you pin a version, pin it.
“Should the server act as one shared account or as each user?” Start with one key and a mostly read-only tool set, as here. Don’t give a shared key more power than every user of the server should have. Once different users should see different tickets, the server needs each user’s own token, which for hosted servers means OAuth (the standard “sign in and grant access” flow). Wherever the key comes from, never log it, and rotate it if it ever shows up in a transcript.
Check your understanding
A teammate adds a list_projects tool that returns every project with all 30 fields. What do you ask before merging?
What job it’s for. If people ask “which projects is Northwind in?”, build that, returning names only. If nobody asks about projects, skip it: its definition costs context on every request.
Claude keeps calling find_tickets three times in a row with the same arguments for "show me all of Northwind's tickets". What's likely wrong?
It gets five tickets and can’t tell whether that’s all of them. Check the result carries the “more match” note when has_more is true, and decide whether “all” should really be a count or a summary.
Kitebase starts returning 503 for an hour. What does the model see, and what should it tell the user?
“Kitebase had a problem on its side (HTTP 503). Try once more; if it fails again, tell the user Kitebase may be down.” After one retry it should say it can’t look up tickets right now, not invent an answer. Your stderr logs have the request ids for Kitebase’s support.
Why does assign_ticket check the username before calling the API, when the API would reject a bad id anyway?
The API’s error would be about an assignee_id, an id the model never saw. The server can say “No Kitebase user ‘sam’. Usernames: ana, marco, priya.”, which the model can act on, and it saves a request against the rate limit.
What to remember
- An API wrapper is a translation layer: task-shaped tools in front, the key and the HTTP details behind.
- Design one tool per job. Take the names people say and resolve ids in the server.
- The model chooses from names, descriptions and schemas only. Say what the tool does, what it returns, and when to use its neighbour.
- Read the key from the environment at startup. Never put it in an argument, a result or an error.
- Ask for one small page, say when there’s more, and trim each item to the fields the task needs.
- Map each HTTP failure to a
ToolErrorsentence with a next step. Wait out short rate limits in the server, and hand long ones back.
What to study next
The server works and fails politely. MCP Server Best Practices turns these patterns into habits for every server: descriptions the model can act on, validating arguments, guarding tools that change or delete things, and testing a server the way a client does.
Further reading
- Model Context Protocol: Tools. The official reference for tool definitions, input schemas, annotations and error results.
- MCP Python SDK. The
mcppackage used here, with examples ofMCPServertools. - HTTPX: Transports. How
MockTransportlets you test HTTP code without a network. - Anthropic: Writing effective tools for agents. Why fewer, well-described tools that return only what matters work better than a full API surface.
- MDN: HTTP response status codes. What each status means, including 429 and the
Retry-Afterheader.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.