Selling Your MCP Server
Kitebase put its MCP server online over Streamable HTTP, so customers’ agents could look up tickets and search the help center without installing anything. A month later the hosting bill has tripled. The logs show search_help called thousands of times overnight, but not by whom. You can’t ask that customer to slow down, you can’t cap them, and you can’t send them a bill.
A server anyone can call is a cost. Before you can charge for it, it has to know who is calling, how much they’re allowed and how much they used. That’s a few hundred lines of ordinary code. This article builds it, and covers the other half: who would pay, and how to pick a price.
What you’ll build: a paid front door for the Kitebase server from article 01: What Is MCP?: an API-key check, per-plan rate limits and monthly quotas, a usage meter per key, and a September bill for three made-up customers. It runs offline: real MCP over HTTP, in one process.
Who pays, and for what
The Kitebase server is about 100 lines of Python that anyone could write. What they can’t copy is the data behind it. So you’re not selling “an MCP server”; you’re selling access to something, and the server is a new door into it. That gives three kinds of buyer:
- Your existing customers. Northwind Studio already pays for Kitebase, and the MCP server is a feature of the product. You might not charge extra, but you still need keys and limits, or one customer’s agent slows the server down for everyone.
- Developers who want your data from their agents. Hollow Pine’s support agent needs Kitebase’s help center. They pay per call, like any metered API.
- Teams who’d rather not run it. If your server is open source, some will pay you to host it and answer email when it breaks. That’s open core: free code, paid hosting.
The person who installs the server is rarely the one who pays: Priya pastes the config, Northwind’s admin gets the invoice. And some servers shouldn’t be sold at all. An internal server has no buyer, and a thin wrapper over a free public API has nothing behind it people can’t get directly. Give those away.
Why can't a local stdio server charge per user?
A stdio server runs on the user’s machine as a subprocess of their host (article 02 covers transports). There’s no login and no way to tell one user from another; the MCP spec says stdio servers get credentials from the environment, and those are the user’s own. Charging needs a server you run, over the network, that sees each caller as an account: a remote Streamable HTTP server.
Pricing models, with real numbers
There are three basic ways to charge, and one plan can mix them:
- Per seat: a fixed price per person (per API key, here) per month. Predictable for the buyer, and you earn from light users.
- Usage-based: a price per unit of use, like per 1,000 tool calls. What customers pay tracks what they cost you to serve.
- A free tier: a small allowance for nothing, so people can try the server before anyone signs off on a purchase.
A common mix is included usage plus overage: the seat price includes some calls, and calls past that cost extra. Here are Kitebase’s plans for the worked example, with prices made up to show the mechanics:
| Plan | Price | Included | Past that | Rate limit |
|---|---|---|---|---|
| Free | $0 | 100 calls a key per month | refused | 10 calls a minute |
| Pay as you go | $0 | nothing | $4 per 1,000 calls | 60 a minute |
| Team | $15 a key per month | 2,000 calls a key | $4 per 1,000 calls | 60 a minute |
Here’s what each customer would pay for September’s usage on each plan:
Northwind would pay less on Pay as you go, so why offer Team? A fixed monthly price is easier to approve than a bill nobody can predict, and seat fees earn from teams that barely call anything.
To pick numbers, start from your cost per call. Say hosting plus Kitebase’s backend comes to $0.0008 a call (a made-up figure; divide your own hosting bill by your call count). At $4 per 1,000, a call sells for $0.004: five times cost, which leaves room for support, failed calls and the free tier.
The gotcha: the model decides how many calls to make, not the customer. An agent that retries a search 40 times on one question costs 40 calls. That’s why every plan here has a rate limit. Default to a free tier plus one paid plan, and add a second when customers ask.
The request path: who, then what
Every tool call needs three answers: who is calling, whether their plan allows it, and what to record for the bill. The example wraps the Kitebase server in two layers of middleware, code that runs around every request without the tools knowing:
The layers split the work by what each one can see:
- HTTP layer: who is this? It reads the
Authorizationheader and can refuse a request before any MCP happens, but it can’t see which tool is being called. - MCP layer: may they, and what did it cost? It sees the method and tool name, so it can limit and meter tool calls and leave
tools/listalone.
Wiring them up is a few lines in gateway.py:
def build_app(mcp_server, gateway: Gateway):
mcp_server.middleware.append(gateway.meter_tool_calls)
inner = mcp_server.streamable_http_app(stateless_http=True, json_response=True)
return inner, gateway.check_api_key(inner)
streamable_http_app() turns the MCPServer into an ASGI app, the interface Python’s async web servers use to call your code. stateless_http=True suits the 2026-07-28 spec revision, where every request stands alone (article 02) and so carries its own key. check_api_key(inner) wraps that app, so every HTTP request hits the key check first.
Step 1: API keys, to know who’s calling
An API key is a long random string that identifies one caller. The host sends it in the Authorization header on every request. Priya adds the Kitebase server to Claude Code with her key:
claude mcp add --transport http kitebase https://mcp.kitebase.example/mcp \
--header "Authorization: Bearer kb_live_Xq3..."
On your side, a key is issued once and never stored in plain text:
def issue_key(self, key_id: str, customer_id: str) -> str:
"""Create a key and return the raw value. Show it to the customer once; it isn't stored."""
raw = "kb_live_" + secrets.token_urlsafe(24)
self._keys_by_hash[hash_key(raw)] = ApiKey(key_id, customer_id, hash_key(raw), raw[:12])
return raw
secrets.token_urlsafe(24) gives 24 random bytes as text, far too many to guess. You keep only the key’s SHA-256 hash (a fixed-length fingerprint that can’t be turned back into the key), so a leaked database doesn’t leak working keys. Checking a key means hashing what arrived and looking that up. The kb_live_ prefix makes a key easy to spot if someone pastes it into a public repo.
The HTTP middleware does that lookup on every request:
auth = dict(scope["headers"]).get(b"authorization", b"").decode()
key = self.accounts.lookup(auth.removeprefix("Bearer ").strip())
if key is None:
response = JSONResponse(
{"error": "Missing or unknown API key. Create one at https://kitebase.example/settings/api-keys"},
status_code=401,
headers={"WWW-Authenticate": 'Bearer realm="kitebase"'},
)
return await response(scope, receive, send)
scope.setdefault("state", {})["api_key"] = key # the MCP layer reads it from here
A request with no key gets this in the demo:
1. A request with no API key
HTTP 401 Missing or unknown API key. Create one at https://kitebase.example/settings/api-keys
401 is the HTTP status for “I don’t know who you are”. Without a known caller nothing should work, so the request never reaches MCP.
Issue one key per person or per agent, never one per company. Northwind has three: nw-priya, nw-sam and nw-lee. When Lee leaves, you revoke one key, and the bill can show who used what.
The gotcha is leaking keys. Never put one in a URL, since URLs end up in logs; the MCP spec forbids access tokens in the query string for the same reason. Never log the header, and never echo a key back in an error.
Shouldn't a paid MCP server use OAuth instead of API keys?
Eventually, maybe. The MCP spec makes authorization optional, and says HTTP servers that support it should follow its OAuth 2.1-based flow. With OAuth, the user signs in through a browser and the host gets a token, so there’s no key to copy and paste. The Python SDK supports it through MCPServer(token_verifier=..., auth=AuthSettings(...)).
It’s also a lot more to run: an authorization server, discovery metadata, token audiences. If your buyers are developers who are used to API keys, start with keys. Move to OAuth when people who aren’t developers need to connect from a host.
Step 2: Meter every tool call
Metering means recording each billable unit as it happens, so you can add them up at the end of the month. The unit here is one tool call. It happens in the MCP layer, in meter_tool_calls, which the SDK runs for every inbound message:
async def meter_tool_calls(self, ctx: ServerRequestContext, call_next: CallNext):
if ctx.method != "tools/call":
return await call_next(ctx) # tools/list and friends are free
key = getattr(ctx.request.state, "api_key", None) if ctx.request else None
if key is None:
return refuse("This server needs an API key.")
customer = self.accounts.customers[key.customer_id]
tool = (ctx.params or {}).get("name", "?")
now = self.clock()
... # rate limit and quota: step 3
result = await call_next(ctx) # runs the actual tool
log("tool_error" if is_error(result) else "ok")
return result
ctx.method and ctx.params are the raw JSON-RPC method and params, and ctx.request is the HTTP request the key check tagged. call_next(ctx) hands the message on to the SDK, which validates the arguments and runs the tool. Code before it can refuse the call; code after it sees the result.
Each call becomes one row in the usage log:
UsageEvent(key_id="nw-priya", customer_id="northwind", tool="get_ticket",
at=datetime(2026, 9, 27, 14, 0, tzinfo=timezone.utc), outcome="ok")
Bill only calls that worked. In the demo, Priya asks for a ticket that doesn’t exist:
3. Priya asks for a ticket that doesn't exist
is_error: true Error executing tool get_ticket: No ticket 'KITE-999'. Ticket ids look like KITE-142.
That’s logged as tool_error and shows up under “Not billed”. The model chose the call, not the customer, and nobody likes paying for errors.
One gotcha bit me while building this. The middleware sees the result as it goes on the wire, a dict with camelCase keys, not the CallToolResult object. My first version read is_error off it, got False every time, and billed every failed call. The is_error() helper in gateway.py reads result["isError"]. Test your meter with a failing call.
The SDK marks this middleware hook as provisional, so its signature may change in a 2.x minor release. The example pins mcp>=2.2,<2.3; check the changelog before you upgrade.
Step 3: Rate limits and quotas per plan
Plans put two different limits on usage, and they do different jobs:
- A rate limit caps how fast one key can call: here, 60 calls in any 60 seconds. It protects your server from bursts, and the customer’s bill from a runaway agent.
- A quota caps how much a customer can use in a billing period: here, the Free plan’s 100 calls a month. It’s what the plan includes.
The rate limiter is a sliding window: keep the times of each key’s recent calls, drop the ones older than 60 seconds, and refuse if what’s left is already at the limit.
def check(self, key_id: str, limit: int, now: float) -> float | None:
"""Record a call and return None, or return how many seconds to wait."""
calls = self._calls[key_id]
while calls and now - calls[0] >= self.window:
calls.popleft()
if len(calls) >= limit:
return self.window - (now - calls[0])
calls.append(now)
return None
calls is a deque, a list that’s cheap to pop from the front. A simpler fixed window (count per clock minute, reset at :00) lets a key make 60 calls at 14:00:59 and 60 more at 14:01:00; the sliding window never allows more than 60 in any 60 seconds.
Hollow Pine’s agent gets stuck in a loop and searches 65 times in the same minute:
5. Hollow Pine's agent gets stuck in a loop: 65 searches in the same minute
60 ran, 5 refused. The first refusal:
is_error: true Rate limit: the Pay as you go plan allows 60 calls a minute per key. Try again in 60 seconds.
Blue Fern Labs, on Free with its 100 calls used, gets the quota message instead:
4. Blue Fern Labs (Free plan, 100 of 100 calls used) searches the help center
is_error: true Blue Fern Labs has used all 100 calls on the Free plan this month. Upgrade at https://kitebase.example/billing, or wait until the 1st.
Both refusals are tool results with is_error: true, built by refuse() in gateway.py, not HTTP errors. The model reads them like any tool result (article 03 covers isError) and can tell the user what to do. An HTTP 429 (“too many requests”), the usual answer in a REST API, would fail the request at the transport level instead, and what the model sees then is up to the host. Say which limit it is and when it resets, so the model doesn’t retry straight away.
The gotcha: a rate limit caps speed, not spending. A loop running at 60 calls a minute for an 8-hour night is 28,800 calls, or $115.20 on Pay as you go, and every one is allowed. Paid plans need a monthly spend cap, or at least an email when usage jumps.
Step 4: The monthly bill
At the end of the month, add up each customer’s billable events and apply their plan. For Northwind on Team:
- 3 keys × $15 = $45.00 in seat fees.
- 3 keys × 2,000 = 6,000 calls included. Northwind made 6,413, so 413 are extra.
- 413 × $4 / 1,000 = $1.652, rounded to $1.65. Total: $46.65.
monthly_bill does that for every customer, and the demo prints it:
Bill for September 2026 (worked-example prices)
Customer Plan Keys Billed Not billed Amount
Northwind Studio Team 3 6,413 1 $45.00 + $1.65 = $46.65
Blue Fern Labs Free 1 100 1 $0.00 + $0.00 = $0.00
Hollow Pine Pay as you go 1 12,440 5 $0.00 + $49.76 = $49.76
“Not billed” is the KITE-999 error, Blue Fern’s refused call and Hollow Pine’s five rate-limited calls. Money is computed with Python’s Decimal and rounded to cents once per line, because floats round in binary: 0.1 + 0.2 isn’t exactly 0.3.
Don’t build the rest of billing yourself: payments, invoices, failed-payment retries, refunds and tax are a product of their own. Send usage to a billing provider and let it invoice. With Stripe, you report usage as meter events, and each event takes an identifier so a resent event isn’t counted twice. That’s one reason a real meter keeps a row per call, each with its own id, rather than a running total.
The provider also decides who deals with tax. With a payment processor like Stripe, you are the seller, so working out, collecting and filing sales tax or VAT is your job (Stripe Tax can do the calculating and collecting). A merchant of record such as Lemon Squeezy or Paddle is the legal seller instead: it sells on your behalf, handles sales tax and payment disputes, and pays you out. Tax rules depend on where you and your customers are, so read the provider’s docs and talk to an accountant before you sell to anyone.
The code is the small part. Paying customers also expect terms of service and a privacy policy (tickets carry customer names, and your server sees every one; this usage log stores tool names and times, not arguments, on purpose), someone to email when it breaks, and a page showing this month’s usage per key, so the invoice is never a surprise.
Try it yourself
The companion example is the whole gateway: server.py from article 01 unchanged, billing.py with the plans, keys, limiter, meter and bill, and gateway.py with the two middlewares. main.py runs the five calls above and prints the bill, offline.
Download the runnable example (zip)
cd 09-selling-your-mcp-server
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
- In
gateway.py, change the lastlog(...)line to always log"ok". Run it again: Northwind now makes 6,414 billed calls, because the KITE-999 error is billed. - Move Hollow Pine to the Team plan in
data/usage-2026-09.json. Its bill becomes $15.00 + $41.76 = $56.76, as in the grid above. - Run
python main.py --serve. It serves onhttp://localhost:8000/mcpand prints a fresh key fornw-priya. Add it to Claude Code withclaude mcp add --transport http, as in step 1, and ask about KITE-142.
pip install pytest && pytest -q runs the offline tests. They go through the real server over HTTP, in-process: a 401 for a wrong key, a failed call not billed, a runaway agent refused and then allowed again 61 seconds later.
Common beginner mistakes
- One key per company. You can’t revoke one person, and you can’t see who ran up the bill. Issue a key per person or per agent.
- Storing keys in plain text. A leaked database becomes a pile of working keys. Store hashes, and show the raw key once.
- Billing failed calls. The model chose the call, and the tool failed. Log it, don’t charge for it.
- Counting usage in memory in production. A restart loses the month’s usage, which is money. Write each event to a database as it happens.
Questions you will face in production
“Where does the usage log live?”
In a database table, written where log() is called now: one row per call with an id, the key, the tool, the time and the outcome. Send usage to your billing provider in batches carrying those ids, so a retried batch doesn’t bill twice.
“How do I stop a customer’s agent from running up a huge bill?” Three layers: the rate limit per key, a monthly spend cap the customer sets (refused like the quota once it’s hit), and an email at 50% and 90% of the cap so it never surprises them.
Check your understanding
A customer says their bill counts calls that returned errors. Where do you look first?
At the outcome recorded for those calls in the usage log. If failed calls are logged as ok, the meter’s error check is wrong. In this SDK, the middleware sees the wire-format dict, so it has to read result["isError"], not result.is_error. A test with a failing call catches it.
Why is a missing API key an HTTP 401, but a rate-limited call a tool result with is_error: true?
With no key, you don’t know who’s calling, so nothing about the connection should work, including tools/list. It’s refused before MCP. A rate-limited call comes from a known customer on a working connection. Returning it as a tool result lets the model read “Try again in 60 seconds” and tell the user.
Northwind adds a fourth key in October and makes 7,500 calls. What's the bill on Team?
4 × $15 = $60.00 in seat fees, and 4 × 2,000 = 8,000 calls included. 7,500 is under that, so there’s no overage: $60.00. The new key raised both the seat fees and the allowance.
You run two copies of the server behind a load balancer. What happens to the 60-calls-a-minute limit?
Each copy keeps its own RateLimiter in memory, so a key can make about 60 calls a minute on each: 120 in total. Move the limiter’s state to a shared store like Redis so every copy counts the same calls.
What to remember
- You’re selling access to what’s behind the server, not the server. Charging needs a remote HTTP server that knows who’s calling.
- Per seat is predictable, usage-based tracks your costs, a free tier gets people to try it. Included usage plus overage mixes them. Work out prices from your cost per call.
- Check the API key at the HTTP layer; limit and meter at the MCP layer, on
tools/callonly. Store key hashes, and issue a key per person or agent. - Rate limits cap speed; quotas cap the month. Refuse with a tool result the model can relay, and add a spend cap, because a rate limit alone doesn’t stop a slow leak.
- Bill only calls that worked, keep one usage row per call, and let a billing provider handle invoices and tax.
What to study next
That’s the end of the MCP track: you can build a server, connect it, ship it and charge for it. The callers on the other end are often agents, programs that decide which tool to call and when, and how they behave is what your limits have to handle. What an AI Agent Actually Is starts that topic.
Further reading
- MCP spec: Authorization. When authorization applies, the OAuth 2.1-based flow for HTTP servers, and the rules on tokens.
- MCP Python SDK.
MCPServer, its middleware hook andstreamable_http_app(). - Stripe: Usage-based billing. Stripe’s options for metered pricing.
- Stripe: Record usage with the API. Meter events, their
identifierfor deduplication, and how late you can report usage. - Lemon Squeezy: Merchant of record. What a merchant of record takes off your hands.
Where this article comes from. This is a synthesis of the MCP specification and common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.