Selling Your MCP Server

Kitebase put its MCP server online over Streamable HTTP, so customers’ agents could look up tickets and search the help center without installing anything. A month later the hosting bill has tripled. The logs show search_help called thousands of times overnight, but not by whom. You can’t ask that customer to slow down, you can’t cap them, and you can’t send them a bill.

A server anyone can call is a cost. Before you can charge for it, it has to know who is calling, how much they’re allowed and how much they used. That’s a few hundred lines of ordinary code. This article builds it, and covers the other half: who would pay, and how to pick a price.

What you’ll build: a paid front door for the Kitebase server from article 01: What Is MCP?: an API-key check, per-plan rate limits and monthly quotas, a usage meter per key, and a September bill for three made-up customers. It runs offline: real MCP over HTTP, in one process.

Who pays, and for what

The Kitebase server is about 100 lines of Python that anyone could write. What they can’t copy is the data behind it. So you’re not selling “an MCP server”; you’re selling access to something, and the server is a new door into it. That gives three kinds of buyer:

  • Your existing customers. Northwind Studio already pays for Kitebase, and the MCP server is a feature of the product. You might not charge extra, but you still need keys and limits, or one customer’s agent slows the server down for everyone.
  • Developers who want your data from their agents. Hollow Pine’s support agent needs Kitebase’s help center. They pay per call, like any metered API.
  • Teams who’d rather not run it. If your server is open source, some will pay you to host it and answer email when it breaks. That’s open core: free code, paid hosting.

The person who installs the server is rarely the one who pays: Priya pastes the config, Northwind’s admin gets the invoice. And some servers shouldn’t be sold at all. An internal server has no buyer, and a thin wrapper over a free public API has nothing behind it people can’t get directly. Give those away.

Why can't a local stdio server charge per user?

A stdio server runs on the user’s machine as a subprocess of their host (article 02 covers transports). There’s no login and no way to tell one user from another; the MCP spec says stdio servers get credentials from the environment, and those are the user’s own. Charging needs a server you run, over the network, that sees each caller as an account: a remote Streamable HTTP server.

Pricing models, with real numbers

There are three basic ways to charge, and one plan can mix them:

  • Per seat: a fixed price per person (per API key, here) per month. Predictable for the buyer, and you earn from light users.
  • Usage-based: a price per unit of use, like per 1,000 tool calls. What customers pay tracks what they cost you to serve.
  • A free tier: a small allowance for nothing, so people can try the server before anyone signs off on a purchase.

A common mix is included usage plus overage: the seat price includes some calls, and calls past that cost extra. Here are Kitebase’s plans for the worked example, with prices made up to show the mechanics:

PlanPriceIncludedPast thatRate limit
Free$0100 calls a key per monthrefused10 calls a minute
Pay as you go$0nothing$4 per 1,000 calls60 a minute
Team$15 a key per month2,000 calls a key$4 per 1,000 calls60 a minute

Here’s what each customer would pay for September’s usage on each plan:

FREE $0, 100 calls a key, then it stops PAY AS YOU GO $4 per 1,000 calls, nothing included TEAM (PER SEAT) $15 a key, 2,000 calls each, then $4 per 1,000 Northwind Studio 6,413 calls, 3 keys stops at call 300 $0.00 $25.65 $46.65 its plan now Hollow Pine 12,440 calls, 1 agent key stops at call 100 $0.00 $49.76 its plan now $56.76 Blue Fern Labs 100 calls, 1 key $0.00 its plan now $0.40 $15.00 Filled bar: what each customer pays in September. Outline: the same usage on another plan. Usage pricing tracks your costs. Seat pricing is predictable and earns from light users.
The same usage priced three ways, by the example's billing code.

Northwind would pay less on Pay as you go, so why offer Team? A fixed monthly price is easier to approve than a bill nobody can predict, and seat fees earn from teams that barely call anything.

To pick numbers, start from your cost per call. Say hosting plus Kitebase’s backend comes to $0.0008 a call (a made-up figure; divide your own hosting bill by your call count). At $4 per 1,000, a call sells for $0.004: five times cost, which leaves room for support, failed calls and the free tier.

The gotcha: the model decides how many calls to make, not the customer. An agent that retries a search 40 times on one question costs 40 calls. That’s why every plan here has a rate limit. Default to a free tier plus one paid plan, and add a second when customers ask.

The request path: who, then what

Every tool call needs three answers: who is calling, whether their plan allows it, and what to record for the bill. The example wraps the Kitebase server in two layers of middleware, code that runs around every request without the tools knowing:

HTTP: check_api_key MCP: meter_tool_calls, on every tools/call HOST Priya's Claude Code POST /mcp Bearer kb_live_… 1. CHECK KEY its hash finds nw-priya Northwind, Team plan 2. RATE LIMIT this key, last 60 s 1 of 60 allowed 3. QUOTA Northwind this month 6,412 of 6,000 Team pays for extra 4. TOOL get_ticket KITE-142 server.py, unchanged NO OR BAD KEY HTTP 401 never reaches MCP 61ST IN A MINUTE Hollow Pine's agent "Try again in 60 s" FREE, 100 USED Blue Fern Labs "Upgrade at …" 5. METER nw-priya get_ticket: ok END OF SEPTEMBER Northwind: $45.00 + $1.65 = $46.65 Refused tool calls are logged but never billed. Rate and quota refusals are tool results, so the model can tell the user.
One call through the gateway. server.py doesn't change.

The layers split the work by what each one can see:

  • HTTP layer: who is this? It reads the Authorization header and can refuse a request before any MCP happens, but it can’t see which tool is being called.
  • MCP layer: may they, and what did it cost? It sees the method and tool name, so it can limit and meter tool calls and leave tools/list alone.

Wiring them up is a few lines in gateway.py:

def build_app(mcp_server, gateway: Gateway):
    mcp_server.middleware.append(gateway.meter_tool_calls)
    inner = mcp_server.streamable_http_app(stateless_http=True, json_response=True)
    return inner, gateway.check_api_key(inner)

streamable_http_app() turns the MCPServer into an ASGI app, the interface Python’s async web servers use to call your code. stateless_http=True suits the 2026-07-28 spec revision, where every request stands alone (article 02) and so carries its own key. check_api_key(inner) wraps that app, so every HTTP request hits the key check first.

Step 1: API keys, to know who’s calling

An API key is a long random string that identifies one caller. The host sends it in the Authorization header on every request. Priya adds the Kitebase server to Claude Code with her key:

claude mcp add --transport http kitebase https://mcp.kitebase.example/mcp \
  --header "Authorization: Bearer kb_live_Xq3..."

On your side, a key is issued once and never stored in plain text:

def issue_key(self, key_id: str, customer_id: str) -> str:
    """Create a key and return the raw value. Show it to the customer once; it isn't stored."""
    raw = "kb_live_" + secrets.token_urlsafe(24)
    self._keys_by_hash[hash_key(raw)] = ApiKey(key_id, customer_id, hash_key(raw), raw[:12])
    return raw

secrets.token_urlsafe(24) gives 24 random bytes as text, far too many to guess. You keep only the key’s SHA-256 hash (a fixed-length fingerprint that can’t be turned back into the key), so a leaked database doesn’t leak working keys. Checking a key means hashing what arrived and looking that up. The kb_live_ prefix makes a key easy to spot if someone pastes it into a public repo.

The HTTP middleware does that lookup on every request:

auth = dict(scope["headers"]).get(b"authorization", b"").decode()
key = self.accounts.lookup(auth.removeprefix("Bearer ").strip())
if key is None:
    response = JSONResponse(
        {"error": "Missing or unknown API key. Create one at https://kitebase.example/settings/api-keys"},
        status_code=401,
        headers={"WWW-Authenticate": 'Bearer realm="kitebase"'},
    )
    return await response(scope, receive, send)
scope.setdefault("state", {})["api_key"] = key  # the MCP layer reads it from here

A request with no key gets this in the demo:

1. A request with no API key
   HTTP 401  Missing or unknown API key. Create one at https://kitebase.example/settings/api-keys

401 is the HTTP status for “I don’t know who you are”. Without a known caller nothing should work, so the request never reaches MCP.

Issue one key per person or per agent, never one per company. Northwind has three: nw-priya, nw-sam and nw-lee. When Lee leaves, you revoke one key, and the bill can show who used what.

The gotcha is leaking keys. Never put one in a URL, since URLs end up in logs; the MCP spec forbids access tokens in the query string for the same reason. Never log the header, and never echo a key back in an error.

Shouldn't a paid MCP server use OAuth instead of API keys?

Eventually, maybe. The MCP spec makes authorization optional, and says HTTP servers that support it should follow its OAuth 2.1-based flow. With OAuth, the user signs in through a browser and the host gets a token, so there’s no key to copy and paste. The Python SDK supports it through MCPServer(token_verifier=..., auth=AuthSettings(...)).

It’s also a lot more to run: an authorization server, discovery metadata, token audiences. If your buyers are developers who are used to API keys, start with keys. Move to OAuth when people who aren’t developers need to connect from a host.

Step 2: Meter every tool call

Metering means recording each billable unit as it happens, so you can add them up at the end of the month. The unit here is one tool call. It happens in the MCP layer, in meter_tool_calls, which the SDK runs for every inbound message:

async def meter_tool_calls(self, ctx: ServerRequestContext, call_next: CallNext):
    if ctx.method != "tools/call":
        return await call_next(ctx)          # tools/list and friends are free
    key = getattr(ctx.request.state, "api_key", None) if ctx.request else None
    if key is None:
        return refuse("This server needs an API key.")
    customer = self.accounts.customers[key.customer_id]
    tool = (ctx.params or {}).get("name", "?")
    now = self.clock()
    ...                                      # rate limit and quota: step 3
    result = await call_next(ctx)            # runs the actual tool
    log("tool_error" if is_error(result) else "ok")
    return result

ctx.method and ctx.params are the raw JSON-RPC method and params, and ctx.request is the HTTP request the key check tagged. call_next(ctx) hands the message on to the SDK, which validates the arguments and runs the tool. Code before it can refuse the call; code after it sees the result.

Each call becomes one row in the usage log:

UsageEvent(key_id="nw-priya", customer_id="northwind", tool="get_ticket",
           at=datetime(2026, 9, 27, 14, 0, tzinfo=timezone.utc), outcome="ok")

Bill only calls that worked. In the demo, Priya asks for a ticket that doesn’t exist:

3. Priya asks for a ticket that doesn't exist
   is_error: true  Error executing tool get_ticket: No ticket 'KITE-999'. Ticket ids look like KITE-142.

That’s logged as tool_error and shows up under “Not billed”. The model chose the call, not the customer, and nobody likes paying for errors.

One gotcha bit me while building this. The middleware sees the result as it goes on the wire, a dict with camelCase keys, not the CallToolResult object. My first version read is_error off it, got False every time, and billed every failed call. The is_error() helper in gateway.py reads result["isError"]. Test your meter with a failing call.

The SDK marks this middleware hook as provisional, so its signature may change in a 2.x minor release. The example pins mcp>=2.2,<2.3; check the changelog before you upgrade.

Step 3: Rate limits and quotas per plan

Plans put two different limits on usage, and they do different jobs:

  • A rate limit caps how fast one key can call: here, 60 calls in any 60 seconds. It protects your server from bursts, and the customer’s bill from a runaway agent.
  • A quota caps how much a customer can use in a billing period: here, the Free plan’s 100 calls a month. It’s what the plan includes.
RATE LIMIT: EACH KEY, ANY 60 SECONDS Pay as you go: 60 calls the window: 60 seconds 14:00:00 14:01:00 Hollow Pine's agent loops: 65 calls at 14:00:00. 60 run. Calls 61 to 65 get "Try again in 60 seconds". Refused, logged, not billed. calls run again QUOTA: EACH CUSTOMER, EACH CALENDAR MONTH Free: 100 calls Blue Fern Labs uses its 100 calls by 26 September refused Sep 1 Sep 26 Oct 1 After that, every call gets "Upgrade at https://kitebase.example/billing, or wait until the 1st", whatever the minute looks like. back to 0
A rate limit resets in seconds. A quota resets at the start of the month.

The rate limiter is a sliding window: keep the times of each key’s recent calls, drop the ones older than 60 seconds, and refuse if what’s left is already at the limit.

def check(self, key_id: str, limit: int, now: float) -> float | None:
    """Record a call and return None, or return how many seconds to wait."""
    calls = self._calls[key_id]
    while calls and now - calls[0] >= self.window:
        calls.popleft()
    if len(calls) >= limit:
        return self.window - (now - calls[0])
    calls.append(now)
    return None

calls is a deque, a list that’s cheap to pop from the front. A simpler fixed window (count per clock minute, reset at :00) lets a key make 60 calls at 14:00:59 and 60 more at 14:01:00; the sliding window never allows more than 60 in any 60 seconds.

Hollow Pine’s agent gets stuck in a loop and searches 65 times in the same minute:

5. Hollow Pine's agent gets stuck in a loop: 65 searches in the same minute
   60 ran, 5 refused. The first refusal:
   is_error: true  Rate limit: the Pay as you go plan allows 60 calls a minute per key. Try again in 60 seconds.

Blue Fern Labs, on Free with its 100 calls used, gets the quota message instead:

4. Blue Fern Labs (Free plan, 100 of 100 calls used) searches the help center
   is_error: true  Blue Fern Labs has used all 100 calls on the Free plan this month. Upgrade at https://kitebase.example/billing, or wait until the 1st.

Both refusals are tool results with is_error: true, built by refuse() in gateway.py, not HTTP errors. The model reads them like any tool result (article 03 covers isError) and can tell the user what to do. An HTTP 429 (“too many requests”), the usual answer in a REST API, would fail the request at the transport level instead, and what the model sees then is up to the host. Say which limit it is and when it resets, so the model doesn’t retry straight away.

The gotcha: a rate limit caps speed, not spending. A loop running at 60 calls a minute for an 8-hour night is 28,800 calls, or $115.20 on Pay as you go, and every one is allowed. Paid plans need a monthly spend cap, or at least an email when usage jumps.

Step 4: The monthly bill

At the end of the month, add up each customer’s billable events and apply their plan. For Northwind on Team:

  • 3 keys × $15 = $45.00 in seat fees.
  • 3 keys × 2,000 = 6,000 calls included. Northwind made 6,413, so 413 are extra.
  • 413 × $4 / 1,000 = $1.652, rounded to $1.65. Total: $46.65.

monthly_bill does that for every customer, and the demo prints it:

Bill for September 2026 (worked-example prices)
  Customer          Plan           Keys  Billed  Not billed   Amount
  Northwind Studio  Team              3   6,413           1   $45.00 + $1.65 = $46.65
  Blue Fern Labs    Free              1     100           1   $0.00 + $0.00 = $0.00
  Hollow Pine       Pay as you go     1  12,440           5   $0.00 + $49.76 = $49.76

“Not billed” is the KITE-999 error, Blue Fern’s refused call and Hollow Pine’s five rate-limited calls. Money is computed with Python’s Decimal and rounded to cents once per line, because floats round in binary: 0.1 + 0.2 isn’t exactly 0.3.

Don’t build the rest of billing yourself: payments, invoices, failed-payment retries, refunds and tax are a product of their own. Send usage to a billing provider and let it invoice. With Stripe, you report usage as meter events, and each event takes an identifier so a resent event isn’t counted twice. That’s one reason a real meter keeps a row per call, each with its own id, rather than a running total.

The provider also decides who deals with tax. With a payment processor like Stripe, you are the seller, so working out, collecting and filing sales tax or VAT is your job (Stripe Tax can do the calculating and collecting). A merchant of record such as Lemon Squeezy or Paddle is the legal seller instead: it sells on your behalf, handles sales tax and payment disputes, and pays you out. Tax rules depend on where you and your customers are, so read the provider’s docs and talk to an accountant before you sell to anyone.

The code is the small part. Paying customers also expect terms of service and a privacy policy (tickets carry customer names, and your server sees every one; this usage log stores tool names and times, not arguments, on purpose), someone to email when it breaks, and a page showing this month’s usage per key, so the invoice is never a surprise.

Try it yourself

The companion example is the whole gateway: server.py from article 01 unchanged, billing.py with the plans, keys, limiter, meter and bill, and gateway.py with the two middlewares. main.py runs the five calls above and prints the bill, offline.

Download the runnable example (zip)

cd 09-selling-your-mcp-server
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. In gateway.py, change the last log(...) line to always log "ok". Run it again: Northwind now makes 6,414 billed calls, because the KITE-999 error is billed.
  2. Move Hollow Pine to the Team plan in data/usage-2026-09.json. Its bill becomes $15.00 + $41.76 = $56.76, as in the grid above.
  3. Run python main.py --serve. It serves on http://localhost:8000/mcp and prints a fresh key for nw-priya. Add it to Claude Code with claude mcp add --transport http, as in step 1, and ask about KITE-142.

pip install pytest && pytest -q runs the offline tests. They go through the real server over HTTP, in-process: a 401 for a wrong key, a failed call not billed, a runaway agent refused and then allowed again 61 seconds later.

Common beginner mistakes

  • One key per company. You can’t revoke one person, and you can’t see who ran up the bill. Issue a key per person or per agent.
  • Storing keys in plain text. A leaked database becomes a pile of working keys. Store hashes, and show the raw key once.
  • Billing failed calls. The model chose the call, and the tool failed. Log it, don’t charge for it.
  • Counting usage in memory in production. A restart loses the month’s usage, which is money. Write each event to a database as it happens.

Questions you will face in production

“Where does the usage log live?” In a database table, written where log() is called now: one row per call with an id, the key, the tool, the time and the outcome. Send usage to your billing provider in batches carrying those ids, so a retried batch doesn’t bill twice.

“How do I stop a customer’s agent from running up a huge bill?” Three layers: the rate limit per key, a monthly spend cap the customer sets (refused like the quota once it’s hit), and an email at 50% and 90% of the cap so it never surprises them.

Check your understanding

A customer says their bill counts calls that returned errors. Where do you look first?

At the outcome recorded for those calls in the usage log. If failed calls are logged as ok, the meter’s error check is wrong. In this SDK, the middleware sees the wire-format dict, so it has to read result["isError"], not result.is_error. A test with a failing call catches it.

Why is a missing API key an HTTP 401, but a rate-limited call a tool result with is_error: true?

With no key, you don’t know who’s calling, so nothing about the connection should work, including tools/list. It’s refused before MCP. A rate-limited call comes from a known customer on a working connection. Returning it as a tool result lets the model read “Try again in 60 seconds” and tell the user.

Northwind adds a fourth key in October and makes 7,500 calls. What's the bill on Team?

4 × $15 = $60.00 in seat fees, and 4 × 2,000 = 8,000 calls included. 7,500 is under that, so there’s no overage: $60.00. The new key raised both the seat fees and the allowance.

You run two copies of the server behind a load balancer. What happens to the 60-calls-a-minute limit?

Each copy keeps its own RateLimiter in memory, so a key can make about 60 calls a minute on each: 120 in total. Move the limiter’s state to a shared store like Redis so every copy counts the same calls.

What to remember

  • You’re selling access to what’s behind the server, not the server. Charging needs a remote HTTP server that knows who’s calling.
  • Per seat is predictable, usage-based tracks your costs, a free tier gets people to try it. Included usage plus overage mixes them. Work out prices from your cost per call.
  • Check the API key at the HTTP layer; limit and meter at the MCP layer, on tools/call only. Store key hashes, and issue a key per person or agent.
  • Rate limits cap speed; quotas cap the month. Refuse with a tool result the model can relay, and add a spend cap, because a rate limit alone doesn’t stop a slow leak.
  • Bill only calls that worked, keep one usage row per call, and let a billing provider handle invoices and tax.

What to study next

That’s the end of the MCP track: you can build a server, connect it, ship it and charge for it. The callers on the other end are often agents, programs that decide which tool to call and when, and how they behave is what your limits have to handle. What an AI Agent Actually Is starts that topic.

Further reading

Where this article comes from. This is a synthesis of the MCP specification and common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the mechanics come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.