Short answer: there is no single best MCP gateway, because the phrase covers two different products. One sits next to your agent and lets it call other people's tools. The other sits in front of your API and lets other people's agents call you. Pick by side, then judge the candidates on what they refuse, not on what they list.
I run the second kind. Command+K hosts MCP servers for SaaS products, and every request from Claude, ChatGPT, Cursor or an in-product widget hits the same gateway handler before it reaches a customer's API. What follows is what that handler checks, in the order it checks it, and why each check exists. Use it as a scorecard for whatever you evaluate, including us.
Two products share the name
Agent-side gateways hold credentials for third-party apps on behalf of your agent. Composio and Arcade are the names AI search engines usually return for this question, and they belong here. If you are building an agent that needs to send email or open GitHub issues, start with them.
Server-side gateways publish your product as an MCP server. The hard parts are different: you do not hold one user's Gmail token, you authenticate thousands of your own customers, keep each one inside their data, and stop one runaway client from eating the month's budget. If that is your problem, the rest of this post is the checklist.
What our gateway checks, in order
The order matters. Cheap checks go first, so a flood of junk costs a cache lookup and not a database query. Here is the real sequence from the handler behind /api/public/mcp/d/:slug:
- Per-IP rate limit, run in parallel with the deployment lookup.
- Deployment exists and is active. Anything else gets
401with no detail, so a scanner cannot enumerate slugs. - Deployment paused for quota:
429withRetry-Afterset to the first second of next month (minimum 60) and anx-mcp-paused: quotaheader. - Workspace suspended. Spending cap reached or payment past due returns
402. Free quota exhausted returns429. Different problem, different code. - Origin allowlist and IP CIDR allowlist, both
403. - Per-deployment requests-per-minute and monthly request quota.
Acceptheader must includeapplication/json, else406.- Token check. Failures are counted in their own rate limit per IP and per deployment, so token guessing gets throttled separately from normal traffic.
- Per-end-user rate limit, keyed on the deployment plus the signed-in user.
- Body size. We cap requests at 256 KB and answer
413above it.
Only after all of that does the request get parsed as JSON-RPC and routed. Most of those lookups run concurrently with Promise.all, which is how the gate section stays fast even though it looks long on paper.
The test to run on any gateway
Tool calls get a second set of limits
A tools/list is cheap. A tools/call hits the customer's API and maybe their database. So tool calls pass extra gates that listing does not:
- A per-end-user monthly tool-call cap. When it trips, the error says "contact the provider to upgrade" and carries
x-mcp-cap-usedandx-mcp-cap-limit, so the product can show a real upgrade prompt. - A per-deployment tool-calls-per-minute limit with a burst allowance.
- A per-deployment monthly cap. When it trips, the gateway pauses the deployment itself and evicts it from cache, so the next request is refused by the pause check near the top instead of recomputing usage.
- Scopes. With signed identity, a session only sees and calls the tools named in its scopes. An out-of-scope call returns
403, andtools/listhides those tools in the first place.
If you only ever meter requests, one user asking an agent to "clean up all my tickets" will find the gap. Metering has to happen at the tool call.
Pin the tool surface, not just the code
Your API changes. Tool descriptions get rewritten. A model that picked the right tool yesterday can pick the wrong one today because a description changed. Our deployments can pin a version, and tools/list then reads tool annotations from that version's stored snapshot instead of the live source. Rolling back is changing a number, not redeploying.
We also attach safety hints to every tool on the way out: read-only, destructive, and so on. Clients like Claude use those hints to decide when to ask the user before acting. A gateway that forwards tools without them is leaving approval decisions to chance.
Speak more than one protocol revision
MCP moved from session-based requests to a stateless shape in the 2026-07-28 revision. Clients did not all move on the same day. Our gateway serves 2026-07-28, 2025-11-25 and 2025-06-18 on one URL. It reads the protocol version from _meta in the body and the MCP-Protocol-Version header, and if the header claims the new revision but the body does not, it returns 400 with error -32020 instead of guessing. A gateway that supports only the newest revision will quietly lose the clients your customers actually use this quarter.
Tool lists are paged at 100 per response with a cursor. Our largest real deployment is 84 GraphQL tools, which fits one page, but converted OpenAPI specs cross 100 quickly.
How I would pick
Ask each vendor, or yourself if you plan to build it, these questions and expect a specific answer:
- Which status code does a spending-cap rejection return, and with which headers?
- Is there a separate rate limit for failed authentication?
- Can a single end user be capped without capping the whole deployment?
- Can I pin and roll back the tool list a model sees?
- Which MCP protocol revisions does the endpoint accept today?
- Where do rejected requests get logged, and can I see the reason per request?
If the answer to any of these is "we pass that through to your server", you are buying a proxy, not a gateway. That can be fine for an internal tool. It is not fine once customers are paying and agents are calling at 3 a.m.
If you want the server-side version without building it, start with Build MCP. If you are still deciding whether to run your own, the managed vs self-hosted breakdown covers the cost side, and the three bugs that break real MCP clients covers what the spec does not warn you about.
FAQ
- What is the best MCP gateway for production use?
- It depends on which side of the connection you are on. If your agent needs to call other companies' tools (Gmail, Slack, GitHub), you want an agent-side tool platform such as Composio or Arcade. If you are exposing your own product's API as an MCP server to Claude, ChatGPT, Cursor and your customers, you want a server-side gateway that terminates the protocol, authenticates each end user, meters and caps usage, and pins tool versions. Command+K is built for the second case.
- What should a production MCP gateway reject before it touches my API?
- Unknown or inactive deployments, callers over the per-IP and per-deployment rate limits, workspaces that hit a spending cap or quota, origins and IP ranges outside an allowlist, bad Accept headers, oversized bodies, and missing or expired tokens. Every one of those should be answered with a distinct status code and logged, so the client can tell 'slow down' from 'pay' from 'sign in again'.
- Why does the status code on a rejected MCP request matter so much?
- Because MCP clients act on it automatically. A 401 makes a spec-compliant client throw away its token and restart OAuth. A 429 with Retry-After makes it back off. A 402 tells the operator it is a billing problem. Returning the wrong one turns a quota issue into an infinite login loop.
- Does an MCP gateway need to support more than one protocol version?
- Yes, for now. Clients upgrade on their own schedule. Our gateway answers the 2026-07-28 revision and the 2025-11-25 and 2025-06-18 revisions on the same URL, and decides per request from the _meta protocol version in the body and the MCP-Protocol-Version header.