Tutorial
How to deploy a remote MCP server to production: 2026 stateless MCP + HTTP load balancer + JSON-RPC
Protocol statelessness is worth money only when you actually use a normal HTTP load balancer. Production remote MCP is about replica scale, proxies that must not buffer SSE, and drains that do not kill long streams—not inventing another session.
After 2026-07-28, remote MCP no longer pins a client to one process with initialize plus Mcp-Session-Id. Every JSON-RPC request carries protocol version and client capabilities in _meta; Streamable HTTP mirrors the useful fields onto HTTP headers so load balancers and gateways can route, rate-limit, and emit metrics without parsing the body. The hard production work is not writing tools/call. It is how you scale replicas, pick a balancer policy, keep nginx from buffering SSE progress, drain subscriptions/listen on a rollout, and where application state lives. We already covered why the protocol is stateless in Stateless MCP and who may call in OAuth 2.1. This article only adds “how it hangs behind an HTTP load balancer.” Transport text: Streamable HTTP. Envelope rules: JSON-RPC 2.0. Do not force this public topology onto a local stdio child process.
The production trap: session affinity collapses horizontal scale
People hear “stateless” and ship one container. The payoff appears only when you use a normal HTTP load balancer—round-robin, least-conn, CPU-weighted—not session affinity. Older revisions used a connection-scoped session: after the handshake, later tools/call had to hit the process that held Mcp-Session-Id. More replicas meant cookie stickiness or IP hash on the ALB or nginx; one dead box killed the whole agent session. 2026-07-28 removed protocol sessions. Any replica should finish any POST on its own.
Cursor, Claude Desktop, OpenAI remote MCP, and home-grown agent hosts do not politely serialize onto one endpoint. The same second can carry tools/list, tools/call, and a long-lived subscriptions/listen. Those three hops need not land on the same instance. Affinity by client IP collapses horizontal scale and also freezes rate limits and canaries to one box. Protocol entry: what MCP is. The agent loop: what an AI agent is.
Against our other layers: statelessness removed the protocol session, not business data. Invoice drafts, carts, unfinished multi-round tool calls still go to Redis or a database and come back via a draftId in arguments. A process-local map is not production state. Auth is not this layer either—Bearer for public remote MCP lives on HTTP headers; see the OAuth article. This piece assumes those three layers are already split and talks only topology and the ops contract.
| Approach | What the load balancer sees | Production outcome |
|---|---|---|
| Cookie / IP affinity + in-process session | Must pin the same client to the same upstream | Hard to scale, rollouts break, single point of failure |
| Stateless replicas + ordinary HTTP LB | Any replica can take any JSON-RPC POST | Horizontal scale, canaries, per-tool limits |
| Serverless cold start + one POST | No long-lived process to “remember” a handshake | Fine for short RPC; long SSE needs its own timeouts |
Target architecture: client → HTTP LB → N stateless replicas
Keep the topology thin: public DNS → TLS termination (ALB, NLB plus sidecar, nginx, Caddy, a cloud LB) → a set of MCP replicas with the same image and same config. Do not hang an “MCP session store” in front of them to paper over the protocol. Health checks use a separate GET /healthz: process up, dependencies reachable. Do not POST an empty body to /mcp, and do not use a live tools/list as the probe—that hits the real registry and can 401, falsely draining the replica.
Each replica must finish one hop alone: validate Origin (DNS rebinding → 403), read MCP-Protocol-Version / Mcp-Method / Mcp-Name, validate Bearer if you protect the resource, parse the JSON-RPC envelope, run the tool, return a single JSON object or a request-scoped SSE stream. The spec wants one MCP endpoint that accepts POST, for example https://mcp.example.com/mcp. GET streams and protocol-level sessions are gone in 2026-07-28. Do not reopen GET /sse “for old health checks.”
What replicas share is application dependencies: database, object storage, third-party APIs, optional Redis. They do not share “which MCP connections are open.” Cloud Run, Cloud Functions, Knative scale-per-request platforms fit this model. What you still tune is idle timeout on long SSE, not session stickiness. Run the process as an unprivileged user behind the proxy; do not bind the MCP process to 0.0.0.0:80 on the public internet.
How JSON-RPC 2.0 crosses the load balancer
MCP encodes messages as JSON-RPC 2.0, UTF-8 required. On Streamable HTTP every client request or notification is a new HTTP POST; servers do not initiate JSON-RPC requests. The load balancer does not need the business meaning of method—it forwards bytes. The 2026 transport mirrors method to Mcp-Method and tool/resource/prompt names to Mcp-Name so intermediaries can rate-limit by tool, shard by method, and canary by version without parsing the body.
The body remains the source of truth. MCP-Protocol-Version on the header must match params._meta.io.modelcontextprotocol/protocolVersion byte for byte or the server MUST return 400 with HeaderMismatch. Clients must also send Accept: application/json, text/event-stream. A notification POST that is accepted returns 202 Accepted with no body; a request returns one JSON object or an SSE stream. The JSON-RPC id pairs this hop’s request and response. It is not a session key and must not be used to “recover context” across replicas.
HTTP/1.1 400 Bad Request
Content-Type: application/json
{
"jsonrpc": "2.0",
"id": 42,
"error": {
"code": -32600,
"message": "HeaderMismatch: MCP-Protocol-Version does not match params._meta"
}
}
Below is a production-shaped tools/call: auth on the HTTP header, envelope in the body, routing keys the gateway can see also on the header. Stuffing an Access Token into params or _meta is not 2026 remote MCP. Argument shape still needs a Schema check: MCP and JSON Schema.
POST /mcp HTTP/1.1
Host: mcp.example.com
Content-Type: application/json
Accept: application/json, text/event-stream
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: searchInvoices
Authorization: Bearer eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9...
{
"jsonrpc": "2.0",
"id": 42,
"method": "tools/call",
"params": {
"name": "searchInvoices",
"arguments": {
"startDate": "2026-01-01",
"endDate": "2026-01-31",
"status": "paid"
},
"_meta": {
"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientCapabilities": {},
"io.modelcontextprotocol/clientInfo": {
"name": "jsonvue-demo",
"version": "1.0.0"
}
}
}
}
Streamable HTTP: buffering, timeouts, and SSE
Short tool calls should return Content-Type: application/json. Long work may return text/event-stream: request-related notifications/progress first, then a final JSON-RPC response that ends the stream. Most production incidents sit in the reverse proxy, not the MCP SDK. nginx proxy_buffering defaults to holding progress events and flushing them as one chunk; the agent looks wedged. The spec tells servers to send X-Accel-Buffering: no for intermediaries. See nginx proxy_buffering.
On Streamable HTTP the cancel signal is the client closing that SSE, not a follow-up notifications/cancelled POST (that is the stdio binding). The LB must propagate backend disconnects to the client and client disconnects to the replica so the worker can stop. Do not put a gateway in the middle that “retries POST”: JSON-RPC requests are not idempotent by default, and tools/call may already have written the database.
subscriptions/listen is a different long stream: the response stays open and carries change notifications such as tools/list_changed, not progress for one call. The spec encourages periodic SSE comment lines (a line starting with a colon) as keep-alives so idle intermediaries do not hang up. Resumable SSE via Last-Event-ID is not supported. Size LB idle/read timeouts above your keep-alive interval; Cloudflare, ALB, and nginx defaults of 60s are often too short. The snippet below is a minimal reverse-proxy sketch—not a security baseline. TLS, rate limits, and WAF are separate.
upstream mcp_replicas {
least_conn;
server 10.0.1.11:8080;
server 10.0.1.12:8080;
server 10.0.1.13:8080;
}
server {
listen 443 ssl;
server_name mcp.example.com;
location /healthz {
proxy_pass http://mcp_replicas;
proxy_connect_timeout 2s;
proxy_read_timeout 3s;
}
location /mcp {
proxy_pass http://mcp_replicas;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header Connection "";
proxy_buffering off;
proxy_cache off;
proxy_read_timeout 3600s;
proxy_send_timeout 3600s;
}
}
Rolling deploys, drain, and subscriptions/listen
Stateless replicas make rollouts easier; SSE is still an in-flight HTTP request. The order: pull the instance from the target group → stop new /mcp traffic → wait for open JSON responses and SSE streams to finish, or hit the drain timeout you published → then SIGTERM the worker. Do not SIGKILL a process that is still pushing progress. Add the instance back only after health checks go green.
The new version must still accept the old request shape. There is no protocol step that “upgrades the handshake, then switches traffic.” A canary is a percentage of POSTs to the new image. MCP-Protocol-Version can be a routing key: only clients that declare 2026-07-28 enter the new pool; older clients stay on a compatibility pool. On a version mismatch return 400 plus UnsupportedProtocolVersionError. Do not silently down-level and keep running tools.
subscriptions/listen will almost always drop during a publish window. Clients should reopen listen and must not assume events were not lost. Servers should not queue undelivered list_changed in process memory. If you need reliable delivery, write an external queue; listen is only the subscription mouth. MRTR (multi-round input) is also independent POSTs: park the intermediate result in shared storage so the next hop can land on another replica.
Gateway limits, auth, telemetry, and JSONVue
If the gateway can see Mcp-Method and Mcp-Name, set QPS per tool, not only per source IP. Expensive tools (writes, payments, long SQL) get their own quota; tools/list can be looser. On the public internet validate Origin and run OAuth 2.1 as a resource server—401, Protected Resource Metadata, Bearer on every hop; see MCP OAuth 2.1. Auth on the transport, limits on the gateway, Schema in the business layer. Do not mash the three into one middleware.
| Telemetry field | Where it comes from | What you use it for |
|---|---|---|
| MCP-Protocol-Version / Mcp-Method / Mcp-Name | Request headers (aligned with body _meta) | Per-tool limits, canaries, dashboards |
| JSON-RPC id | This hop’s envelope | Join client retries to replica logs |
| HTTP status + JSON-RPC error.code | Transport vs method layer | Tell 401 / HeaderMismatch / business error apart |
Access logs should keep at least those three groups. Replica logs add whether jsonrpc is 2.0 and whether the tool ran before a side effect. Header/body mismatch, missing Accept, and a bad Origin should die at the gateway or replica edge, not inside the tool function. Keep one sample of each failure: HeaderMismatch, 401, missing argument fields, buffered SSE (client only saw the last chunk).
You can finish the lab in the browser:JSON formatter to see if the envelope parses;JSON Schema validator for params.arguments;JSON Diff to compare two error responses;JWT decoder for the Bearer aud on a public deploy. Data stays on this machine.
Further reading: what MCP is, Stateless MCP, OAuth 2.1, MCP and JSON Schema, A2A vs MCP.
FAQ
Do I still need sticky sessions in production?
Not at the protocol layer for 2026-07-28. Sticky sessions only make you believe you still have a session. If the app must stick to a region or a data shard, route on Mcp-Param-* or a tenant key in arguments—application routing—not cookie affinity to an MCP process.
Can I health-check the load balancer with GET /mcp?
No. The modern MCP endpoint accepts POST only; GET streams are gone. Random GETs to /mcp yield 405 or confuse leftover compatibility logic. Probe a separate /healthz that checks process and dependencies and does not run tools.
If SSE drops, should the client resume with Last-Event-ID?
The spec does not support resumable SSE. Clients reopen the matching request (listen → a new subscriptions/listen; a long tool → retry only if the business is idempotent). Keep-alive comment lines plus an unbuffered proxy match the contract better than rolling your own event-id cache.
Does a local stdio MCP also sit behind a load balancer?
No. stdio is a child process the client launched; bytes ride stdin/stdout with no HTTP hop. Load balancing, Origin checks, Bearer, and X-Accel-Buffering belong to Streamable HTTP / remote. The same tools can expose both transports; the production topology wraps only the HTTP face.
Summary and next steps
The production architecture of remote MCP is one sentence: a stateless JSON-RPC request crosses an ordinary HTTP load balancer and lands on any identical replica; long streams open in request scope, not connection scope. Headers are for the gateway, the body is the truth, application state lives in external storage.
Ship in this order: any replica can finish one tools/call alone, then disable buffering and raise timeouts and add drain, then per-tool limits and canaries. Keep a success envelope, a HeaderMismatch, and a 401 in JSONVue on this machine. Protocol semantics: Stateless MCP. Who may call: OAuth 2.1.