Content

MCP Went Stateless. So Where Do Idempotency Keys Live Now?

The MCP 2026-07-28 spec removed session state from the protocol, which is the right call for scaling but leaves retry safety completely unaddressed. We walk through the duplicate-refund failure, the three places an idempotency key can actually live, and the trade-offs of each.

EngineeringOct 13, 2026

Key takeaways

  • -The MCP 2026-07-28 spec removed the initialize handshake and the Mcp-Session-Id header, so every request is now self-describing and any server instance can serve any call.
  • -The spec says nothing at all about idempotency, retry safety, or request-ID semantics, and that silence is an absence rather than a design feature.
  • -Explicit handles returned by a tool solve application state, but they do not solve retry safety, because the retry is the client's decision and the side effect has already happened.
  • -An idempotency key can live in the client, in your tool handler, or in the execution layer underneath the tool, and each of those placements has a different failure mode.
  • -Swytchcode sets idempotency per integration in the execution_policy block of manifest.json, with mode none as the default and mode dynamic generating and reusing a key across retries.

A customer support agent decides to refund an order. It issues a tools/call against your MCP server, your handler calls Stripe, Stripe starts processing, and then the connection between the client and your server drops at seven seconds. The client's HTTP request fails. The client has no response, no result, no error body from the provider. It has a timeout.

So it retries. That is the correct, boring, textbook thing for an HTTP client to do.

Now the customer has two refunds.

Nothing in the Model Context Protocol told the client that retrying was unsafe. Nothing in the protocol gave it a way to ask whether the first call landed. And after the 2026-07-28 revision of the spec, there is no session the client could have consulted even if it wanted to. The protocol is stateless now, by design, and retry safety was never in it to begin with.

This post is about where the key goes instead.

What the 2026-07-28 spec actually changed

The short version: MCP dropped the idea that a client and a server have a relationship. They have requests.

The initialize / initialized handshake is gone, and so is the Mcp-Session-Id header (SEP-2575 and SEP-2567). There is no negotiated session to establish and no identifier to carry forward. Instead, every request is self-describing. Protocol version, client identity and client capabilities travel in _meta, for example under the io.modelcontextprotocol/clientInfo key, and requests carry an MCP-Protocol-Version header.

If a client wants to know what a server can do before it starts calling, there is an optional server/discover RPC it can use. Optional is the important word. Clients may skip it entirely and just start making calls, so a server cannot assume discovery happened.

A few other changes matter for anyone operating a server:

  • Streamable HTTP requests must now include Mcp-Method headers (for example tools/call) and Mcp-Name headers (for example the tool name), per SEP-2243. That is a real gift for routing, metrics and rate limiting at the edge, because a proxy can see what a request is without parsing the JSON-RPC body.
  • tools/list, prompts/list, resources/list and resources/read responses now carry ttlMs and cacheScope (SEP-2549). List results are cacheable, with the server stating the terms.
  • Server-initiated elicitation/create and sampling/createMessage are replaced by Multi Round-Trip Requests, or MRTR (SEP-2322). A server that needs more input returns resultType: "input_required" along with the requests it needs answered, and the client retries the original call with answers in inputResponses. The server does not call back into the client. It asks, by returning.
  • Notifications no longer come from the old HTTP GET endpoint. They move to a subscriptions/listen stream under the Tasks extension, which also adds tasks/get and tasks/update.
  • Roots, Sampling, Logging (SEP-2577) and the HTTP+SSE transport are deprecated but supported for at least twelve months.
  • Dynamic Client Registration is deprecated in favour of Client ID Metadata Documents, and clients must validate the iss parameter per RFC 9207 before redeeming an authorization code (SEP-2468).

We read the change set through Appwrite's summary of the 2026-07-28 specification, which is a secondary source rather than the spec itself. Treat it, and this post, as orientation. Check the official specification for exact wire semantics before you ship anything that depends on them.

Where state genuinely has to survive across calls, the spec's answer is explicit handles. A tool returns an identifier as part of its result, the model carries that identifier around in context, and it passes it back as an argument on the next call. The state lives in the tool's own storage, keyed by something the model holds. The protocol carries nothing.

Statelessness is the right call

We want to be clear that we think this change is good, because the rest of this post is about a gap it exposes, and we do not want that read as a complaint.

The old model meant a session was pinned to whichever server process handled initialize. That gave you three unpleasant options: sticky sessions at the load balancer, a shared session store that every instance reads and writes, or an accidental single-instance deployment that nobody admits is a single-instance deployment. All three are operational tax. All three fail in ways that are annoying to debug, because the symptom is usually "it works until we scale out."

With session state removed, any instance behind a plain round-robin load balancer can serve any request. You can deploy MCP servers the way you deploy any other stateless HTTP service: horizontal autoscaling, rolling restarts, no drain logic for in-flight sessions, no session affinity cookies. A pod can die mid-traffic and the next request just goes somewhere else.

The ttlMs and cacheScope fields make that even cheaper, because tools/list stops being a per-client round trip on every connection and starts being something a CDN or a client-side cache can hold.

This is a straightforwardly better deployment story. We are not arguing against it.

The gap: handles solve state, not safety

Here is the distinction that we think gets blurred in the "state moved to handles" summary.

Handles solve application state. A pagination cursor, a draft ID, a multi-step workflow position: all of those are facts about where the conversation is, and all of them are fine to hand to the model and get back later.

Handles do not solve retry safety, and they cannot, for two reasons.

The first is timing. A handle is returned in a response. If the request failed before the response arrived, there is no handle. The exact window where you need protection is the window where the mechanism is unavailable.

The second is authority. The retry is the client's decision. A client times out, decides the call did not complete, and sends it again. Your server is not consulted. There is no handshake in which the server could say "that request already executed, here is the original result." The protocol has no request-deduplication semantics, because the protocol has no opinion about request identity surviving a failed transport.

And in between those two, the side effect has already happened. Stripe charged the card. The email went out. The invoice exists. Statelessness at the protocol layer does not make the world stateless. It just means the protocol is no longer the place where you can find out what the world did.

This is the part worth being blunt about: the MCP specification says nothing about idempotency, retry safety, or request-ID semantics. Not in 2026-07-28, not before it. This is an absence in the spec, not a feature of it. Nobody removed idempotency from MCP, because it was never there. What 2026-07-28 did was make the absence impossible to paper over, because there is no longer a session you could have quietly hung a dedupe cache off.

MRTR makes this slightly sharper, incidentally. MRTR works by the client retrying the original call with inputResponses attached. That is a legitimate, spec-sanctioned re-send of a request the server has already partially processed. If your tool does work before it returns input_required, you need to think about what happens on that retry. The protocol is handing you a replay as normal operation.

Three places an idempotency key can live

An idempotency key is just a caller-supplied token that tells the provider "if you have seen this token before, return what you returned last time instead of doing the thing again." Stripe, Square, Adyen, PayPal, Shopify, SendGrid and most other payment and messaging APIs support some version of it. The question is who generates it and who remembers it.

There are three honest answers.

Where the key livesHow it worksWhat it gets rightWhat it gets wrong
In the MCP clientThe client generates a key per logical operation and passes it as a tool argument. The tool forwards it to the provider.The client is the only component that knows a retry is a retry, so it is the only one that can reuse a key across a transport failure.Requires every client to implement it, and MCP does not define an argument for it, so each server invents its own parameter name. A model that re-reads its own transcript and decides to call the tool again will mint a new key, because to the model it is a new intent.
In your tool handlerThe handler generates a key, stores a mapping from some request fingerprint to that key, and reuses it when it sees the same fingerprint again.Fully under your control. No client cooperation needed. Works with any client, including ones that know nothing about idempotency.You are now running a store that must be shared across instances, which is the sticky-session problem the spec just removed, wearing a different hat. You also have to define "the same request," and argument hashing is a bad proxy for intent. Two genuine refunds of the same amount for the same customer look identical.
In the execution layer underneath the toolThe layer that actually performs the outbound HTTP call owns the key: it generates one per execution, attaches it to mutating requests, and reuses the same key for its own retries.The retry and the key are owned by the same component, which is the only arrangement where reuse is guaranteed rather than hoped for. It applies uniformly across providers, and your tool handler contains no retry bookkeeping at all.It can only protect retries it performs itself. If the MCP client times out and re-calls your tool from scratch, that is a new execution and gets a new key, unless the client supplied one. It is also another dependency in the path.

Read that last cell carefully, because it is the honest limit and we would rather state it than let you discover it. Nothing below the client can see a client-level retry. If your client times out at seven seconds and re-issues tools/call, any layer underneath is looking at a request it has never seen. Defence in depth means the client should pass a key through when it can, and the layer underneath should guarantee key stability for everything else.

What an execution layer does buy you is that the much more common failure is covered without anyone writing code for it. Most duplicate side effects we see are not clean client-level retries. They are a socket reset mid-flight, a 502 from a provider's edge, a connection closed after the provider committed but before it responded. Those are retries performed by whatever component is holding the HTTP connection, and that component is the one that should be carrying the key.

What we do about it

Swytchcode is an execution layer: the model or the tool handler says which call to make, and Swytchcode performs it with credentials, policy checks, retries and an audit trail. We wrote about the general shape of that in what an execution layer is and why AI agents need one.

Idempotency is configured per integration, in the execution_policy block of that integration's manifest.json. There are two modes.

none is the default. No key is attached, and behaviour is exactly what you would get calling the provider yourself.

dynamic generates a key and attaches it to mutating requests for providers that support it.

{
  "execution_policy": {
    "idempotency": {
      "mode": "dynamic",
      "header_name": "Idempotency-Key"
    }
  }
}

Every field is optional. Omit them and you fall back to mode: "none" and header_name: "Idempotency-Key", so the header name above is the default and you only set it for a provider that expects something else.

The scoping rule is the one to internalise: each exec call gets its own key, and an unrelated call gets a new one. That is what makes this safe to leave on. Turning it on does not mean two deliberate refunds collapse into one.

The retry path is the whole point:

  1. The request goes out with a generated key.
  2. Something fails before the response arrives. A dropped connection, for example.
  3. The retry reuses the same key.
  4. The provider recognises the key and returns the original result rather than running the operation twice.

Step three is the part that is hard to get right by hand, because the natural place to generate a key is inside the function you are retrying. The write-up on how Swytchcode retries failed API calls covers how the backoff itself behaves. The thing to note here is that retries and keys are owned by the same code path, so they cannot drift apart.

Claude Code running swytchcode list and swytchcode info commands to find the right methods before executing

Finding a method with swy list methods and inspecting its inputs with swy info before executing it. The canonical ID you look up here is the unit that gets its own idempotency key at execution time, which is why "one key per exec call" is a rule you can reason about straight from the tool definition.

From a tool handler, the call looks like this. For a Stripe refund, the verified canonical method is stripe.refund.create3, described as "Refund a charge or payment intent":

import { exec } from "@swytchcode/runtime";

export async function refundTool({ paymentIntentId, amount }) {
  const res = await exec("stripe.refund.create3", {
    body: { payment_intent: paymentIntentId, amount }
  });
  return res;
}

There is no key in that code, no retry loop, and no fetch. There is also no .env file in the project. Swytchcode keeps credentials in its own local store at ~/.swytchcode/credentials.db, outside the project folder, and reads them when a call runs, so the handler never sees a secret. You connect a provider once, interactively:

swy auth connect stripe
swy auth status

Execution mode is a project setting rather than something your code touches. swy init writes it into .swytchcode/tooling.json:

swy init --mode=sandbox
Terminal running swytchcode init and choosing an editor and execution mode

swy init asking for an editor and an execution mode. Mode is answered once and stored in .swytchcode/tooling.json, which is why sandbox versus production never appears as an environment variable or as a flag in the handler code above.

The documentation worth reading rather than taking our word for is the idempotency guide and the manifest.json reference, where per-integration idempotency sits alongside endpoints, retries and timeouts. If you want the config files explained end to end, we have a breakdown of tooling.json, manifest.json and policies.json, and swy exec explained covers what a single execution's inputs, outputs and exit codes look like.

The docs' own best-practice guidance is short, and we agree with it: turn it on for mutating APIs such as payments, invoices and emails, leave it off for read-only GETs where it buys nothing, and never generate a new key during a retry. That last rule is the one the whole mechanism rests on. For a longer walkthrough of the failure modes, see idempotency explained: how Swytchcode prevents duplicate charges and duplicate emails.

One caveat we will state plainly: the provider has to support idempotency keys for any of this to do anything. Attaching a header to an API that ignores it changes nothing. Check the provider's documentation, including how long it remembers a key, because retention windows vary and a retry outside the window is a fresh operation.

You can also just do this yourself

We would rather make the real argument than the convenient one, so: you can solve this in your tool handler. It is not exotic.

Generate a key before the operation, store it keyed by something you trust as an intent identifier, attach it to the outbound call, and reuse it if you retry. If your tool handler already talks to a database, the storage is close to free.

The reasons we think it belongs lower down are specific rather than sweeping. The first is that per-provider header names and semantics are a long tail of small differences, and each one is a chance to get it slightly wrong. The second is that the store has to be shared across instances to survive the stateless deployment model, which means you have reintroduced a piece of shared state into a server you just made stateless. The third is that it will not be the only cross-cutting concern: retries, timeouts, credentials, policy and audit all want to live in the same place, and once four of them do, the fifth is cheaper there too.

If you have one tool that calls one provider, write it yourself. If you have fifteen tools calling eight providers and three of them move money, we think the key belongs under the tool rather than in it.

A migration checklist for server authors

If you are moving an existing server off Mcp-Session-Id, here is what we would work through. We have ordered it roughly by how likely each item is to bite.

  1. Find every read of Mcp-Session-Id. Grep for the header name and for whatever your framework calls its session object. Each hit is either state that moves into an explicit handle, or state that was never needed.
  2. Convert session-held state into handles. Decide what the handle identifies, where it is stored, and how long it lives. Return it in the tool result and accept it as a named argument. Give it a TTL, because nothing will clean it up for you now.
  3. Stop assuming initialization ran. There is no handshake. Read protocol version, client identity and capabilities from _meta and the MCP-Protocol-Version header on each request, and have a defined behaviour for when they are missing or unfamiliar.
  4. Do not require server/discover. It is optional for clients. A client may call a tool it never discovered, so your validation has to stand on its own.
  5. Replace server-initiated elicitation and sampling with MRTR. Return resultType: "input_required" with the requests you need, and handle the client's retry carrying inputResponses. Make sure the work you do before returning input_required is safe to see twice.
  6. Move notifications to subscriptions/listen. The HTTP GET notification endpoint is gone. If you are adopting the Tasks extension, tasks/get and tasks/update come with it.
  7. Emit Mcp-Method and Mcp-Name on Streamable HTTP, and then use them. These headers are required, and they are also the cheapest observability win in the change set, because your proxy can now tell a tools/list from a tools/call without reading a body.
  8. Set ttlMs and cacheScope deliberately on your list responses. Deciding to think about it later means clients guess.
  9. Audit OAuth. Dynamic Client Registration is deprecated in favour of Client ID Metadata Documents, and if you act as a client you must validate iss per RFC 9207 before redeeming a code.
  10. Keep the deprecated paths alive. Roots, Sampling, Logging and HTTP+SSE have at least twelve months. Removing them early only breaks your own users.
  11. Then, separately, go through every mutating tool and decide where its idempotency key comes from. Client argument, handler-managed store, or execution layer. Write the answer down per tool. Any tool where the answer is "we have not decided" is a tool that will produce a duplicate eventually.
  12. Drop sticky sessions and session stores from your deployment. This is the payoff. Once nothing reads a session, the load balancer config gets simpler and the autoscaler stops being a liability.

Item eleven will not appear in a migration guide written from the spec, because the spec does not mention it. That is exactly why it is worth putting on the list.

FAQ

Does MCP 2026-07-28 define an idempotency key?

No. The specification does not address idempotency, retry safety, or whether a request ID means anything after a transport failure. This is an absence rather than a deliberate feature, and it was true before this revision too. The 2026-07-28 changes make it more visible because there is no session left to attach a deduplication cache to.

Can I just use the JSON-RPC request ID for deduplication?

Not safely. JSON-RPC IDs exist to match a response to a request, and the protocol says nothing about their uniqueness or meaning across connections. A client that retries after a timeout may reuse an ID or mint a new one, and both are defensible readings. Building dedupe on top of it means depending on behaviour no client ever promised.

Do explicit handles solve this?

They solve application state, which is what they are for. They do not solve retry safety, because a handle arrives in a response, and the dangerous case is the one where no response arrived. The side effect can happen without you ever getting a handle back.

Where should the key live if I only control the server?

In the execution path that performs the outbound call, whether that is your own code or a layer under it. Key generation and retry logic have to be owned by the same component, or they will drift. If the client can also pass a key as a tool argument, accept it and prefer it, since the client is the only component that can see a client-level retry.

Does enabling dynamic idempotency break deliberate repeat calls?

It should not, and in Swytchcode's case it does not, because each exec call gets its own key and an unrelated call gets a new one. Two separate refunds are two executions with two keys. The reuse only happens within one execution's own retry sequence.

Should I turn idempotency on for read requests?

There is no benefit. A GET that runs twice returns data twice. The documented guidance is to enable it for mutating APIs such as payments, invoices and emails, and to skip it for read-only GETs. Leaving it off where it does nothing also keeps your request logs easier to read.

What if the provider does not support idempotency keys?

Then the header is ignored and you get no protection, so you need a different strategy. Options include a provider-side uniqueness constraint you can rely on, a pre-check read before the write where the race window is acceptable, or treating the operation as something that requires human confirmation. Check the provider's retention window too, because a key the provider has forgotten is the same as no key at all.

Does MRTR create a replay problem?

It can. MRTR works by the client re-sending the original call with inputResponses attached, which means your server sees a request it has already partially handled. Treat anything your handler does before returning input_required as code that will run more than once, and keep side effects after the point where you have all the input you need.

Is statelessness actually an improvement, given all this?

Yes. The problems in this post are not caused by statelessness. They were always there, and they were previously easy to hide behind a pinned session. What the change buys you is real: any instance serves any request, no sticky sessions, no shared session store, cacheable list results, and ordinary horizontal scaling. The cost is that retry safety is now visibly your job rather than invisibly your job.

Wrapping up

MCP 2026-07-28 made the protocol honest. It carries no session, it says nothing about retries, and it does not pretend to. That is a cleaner contract than the one it replaced.

The work it leaves you is to decide, per mutating tool, where the idempotency key comes from. Our view is that it belongs in the execution layer underneath the tool, because that is where the retries happen, and a key is only useful if the thing retrying is the thing that remembers it. Your view might reasonably be that your single payment tool does not need another dependency, and that is a fine answer as long as it is an answer someone actually gave.

The failure mode we want you to avoid is the one at the top of this post, where nobody decided and the client's perfectly reasonable timeout turned into a second refund.

If you want to try the configured version, npx swytchcode installs the CLI, and swytchcode.com/skills.md is the agent-readable file you can hand to a coding agent so it sets the project up with the right commands instead of guessing at them.

More content