Idempotency: Making Retries Safe

Networks fail, so we retry — but a retry can double-charge a credit card. How idempotency keys and idempotent operations make retries safe.

Intermediate · 17 min read

Why this matters

Picture this: your app charges a customer's card, the network hiccups, and the response never arrives. Did the charge go through? You have no idea — so you retry. The retry goes through. The first request went through too. Your customer just got charged twice, and your support queue just got its worst ticket of the week.

This is the most expensive two-line bug in distributed systems, and it happens because of a perfectly reasonable instinct: when something fails, try again. Retrying is correct — networks are flaky, and giving up on the first timeout would be worse. The fix isn't to stop retrying. It's to make retries safe. That property has a name: idempotency.

Video: Idempotency on AWS — Explained — System Design Lab
Joud opens on every backend dev's nightmare — the double charge — then shows why retries make it inevitable and how to stop it

Pressing the elevator button twice

This lesson's running analogy: the call button outside an elevator. You press it, the button lights up, and nothing seems to happen — so you press it again. And maybe a third time, because elevators test everyone's patience. The elevator still sends one car. Pressing the lit button again does nothing new. That's idempotency in the physical world: repeat the action all you want, the effect happens once.

Map it to your system:

The elevator doesn't achieve this by being clever. It achieves it because "come to this floor" is an operation where repeats are harmless. Not every operation is built that way — and that's where the trouble starts.

Video: Why Idempotency is very critical in Backend Applications — Hussein Nasser
Hussein Nasser builds the core intuition — repeat the same call, get the same result — the idea behind the elevator button

The precise definition

An operation is idempotent if performing it once has the same effect as performing it multiple times. The textbook shorthand: f(f(x)) = f(x).

Some operations are naturally idempotent — repeats are harmless by construction:

And some are stubbornly not:

Notice the pattern: absolute operations ("set it to X") are idempotent; relative operations ("add X to it") are not. This maps directly onto HTTP, and it's not a coincidence: the HTTP spec declares PUT and DELETE idempotent and POST not. PUT /settings {"theme": "dark"} — send it five times, the theme is dark. POST /charges — send it five times and find out what your refund policy is.

Drop the analogy when precision matters: idempotency is a property of the operation's effect, not of the transport. A retry that arrives twice must not apply its effect twice.

Video: Idempotency in Cows and REST APIs — Todd Fredrich (RestApiTutorial)
A short classic from REST tutorial author Todd Fredrich that nails the plain-English definition: same call, same effect, every time

Why networks force this on you

You might wonder why we can't just avoid retries. Because the alternative is worse, and because timeouts are fundamentally ambiguous. When you send a request and hear nothing back, three things might have happened:

  1. The request never arrived. Safe to retry.
  2. The request arrived and the reply got lost. Not safe to retry — unless the operation is idempotent.
  3. The request is still being processed, slowly. Retrying now creates a duplicate.

You cannot distinguish these from the client side. The network swallowed the evidence. This is the same gremlin from the CAP theorem lesson wearing a different hat: the network doesn't just partition, it lies by omission. So you must retry — a system that gives up on every timeout is a system that fails constantly — and therefore your operations must tolerate being executed more than once.

There's a classic way to frame this, and it shows up in the message-queues lesson too: at-least-once delivery plus an idempotent consumer equals effectively-once processing. The queue promises "I'll deliver this at least once" (it may deliver twice — Amazon's SQS standard queues document exactly this; their FIFO queues even dedupe for you via message deduplication IDs, the queue acting as the lit button). Your consumer promises "processing the same message twice has no extra effect." Multiply those two promises and you get the behavior everyone actually wants: each message takes effect once.

Video: Distributed Systems 2.1: The two generals problem — Martin Kleppmann
Kleppmann proves why unreliable networks make certainty impossible, forcing retries and duplicate handling.

Idempotency keys: the lit button

For operations that aren't naturally idempotent — the charges, the orders, the "add to balance" writes — you need a mechanism. Enter the idempotency key: a unique token the client generates for each logical operation (a UUID works), sent along with the request, typically in a header like Idempotency-Key.

The server's contract is simple:

  1. First time it sees the key: execute the operation, store the key and the response, return the response.
  2. Any repeat with the same key: skip execution entirely, return the stored response.

Now walk through the double-charge scenario. Your app generates key k-9f2 and POSTs the $50 charge. Timeout — no response. Was it case 1, 2, or 3 from above? Doesn't matter. You retry with the same key k-9f2. If the first request executed, the server finds the key, skips the charge, and returns the stored "charged $50" receipt. If the first request never arrived, the server executes it now. Either way: one charge, one receipt. The button was lit; the extra presses did nothing.

A few engineering details worth knowing:

sequenceDiagram
    participant C as Client
    participant S as Server
    participant DB as Database

    C->>S: POST /charges<br/>Idempotency-Key: k-9f2
    S->>DB: execute charge, store (k-9f2 → receipt)
    S--xC: response lost in the network
    Note over C: Timeout. Did it go through?<br/>Retry with the SAME key.
    C->>S: POST /charges<br/>Idempotency-Key: k-9f2
    S->>DB: key k-9f2 already stored?
    DB-->>S: yes → return stored receipt
    S-->>C: 200 (stored receipt, no second charge)

Note what the diagram doesn't show: the client never needed to know which failure case it was in. That's the whole point. The key collapses all three ambiguous outcomes into one safe behavior.

The server-side logic, as a decision the code makes on every request:

flowchart TD
    R["Request arrives<br/>with Idempotency-Key"]:::client --> L["Key seen before?"]:::service
    L -->|Yes| S["Return stored response<br/>no re-execution"]:::service
    L -->|No| E["Execute the operation"]:::service
    E --> W["Store key → response<br/>with a TTL"]:::data
    W --> R2["Return the fresh response"]:::service

One subtlety the diagram hides: the "key seen before?" check and the store must be atomic. Two retries racing each other could both see "not seen," both execute, and both store — the double-charge you were trying to prevent. The standard fix is a uniqueness constraint on the key column: the loser's insert fails, and once the winner finishes it reads the winner's stored response and returns that (if the winner is still in flight, the loser returns 409 instead — see below). Your database already knows how to do atomic inserts; let it.

Video: Idempotency Keys - Redis for Developers — The Software Mentor
One atomic SET NX claims the slot — like the elevator button lighting up: claimed means the work is already handled, so skip it

The key has three states, not two

That flowchart asks "key seen before?" — but "seen" hides two very different situations, and handling them the same way is a bug. A key is really in one of three states:

  1. Unseen. Execute the operation, store key → response, return it. The normal path.
  2. Seen and completed. Skip execution, replay the stored response. The lit button.
  3. Seen but still in progress. The first request hasn't finished — remember timeout case 3, slow, not lost. There's no stored response to replay yet, and executing again would double-apply the effect. So what do you return?

Stripe's answer, and the industry convention: 409 Conflict — "a request with this key is already being processed." The client treats a 409 like a timeout: back off, wait, retry later with the same key. Eventually the first request finishes and stores its response, and a subsequent retry gets the replay. Back to the elevator: the button is lit, but the car hasn't arrived yet. Pressing again doesn't summon a second car — it just reminds you one is already coming.

flowchart TD
    R["Request arrives<br/>with Idempotency-Key"]:::client --> L{"Key state?"}:::service
    L -->|"Unseen"| E["Execute, store key → response"]:::service
    L -->|"Completed"| P["Replay stored response"]:::service
    L -->|"In progress"| C["409 Conflict<br/>client backs off, retries same key"]:::service
    E --> R2["Return fresh response"]:::service

Two more details that separate a real implementation from a sketch:

And the test that proves all of this works: fire N requests carrying the same key at the same instant and assert exactly one side effect. Exactly one charge, one order, one row. If your system fails that test, the key column is missing its uniqueness constraint — go back and add it. Concurrency is where idempotency designs go to be graded, and the grader is strict.

Video: Apache Kafka Idempotency Explained / Prevent Duplicate Messages / Kafka Idempotent Producer — Coding with Neeraj
Neeraj shows how Kafka's idempotent producer tags every message with a producer ID and sequence number, so the broker can tell a fresh send from a retried one

Failure modes: where idempotency breaks

Idempotency keys are simple, which means the bugs hide in the details:

Video: Your Background Jobs Are Silently Running Twice — The Merge Log
Four ways jobs go wrong — lost jobs, double-runs, poison messages, thundering herds — and why idempotency, not exactly-once, is the actual fix

Retries need manners, too

Idempotency makes retries safe. It doesn't make them polite. A client that retries a timed-out request instantly, ten times in a row, turns one failure into a small DDoS — and if every client does it during an outage, the retry storm finishes what the outage started. The fix is the rate-limiting lesson's cousin: exponential backoff with jitter — wait 1s, then 2s, then 4s, with randomization so a thousand clients don't retry in lockstep.

Think of it as elevator etiquette: pressing the lit button again is harmless, but jabbing it forty times a second doesn't summon the car any faster. It just wears out the button. Idempotency is the lit button; backoff is the patience.

Video: Why a Timeout Doesn't Mean It Failed — The Hot Path
A timeout means "I don't know," not "no" — so every retry is a bet, and polite retries bring backoff, jitter, and an idempotency key

Idempotent vs. not: a toggle

Before you go further, build the instinct. Flip between the two and feel the difference — it's the same distinction as "set the thermostat" versus "nudge it up":

Interactive diagram: VsToggle (loads in the app)

The redesign hint in that footnote is worth internalizing: when you can, prefer the absolute form. "Set balance to $100" beats "add $10" every time a retry is possible — and a retry is always possible. The best idempotency mechanism is an operation that doesn't need one.

Video: Idempotency in ASP.NET Core (.NET 10): The Definitive Guide for 2026 — Frontend to Backend with Rohan
Rohan lines up GET vs POST vs PUT to show which methods are idempotent by nature — then builds the key pattern in Redis for the ones that aren't

The exactly-once illusion

A word of caution before you feel too safe: nobody can promise exactly-once delivery across a network. Not your queue, not your framework, not a vendor's marketing page. The network can always duplicate or delay, and "exactly once" would require the sender and receiver to agree on what happened during a partition — the CAP theorem already told you how that story ends.

What you can build is effectively-once: at-least-once delivery where duplicates are harmless because the consumer is idempotent. Webhooks are the classic arena for this. A payment provider will retry a payment.succeeded webhook until you acknowledge it — their docs will say "expect duplicates." Your endpoint must therefore treat a webhook like a lit elevator button: record the event ID, and if you've seen it, acknowledge without re-applying it. The distributed-transactions lesson wrestles with the same impossibility from the coordination side; here the answer is humbler and cheaper — just make repeats harmless.

So when someone says their pipeline is "exactly once," translate it charitably: they mean effectively-once, built on at-least-once delivery and idempotent consumers, with the duplicates quietly absorbed. Now you know what to ask them: "where do you dedupe, and what's the key?"

Video: The Two Generals' Problem — Tom Scott
Tom Scott tells the classic story of two generals who can never be sure a message arrived — the same math that pops the exactly-once bubble

Takeaways

  1. Timeouts are ambiguous — you can't tell whether a request was lost, processed, or is just slow — so you must retry, and therefore your operations must tolerate repeats.
  2. Idempotency means f(f(x)) = f(x): repeating the operation has the same effect as doing it once. The elevator sends one car no matter how many times you press the lit button.
  3. Prefer naturally idempotent operations: absolute writes ("set to X") over relative ones ("add X"). HTTP already encodes this: PUT and DELETE are idempotent, POST is not.
  4. For operations that can't be naturally idempotent, use idempotency keys: the client generates one key per logical operation, the server stores key → response and replays the stored response on duplicates. Same key across retries — a new key per retry defeats the mechanism.
  5. Exactly-once delivery is impossible over a real network; build effectively-once instead: at-least-once delivery (like SQS standard queues) plus idempotent consumers that dedupe on a key.
  6. A key has three states, not two: unseen (execute and store), in progress (return 409 — the client backs off and retries with the same key), and completed (replay the stored response). A key reused with different parameters is a bug, not a retry — reject it.

Check your understanding

  1. You press the elevator call button, nothing happens, so you press it again. In an idempotent system, what arrives?

    • Exactly one elevator — repeat presses have no additional effect
    • One elevator per press, but the extras are refunded
    • No elevator at all, because duplicate presses cancel each other out
    • One elevator per press — the system dispatches every request it receives
  2. Which of these operations is naturally idempotent?

    • Increment a page-view counter by one
    • Append a line to a log file
    • POST /orders — create a new order
    • PUT /profile — replace the profile with this exact content
  3. Your charge request times out, so you retry with the same idempotency key. The first request actually succeeded. What does the server do?

    • It charges the card again, because a retry is a new request
    • It rejects the retry with an error, since the key is already used
    • It ignores the key and executes the operation only if the amount differs
    • It skips execution and returns the stored response from the first request
  4. A queue guarantees at-least-once delivery and your consumer is idempotent. What processing semantics do you get?

    • At-most-once: some messages will be silently dropped
    • Effectively-once: every message takes effect exactly one time
    • Unordered delivery with no duplicate protection
    • Exactly-once delivery, guaranteed by the queue alone
  5. A retry carrying the same idempotency key arrives while the first request is still being processed. What should the server do?

    • Reject the key as expired and ask the client to generate a new one
    • Execute the operation again so the client gets a faster response
    • Return 409 Conflict, so the client backs off and retries later with the same key
    • Replay the stored response from the first request

Go deeper

Want to keep pulling this thread? These talks and tutorials go further than we did here:

Sources & further reading