Message Queues

Decoupling producers from consumers with queues: delivery guarantees, ordering, backpressure, and how Kafka, RabbitMQ, and SQS differ.

Intermediate · 18 min read

Why this matters

Sooner or later, every growing system hits the same wall: Service A needs Service B to do something, but Service B is slow, or down, or drowning. If A calls B directly and waits, A's fate is chained to B's worst day. A message queue breaks that chain. A drops a message in the queue and walks away; B picks it up when it's ready. The queue absorbs spikes, survives outages, and lets each side scale independently. Email notifications, order processing, image resizing, fraud checks — the unglamorous plumbing of the internet runs on queues.

Think of it like a restaurant's order counter. The waitstaff (producers) scribble orders and clip them to the rail — they don't stand there watching the cooks. The kitchen (consumers) grabs tickets at its own pace. If the kitchen gets slammed, tickets pile up on the rail instead of waiters freezing mid-dining-room. The rail is the queue, and the whole restaurant would collapse without it.

Video: What is a Message Queue? — IBM Technology
Shows how a queue turns an 8-second upload into async work, with buffering, retries, and decoupling benefits

The basic shape

flowchart LR
    P1[Producer A] --> Q[Queue<br/>durable buffer]
    P2[Producer B] --> Q
    Q --> C1[Consumer 1]
    Q --> C2[Consumer 2]
    Q --> C3[Consumer 3]

Producers and consumers never talk to each other directly — they only know the queue. That decoupling buys you three things: temporal decoupling (B can be down when A sends), load leveling (spikes become a backlog, not an outage), and independent scaling (add consumers without touching producers).

Interactive diagram: PacketFlow (loads in the app)

Video: What is a MESSAGE QUEUE and Where is it used? — Gaurav Sen (InterviewReady)
Pizza-shop analogy intro: order, pay, wait for your number — the queue takes your request and calls you when it's ready. ID verified via third-party summary page carrying the exact title; channel attributed via InterviewReady links

Delivery guarantees: pick your poison

This is the section that separates "we use a queue" from "we understand our queue":

The practical rule: assume at-least-once, make your handlers idempotent (dedupe keys, upserts instead of inserts), and sleep well.

Video: Kafka Offsets Explained: Why Exactly Once Stops At Your Database — Code with Sam
Your consumer processed message 42 but the offset says 40 — the gap is where duplicates and lost messages live; commit order is the whole game. Caveat: channel unidentified; description is expert-level human writing checked against Kafka 4.3 docs

Ordering and partitioning

Most queues guarantee order within a partition or queue, not globally. Kafka's model is the clearest: a topic is split into partitions, each an ordered log; messages with the same key always land in the same partition. Consumers in a consumer group split partitions among themselves — add a consumer, and partitions get rebalanced.

flowchart LR
    T[Topic: orders<br/>3 partitions] 
    T --> P0[Partition 0<br/>ordered log]
    T --> P1[Partition 1<br/>ordered log]
    T --> P2[Partition 2<br/>ordered log]
    P0 --> C1[Consumer A]
    P1 --> C2[Consumer B]
    P2 --> C2
    Note[("Key = user_id<br/>same key → same partition")]

The catch: a partition is consumed by exactly one consumer in a group, so your parallelism is capped by your partition count. Pick too few partitions and you can't scale consumers; pick too many and rebalancing gets chatty. This is one of those decisions that's cheap to make early and painful to change later.

Video: Kafka Ordering Explained: Why Your Messages Still Arrive Out of Order — Code with Sam
Builds a payment flow, watches events arrive out of order, then pinpoints where Kafka's guarantee ends — with runnable code you can try yourself

Backpressure: when the rail overflows

Queues absorb spikes — until they can't. Every queue is finite, and "what happens when it's full" is a design decision, not an accident:

Interactive diagram: StepThrough (loads in the app)

Video: Your Stream Is Falling Behind: Backpressure, Lag, and Batching — Swetank Srivastava
Separates consumer lag from backpressure and shows how slowdowns travel upstream until spare capacity recovers

Kafka vs. RabbitMQ vs. SQS

KafkaRabbitMQSQS
ModelDurable distributed logSmart broker, flexible routingFully managed queue
RetentionDays/weeks, replayableUntil acknowledgedUp to 14 days
OrderingPer partitionPer queue (mostly)FIFO queues only
Best forEvent streaming, high throughputComplex routing, low latencyZero-ops AWS workloads
You operateBrokers, ZooKeeper/KRaftThe broker clusterNothing

Rough heuristic: need to replay history or stream millions of events/sec → Kafka. Need fancy routing (topic exchanges, per-message TTLs) → RabbitMQ. Need a queue in the next five minutes with no servers → SQS. Many architectures use two of the three and feel zero shame about it.

Video: Kafka vs RabbitMQ vs SQS: Which One Should You Choose? — RandomTalks
Compares the three systems side by side across replay, ordering, retries, and real workloads

Takeaways

  1. Queues decouple producers from consumers in time and scale — spikes become backlogs, not outages.
  2. Assume at-least-once delivery and make consumers idempotent; true exactly-once rarely survives contact with the real world.
  3. Ordering is per-partition, and partitions cap your consumer parallelism — choose counts deliberately.
  4. Plan for full queues and poison messages: backpressure policies and dead-letter queues are not optional.

Check your understanding

  1. What does a message queue primarily decouple?

    • Synchronous from asynchronous code in a single process
    • The frontend from the backend
    • Producers from consumers, in both time and scaling
    • The database from the cache
  2. With at-least-once delivery, what must your consumers be?

    • Limited to one message per second
    • Idempotent — safe to process the same message twice
    • Synchronous
    • Written in the same language as the producer
  3. In Kafka, what limits how many consumers in a group can work in parallel?

    • The number of brokers in the cluster
    • The retention period of the log
    • The size of each message
    • The number of partitions in the topic
  4. What is a dead-letter queue (DLQ) for?

    • Shunting aside messages that repeatedly fail processing so they don't clog the main queue
    • Holding messages during a planned broker upgrade
    • Storing messages that were successfully processed, for auditing
    • Encrypting messages that contain sensitive data

Go deeper

Want to keep pulling this thread? These talks and tutorials go further than we did here:

Sources & further reading