Observability for Distributed Systems

Metrics, logs, and traces that actually answer questions: RED and USE, sampling math, cardinality explosions, and SLIs, SLOs, and error budgets.

Advanced · 22 min read

Why this matters

"Is the system healthy?" is a surprisingly hard question when "the system" is 40 services, 3 regions, and a queue with opinions. Observability is the discipline of making systems answer questions about themselves — especially the questions you didn't think to ask. Monitoring tells you when something is wrong; observability helps you figure out why. At 3 AM, the difference between a dashboard that says "error rate spiked" and a trace that says "the spike is exactly these requests hitting exactly this shard" is the difference between a 10-minute fix and a 4-hour war room.

Analogy: monitoring is the check-engine light — it tells you something's wrong, in one bit. Observability is the full diagnostic port a mechanic plugs into: live sensor readings, freeze-frame data from when the fault occurred, the ability to ask "show me fuel pressure during the last misfire." You need the light to know to look; you need the diagnostics to actually fix it.

Video: Charity Majors — Observability and the Glorious Future — Chariot Solutions
The Honeycomb co-founder who defined modern observability on why distributed systems broke old monitoring — like navigating rapids with a map drawn for a canal.

The three pillars (and why traces win arguments)

flowchart TD
    R["Request: GET /checkout<br/>total 412ms"] --> A["gateway: 8ms"]
    R --> B["orders svc: 380ms"]
    B --> C["inventory check: 12ms"]
    B --> D["payments svc: 350ms"]
    D --> E["fraud check: 340ms ⚠️"]
    R --> F["notify svc: 24ms"]

One look at that trace and you know the fraud check owns the latency — no guessing, no SSHing into four boxes. Logs would show you four separate "slow request" lines; the trace shows you the relationship. In microservices, traces are the closest thing to a stack trace you'll ever get.

Video: What is Observability? | Grafana for Beginners Ep. 1 — Grafana
Walks through metrics, logs, and traces — and shows why traces settle debugging arguments.

RED and USE: what to measure

Two mnemonics keep dashboards focused:

And measure latency as percentiles, not averages. An average latency of 50 ms can hide 1% of users waiting 5 seconds — the average lies because it lets the happy majority vote down the miserable minority. Track p50, p95, p99 (and p99.9 if you're fancy); alert on the tail, because the tail is where your angriest users live.

Video: The RED Method: How to Instrument Your Services — GrafanaCon EU (Tom Wilkie)
Grafana's Tom Wilkie on the RED method: USE checks the engine, RED asks the passengers how the ride feels — together they tell you what to measure.

Sampling: the math of seeing enough

You cannot keep every trace. A service doing 10k requests/sec emitting 2 KB spans per request generates ~1.7 TB of trace data per day — per service. So you sample:

The back-of-envelope: if your error rate is 0.1% and you want every error trace, tail-based sampling at "all errors + 1% of successes" stores ~1.1% of traffic — a 90× cost reduction over keeping everything, while retaining 100% of the interesting cases. Do the arithmetic for your own traffic; the answer is always "sample, but sample smart."

Video: How to find failures without drowning in tracing data — The New Stack
Trace-sampling strategies that catch failures without drowning your budget in span data.

Cardinality: the bill you didn't expect

Metrics systems index every unique label combination as a separate time series. http_requests_total{path="/checkout", status="500"} is one series. Now add user_id as a label with a million users: you've created millions of series, and your Prometheus is having a very bad day — slow queries, exploding memory, and a storage bill that arrives like a ransom note.

Rules of thumb: labels should be low-cardinality (region, service, status code, route template like /users/:id — never the raw path /users/8472, never user IDs, never request IDs). High-cardinality data belongs in logs and traces, which are designed for it. The cardinality explosion is the #1 self-inflicted observability outage; the second is alerting on the wrong thing (see below).

Interactive diagram: StepThrough (loads in the app)

Video: Containing Your Cardinality — Prometheus Monitoring
Prometheus labels are filing tags — handy until every request gets its own tag and the cabinet explodes; a PromCon talk on when cardinality helps and how to tame it.

SLIs, SLOs, and error budgets: alerting like an adult

Raw metrics don't tell you what's acceptable. The SRE vocabulary fixes that:

The numbers make trade-offs concrete. 99.9% ("three nines") = 43 min/month downtime; 99.99% = 4.3 min/month; 99.999% = 26 sec/month. The engineering cost of each additional nine grows steeply and non-linearly while the user-visible difference shrinks. This is why "how many nines" is a business negotiation, not an engineering aspiration — and why alerting should fire on burn rate (how fast you're spending the budget) rather than on every blip. A 5-minute spike that consumes ~12% of the monthly budget is not a page; a sustained burn that eats 20% in an hour is.

Video: Mathematics of SLOs — SRECon EMEA 2023 (Heinrich Hartmann)
An SRE veteran turns "reliable enough" into arithmetic: your error budget is the allowance you may spend on shipping fast, with burn-rate alerting that pages humans only when it matters.

Takeaways

  1. Metrics tell you when, logs tell you what, traces tell you why across services — invest in all three, correlate them.
  2. Dashboard by RED (services) and USE (resources); measure latency in percentiles, because averages hide the tail.
  3. Sample traces tail-based (keep all errors, few successes) and never put high-cardinality values in metric labels.
  4. Define SLIs/SLOs, compute the error budget in minutes, and alert on burn rate — each extra nine costs steeply more, and non-linearly.

Check your understanding

  1. What does the RED mnemonic stand for when monitoring a service?

    • Redundancy, Encryption, Durability
    • Retries, Escalations, Downtime
    • Reads, Events, Deploys
    • Rate, Errors, Duration
  2. Why are latency percentiles (p99) preferred over averages?

    • Averages hide tail latency — a good average can conceal a terrible experience for 1% of users
    • Averages are harder to compute
    • Percentiles require less storage
    • Averages cannot be graphed over time
  3. What is the key advantage of tail-based over head-based trace sampling?

    • It uses less memory in the collector
    • It requires no collector infrastructure
    • It keeps 100% of errors and slow requests instead of a random subset, so the traces you actually need exist
    • It samples before the request starts, reducing overhead
  4. What is a 'cardinality explosion' in metrics?

    • Too many dashboards being created
    • A high-cardinality label (like user_id) creates one time series per value, exploding memory and storage
    • Metrics being sampled too aggressively
    • Alerts firing in a cascading pattern
  5. A 99.9% SLO over 30 days allows roughly how much downtime (the error budget)?

    • About 7 hours per month
    • About 4 minutes per month
    • About 43 minutes per month
    • Zero — 99.9% means no downtime

Go deeper

Want to keep pulling this thread? These talks and tutorials go further than we did here:

Sources & further reading