Why this matters
"Is the system healthy?" is a surprisingly hard question when "the system" is 40 services, 3 regions, and a queue with opinions. Observability is the discipline of making systems answer questions about themselves — especially the questions you didn't think to ask. Monitoring tells you when something is wrong; observability helps you figure out why. At 3 AM, the difference between a dashboard that says "error rate spiked" and a trace that says "the spike is exactly these requests hitting exactly this shard" is the difference between a 10-minute fix and a 4-hour war room.
Analogy: monitoring is the check-engine light — it tells you something's wrong, in one bit. Observability is the full diagnostic port a mechanic plugs into: live sensor readings, freeze-frame data from when the fault occurred, the ability to ask "show me fuel pressure during the last misfire." You need the light to know to look; you need the diagnostics to actually fix it.
Video: Charity Majors — Observability and the Glorious Future — Chariot Solutions
The Honeycomb co-founder who defined modern observability on why distributed systems broke old monitoring — like navigating rapids with a map drawn for a canal.
The three pillars (and why traces win arguments)
- Metrics — numbers over time: request rate, error rate, latency histograms, CPU, queue depth. Cheap, aggregatable, and the backbone of alerting (Prometheus, Datadog).
- Logs — discrete events with context: "order 8472 failed: payment declined." Indispensable for the specific incident, expensive to retain at scale.
- Traces — the journey of one request across services, as a tree of spans with timings. This is the only pillar that shows causality across service boundaries (OpenTelemetry, Jaeger, Tempo).
flowchart TD
R["Request: GET /checkout<br/>total 412ms"] --> A["gateway: 8ms"]
R --> B["orders svc: 380ms"]
B --> C["inventory check: 12ms"]
B --> D["payments svc: 350ms"]
D --> E["fraud check: 340ms ⚠️"]
R --> F["notify svc: 24ms"]
One look at that trace and you know the fraud check owns the latency — no guessing, no SSHing into four boxes. Logs would show you four separate "slow request" lines; the trace shows you the relationship. In microservices, traces are the closest thing to a stack trace you'll ever get.
Video: What is Observability? | Grafana for Beginners Ep. 1 — Grafana
Walks through metrics, logs, and traces — and shows why traces settle debugging arguments.
RED and USE: what to measure
Two mnemonics keep dashboards focused:
- RED (for services, from Tom Wilkie): Rate (requests/sec), Errors (failed/sec), Duration (latency distribution). If your service dashboard doesn't lead with these three, it's decoration.
- USE (for resources, from Brendan Gregg): Utilization, Saturation, Errors — for every CPU, disk, and network interface. Saturation (queueing) is the one people forget: a disk at 70% utilization with a deep queue is slower than one at 90% with none.
And measure latency as percentiles, not averages. An average latency of 50 ms can hide 1% of users waiting 5 seconds — the average lies because it lets the happy majority vote down the miserable minority. Track p50, p95, p99 (and p99.9 if you're fancy); alert on the tail, because the tail is where your angriest users live.
Video: The RED Method: How to Instrument Your Services — GrafanaCon EU (Tom Wilkie)
Grafana's Tom Wilkie on the RED method: USE checks the engine, RED asks the passengers how the ride feels — together they tell you what to measure.
Sampling: the math of seeing enough
You cannot keep every trace. A service doing 10k requests/sec emitting 2 KB spans per request generates ~1.7 TB of trace data per day — per service. So you sample:
- Head-based sampling — decide at request start (keep 1% randomly). Simple, but the 1% you keep is random — the one failed checkout in a million successes is almost certainly not in your sample. Precisely when you need traces most, you don't have them.
- Tail-based sampling — buffer spans, decide after seeing the outcome: keep 100% of errors and slow requests, 1% of the boring successes. Needs a collector (OpenTelemetry Collector) with memory to hold the buffer, but it keeps exactly the traces you'll actually open.
The back-of-envelope: if your error rate is 0.1% and you want every error trace, tail-based sampling at "all errors + 1% of successes" stores ~1.1% of traffic — a 90× cost reduction over keeping everything, while retaining 100% of the interesting cases. Do the arithmetic for your own traffic; the answer is always "sample, but sample smart."
Video: How to find failures without drowning in tracing data — The New Stack
Trace-sampling strategies that catch failures without drowning your budget in span data.
Cardinality: the bill you didn't expect
Metrics systems index every unique label combination as a separate time series.
http_requests_total{path="/checkout", status="500"} is one series. Now add
user_id as a label with a million users: you've created millions of series, and your
Prometheus is having a very bad day — slow queries, exploding memory, and a storage
bill that arrives like a ransom note.
Rules of thumb: labels should be low-cardinality (region, service, status code, route
template like /users/:id — never the raw path /users/8472, never user IDs, never
request IDs). High-cardinality data belongs in logs and traces, which are designed for
it. The cardinality explosion is the #1 self-inflicted observability outage; the second is
alerting on the wrong thing (see below).
Interactive diagram: StepThrough (loads in the app)
Video: Containing Your Cardinality — Prometheus Monitoring
Prometheus labels are filing tags — handy until every request gets its own tag and the cabinet explodes; a PromCon talk on when cardinality helps and how to tame it.
SLIs, SLOs, and error budgets: alerting like an adult
Raw metrics don't tell you what's acceptable. The SRE vocabulary fixes that:
- SLI (Service Level Indicator) — what you measure: "99th-percentile latency" or "fraction of successful requests."
- SLO (Objective) — the target: "99.9% of requests succeed over 30 days."
- Error budget — the allowed failure: 0.1% of 30 days ≈ 43 minutes of downtime per month. That budget is a feature: it's permission to deploy, experiment, and break things — until it's spent.
The numbers make trade-offs concrete. 99.9% ("three nines") = 43 min/month downtime; 99.99% = 4.3 min/month; 99.999% = 26 sec/month. The engineering cost of each additional nine grows steeply and non-linearly while the user-visible difference shrinks. This is why "how many nines" is a business negotiation, not an engineering aspiration — and why alerting should fire on burn rate (how fast you're spending the budget) rather than on every blip. A 5-minute spike that consumes ~12% of the monthly budget is not a page; a sustained burn that eats 20% in an hour is.
Video: Mathematics of SLOs — SRECon EMEA 2023 (Heinrich Hartmann)
An SRE veteran turns "reliable enough" into arithmetic: your error budget is the allowance you may spend on shipping fast, with burn-rate alerting that pages humans only when it matters.
Takeaways
- Metrics tell you when, logs tell you what, traces tell you why across services — invest in all three, correlate them.
- Dashboard by RED (services) and USE (resources); measure latency in percentiles, because averages hide the tail.
- Sample traces tail-based (keep all errors, few successes) and never put high-cardinality values in metric labels.
- Define SLIs/SLOs, compute the error budget in minutes, and alert on burn rate — each extra nine costs steeply more, and non-linearly.
Check your understanding
What does the RED mnemonic stand for when monitoring a service?
- Redundancy, Encryption, Durability
- Retries, Escalations, Downtime
- Reads, Events, Deploys
- Rate, Errors, Duration
Why are latency percentiles (p99) preferred over averages?
- Averages hide tail latency — a good average can conceal a terrible experience for 1% of users
- Averages are harder to compute
- Percentiles require less storage
- Averages cannot be graphed over time
What is the key advantage of tail-based over head-based trace sampling?
- It uses less memory in the collector
- It requires no collector infrastructure
- It keeps 100% of errors and slow requests instead of a random subset, so the traces you actually need exist
- It samples before the request starts, reducing overhead
What is a 'cardinality explosion' in metrics?
- Too many dashboards being created
- A high-cardinality label (like user_id) creates one time series per value, exploding memory and storage
- Metrics being sampled too aggressively
- Alerts firing in a cascading pattern
A 99.9% SLO over 30 days allows roughly how much downtime (the error budget)?
- About 7 hours per month
- About 4 minutes per month
- About 43 minutes per month
- Zero — 99.9% means no downtime
Go deeper
Want to keep pulling this thread? These talks and tutorials go further than we did here:
- The RED Method: How To Instrument Your Services — Tom Wilkie, KubeCon + CloudNativeCon Europe 2019 (~30 min). The original RED-method talk behind the lesson's RED dashboards.
- Three Pillars, Zero Answers: We Need to Rethink Observability — Ben Sigelman, KubeCon NA 2018 (~35 min). Beyond metrics/logs/traces: cardinality limits and tail-based sampling.
- SRE - Using Error Budgets to Prioritize Work — Nathen Harvey, All The Talks 2020 (~30 min). SLIs, SLOs, and error budgets as a prioritization framework.
- Learn Kubernetes in 6 Hours – Full Course with Real-World Project — freeCodeCamp.org (~6h). Skip to the full-stack observability chapter (Prometheus + Grafana) to see metrics, dashboards, and alerts wired into a real cluster.
Sources & further reading
- Betsy Beyer et al., Site Reliability Engineering (Google / O'Reilly, 2016), Ch. 4 ("Service Level Objectives") — SLIs, SLOs, and error budgets; free online at sre.google.
- Brendan Gregg, "The USE Method" — utilization/saturation/errors for resource analysis.
- OpenTelemetry documentation — traces, spans, and the Collector for tail-based sampling.
- Cindy Sridharan, "Distributed Systems Observability" (O'Reilly, 2018) — the metrics/logs/traces synthesis for microservices.