- From Zero to Millions: Scaling a Web Service
The classic journey from one server to millions of users: vertical and horizontal scaling, load balancers, read replicas, CDNs, and sharding — with the napkin math to back it up.
- SOLID: Object-Oriented Design That Ages Well
Five principles for code that survives contact with the future: one job per class, extend without rewriting, substitutes that keep their promises, small interfaces, and dependencies that point at abstractions — told through a home kitchen.
- From MVC to React: How Frontend Architecture Evolved
How forty years of UI architecture fits together: the OOP ideas underneath, Smalltalk's Model-View-Controller, its MVP and MVVM descendants, the single-page app, and React's component model — told through a diner.
- The CAP Theorem
Why a distributed data store can only pick two of consistency, availability, and partition tolerance — and what that trade-off means in practice.
- Load Balancers
How traffic gets distributed across a fleet of servers, the layers at which it happens, and the algorithms used to pick a target.
- NoSQL in Plain English: Picking the Right Database
Why one database shape doesn’t fit every job: the four NoSQL families — key-value, document, wide-column, and graph — what each is good at, and how to choose between them and good old relational.
- Data Partitioning (Sharding)
Splitting a dataset across multiple machines so no single node has to hold — or serve — all of it.
- Design a URL Shortener
Our first case-study lesson: apply hashing, sharding, caching, and rate limiting to a real interview classic.
- Idempotency: Making Retries Safe
Networks fail, so we retry — but a retry can double-charge a credit card. How idempotency keys and idempotent operations make retries safe.
- AuthN & AuthZ: Who Are You, What Can You Do?
From passwords to JWTs to OAuth2 — how systems verify who you are (authentication) and decide what you are allowed to do (authorization).
- Caching Strategies
Cache-aside, read-through, write-through, and write-behind: where caches live, how data gets in and out, and how to survive eviction, TTLs, and stampedes.
- Content Delivery Networks
How CDNs put your content on every continent: edge PoPs, cache hierarchies, TTLs and purging, origin shields, and accelerating dynamic content.
- Cloud-Native Architecture: Building for Someone Else's Computer
Regions and availability zones, managed services vs self-hosting, Kubernetes essentials, serverless cold starts, the 12-factor rules, and the bill at the end — how to build on someone else's computer without getting burned.
- Message Queues
Decoupling producers from consumers with queues: delivery guarantees, ordering, backpressure, and how Kafka, RabbitMQ, and SQS differ.
- API Gateways
The front door of your system: routing, authentication, rate limiting, and request shaping — and how gateways differ from plain reverse proxies.
- BFF: One API Per Client
Why one generic API fails web, mobile, and partners differently — and how a backend-for-frontend per client type fixes payload size, chattiness, and release coupling, plus the sprawl pitfalls to avoid.
- Replication & Consistency Models
How copies of your data stay in sync (or don’t): synchronous vs. asynchronous replication, leader-based and leaderless designs, quorums, and the consistency models clients actually observe.
- Rate Limiting
Keeping abusive and accidental traffic from ruining everyone’s day: token bucket, leaky bucket, sliding windows, and distributed rate limiting with Redis.
- Microservices vs. Monolith
The most argued-over trade-off in backend architecture: when splitting services pays off, when a modular monolith wins, and how to avoid the distributed monolith.
- Microservices Patterns & Pitfalls: Avoiding the Distributed Monolith
The loopholes that turn microservices into a distributed monolith: the independent-deployability test, chatty services and latency budgets, the shared-database cardinal sin, temporal coupling — and when to start with a modular monolith instead.
- Resilience Patterns: Breakers, Bulkheads & Backoff
How to stop one slow dependency from taking down everything: timeouts, retries with backoff and jitter, circuit breakers, bulkheads, and graceful degradation.
- Stream vs. Batch: Processing Data at Scale
Event streams vs batch jobs, lambda vs kappa architectures, and how to choose the right one when the data never stops arriving.
- Consensus with Raft
How a cluster of unreliable machines agrees on one truth: Raft leader election, log replication, commit rules, and the failure scenarios that keep it honest.
- Sharding Strategies
Splitting a dataset across machines without regret: key-based, range-based, and directory-based sharding, consistent hashing with virtual nodes, and the true cost of resharding.
- Distributed Transactions
Atomicity across machines: two-phase commit, why coordinators block, saga patterns with compensating actions, and the outbox pattern for reliable messaging.
- Observability for Distributed Systems
Metrics, logs, and traces that actually answer questions: RED and USE, sampling math, cardinality explosions, and SLIs, SLOs, and error budgets.
- Domain-Driven Design: Let Events Tell the Story
Model the business the way the business talks: ubiquitous language, bounded contexts, aggregates, and domain events — plus event storming, the transactional outbox, and a peek at event sourcing.
- Anti-Corruption Layer: Protecting Your Domain
The translation boundary between your clean domain and the messy outside world: adapters and translators in DDD, a normalization layer between your React SPA and legacy APIs, and the strangler-fig migration that retires the legacy for good.
- Design Instagram: A Photo-Sharing System
The capstone case study: upload pipelines, sharded metadata, CDNs, and the fan-out problem behind a feed that serves a billion photos.
- Designing AI Systems: RAG, Agents & Inference at Scale
The systems engineering underneath the magic trick: how a request flows from prompt to answer, RAG vs fine-tuning, agent loops, vector databases, the economics of inference (batching, KV-cache, quantization, caching), and why evals replace dashboards.
- Agentic AI Patterns: Why Multi-Agent Systems Fail
Multi-agent demos feel magical until production: error cascades, retry loops, exploding context bills, and prompt injection smuggled in through tools. The agent loop, orchestration patterns, the failure math, and the guardrails that keep a team of agents honest.
- Event Sourcing & CQRS: Your Database Is a Diary
Stop storing the current state and start storing the history: the append-only event log as source of truth, snapshots and projections for speed, and CQRS to split reads from writes — with the replay math, ordering guarantees, and failure modes that decide whether the pattern earns its keep.
- Designing Chat at Scale: Messaging Like WhatsApp
The realtime interview classic: persistent connections, per-conversation ordering, delivery receipts, presence, and group fan-out at WhatsApp scale — and what end-to-end encryption takes away from the server.
- GraphQL vs REST: Choosing Your API Shape
REST hands every client the same fixed menu; GraphQL lets each client order exactly what fits on its plate. How the two models differ on fetching, caching, versioning, and security — and how to keep a flexible query language from becoming a denial-of-service cannon.
- Multi-Region Active-Active: Surviving the Datacenter Fire
Your whole region goes dark at 2 a.m. — now what? RPO and RTO with real numbers, active-passive vs active-active, replication lag, write conflicts and how to survive them, DNS failover, data borders, and why you should set fire to your own plans on purpose.
- Zero Trust: Never Trust, Always Verify
The castle-and-moat network is dead: why modern systems authenticate and authorize every single request, how workload identity and mutual TLS make it work, and what it costs when the front desk goes down.