Almost every system design question is assembled from the same thirty-odd parts. Once you can draw them from memory and name where each one breaks, you stop reinventing the architecture on the whiteboard and start arguing about the pieces that decide the grade. This is the grid I keep in my head walking into a design round, grouped the way a request actually moves through a system.
One warning before the tables. Knowing the names is the floor, not the bar. In 2026 loops at the bigger companies, naming a load balancer and three AWS services gets you nowhere; the score comes from picking the right block for the constraint in front of you and being specific about how it fails. The tables give you the block and the gotcha. The interview is you connecting them under pressure.
At the edge, before a request hits your code
The first few hops fix most of the latency you will quote later, and they are where a design quietly becomes a single point of failure. Round-robin load balancing is the default people reach for, but least-connections handles uneven request durations better, and consistent-hash routing is what you want when a client should stick to the same cache node. Say which one and why; that sentence alone separates candidates.
| Block | What it does | Where it bites |
|---|---|---|
| DNS | Resolves a hostname to an IP, often to the nearest region via GeoDNS | TTLs make changes propagate slowly, so it is a poor tool for fast failover |
| L4 / L7 load balancer | Spreads TCP connections (L4) or HTTP requests (L7) across a server pool | Health checks decide everything; a bad check black-holes live traffic |
| Reverse proxy / API gateway | Terminates TLS, routes, authenticates and rate-limits in one place | Turns into a chokepoint and a graveyard for stale routing config |
| CDN | Caches static assets and video segments at edge POPs near users | Global invalidation is slow; version asset URLs instead of purging |
| Rate limiter | Caps requests per client to protect everything downstream | One Redis counter per limit is a hot key; shard it or buffer locally |
Where the data actually lives
The storage choice is usually the whole interview wearing a trench coat. Pick wrong and every later decision fights you. My default is a relational database until a real number forces me off it: a write rate a single primary cannot take, a blob too big for a row, a query a B-tree cannot serve. Reach for the specialized stores when the access pattern demands it, not because the name sounds impressive.
| Store | Good at | Where it bites |
|---|---|---|
| Relational (Postgres, MySQL) | ACID, joins, secondary indexes, sane defaults | One primary caps write throughput; you will shard eventually |
| Key-value (Redis, DynamoDB) | O(1) get and put by key, trivial to shard | No joins; you commit to your access patterns up front |
| Document (MongoDB) | Flexible schema, nested records read in one shot | Multi-document transactions are awkward; denormalization rots over time |
| Wide-column (Cassandra, ScyllaDB) | Huge write volume, tunable consistency, no single primary | Every query pattern has to be known at table-design time |
| Search index (Elasticsearch, OpenSearch) | Inverted index, full-text, faceted filters | Not a source of truth; it lags and sheds writes under load |
| Object store (S3, GCS) | Cheap, durable blobs from kilobytes to terabytes | High per-object latency; the wrong choice for small hot reads |
| Time-series (Prometheus, InfluxDB) | Append-heavy metrics, downsampling, retention windows | Label cardinality blowups will take it down |
| Vector DB (pgvector, Pinecone) | Approximate nearest-neighbor search over embeddings | Recall trades against latency; HNSW index builds are expensive |
| Graph (Neo4j) | Cheap multi-hop traversals over relationships | Sharding a graph is genuinely hard; many graph problems fit in Postgres |
The vector database is the block that moved from exotic to expected. Retrieval-augmented generation and recommendation questions now assume you can talk about embedding dimensions, an HNSW or IVF index, and the recall-versus-latency dial. You do not always need a dedicated system; pgvector inside the Postgres you already run is the right answer more often than the vendors would like.
Turning one database into many
Scaling reads is replication. Scaling writes is sharding. They are different problems and interviewers watch for candidates who blur them. Replication gives you read capacity and a failover target, at the cost of stale reads from lagging followers. Sharding gives you write capacity, at the cost of cross-shard queries and the misery of rebalancing around a key you chose badly.
| Block | What it does | Where it bites |
|---|---|---|
| Leader-follower replication | Streams writes from a primary out to read replicas | Replication lag serves stale data; read-your-writes needs handling |
| Sharding / partitioning | Splits data across nodes by a partition key | Cross-shard joins and resharding are the cost; the key choice is near-permanent |
| Consistent hashing | Maps keys to nodes on a ring so adding one moves few keys | Skew still creates hot shards; virtual nodes even it out |
| Quorum reads/writes | Requires R + W > N overlapping nodes for consistency | Higher quorum means higher latency; it is a dial, not a default |
Caching, and the four ways it goes wrong
Caching is easy to add to a diagram and hard to run in production. The interesting questions are never whether to cache but invalidation, eviction, and what happens the instant a popular key expires. That last one is the cache stampede: a hot key drops out, a thousand requests miss at once, and they all hammer the database together. Request coalescing and jittered TTLs are the standard defenses, and naming them before you are asked reads as production scars rather than book knowledge.
| Block | What it does | Where it bites |
|---|---|---|
| In-memory cache (Redis, Memcached) | Sub-millisecond reads for hot keys and sessions | Memory is finite; the eviction policy quietly sets your hit rate |
| Cache-aside | App checks the cache, falls back to the DB, fills on a miss | Data stays stale until the TTL expires; still the right default |
| Write-through / write-back | Writes land in the cache and the DB together, or the DB later | Write-back can lose data on a crash; write-through adds write latency |
| Bloom filter | Cheap “definitely not here” check before an expensive lookup | False positives only; sizing the bit array is the tradeoff |
Doing work later
The moment a request kicks off work the user does not need to wait for, you reach for asynchronous processing. The distinction interviewers probe is a queue versus a log. A queue (SQS, RabbitMQ) hands each message to one consumer and forgets it. A log (Kafka, Pulsar) keeps an ordered, replayable record that many consumers read at their own pace. Picking the log when you need replay or fan-out, and the queue when you only need work taken off the hot path, is most of the answer.
| Block | What it does | Where it bites |
|---|---|---|
| Message queue (SQS, RabbitMQ) | Buffers work and decouples producer from consumer | At-least-once delivery means consumers have to be idempotent |
| Event log (Kafka, Pulsar) | Durable, replayable stream, ordered within a partition | Ordering holds only per partition; consumer lag is the metric to watch |
| Stream processor (Flink, Spark) | Windowed aggregation and joins over a live stream | Exactly-once needs checkpoints; unbounded state needs TTLs |
| Worker pool (Celery, Sidekiq) | Pulls jobs off a queue and runs them off the request path | Retries plus side effects produce duplicates without dedup keys |
Getting independent servers to agree
This last group is where senior candidates pull ahead, because most people skip it. The second your system has more than one of anything, you need a story for how the copies agree on who is in charge and what the current truth is. You rarely implement Raft on a whiteboard, but you should know you are standing on it when you say “etcd” or “leader election,” and you should know that a distributed lock without fencing tokens is a split-brain waiting to happen.
| Block | What it does | Where it bites |
|---|---|---|
| Consensus (Raft, Paxos) | Agrees on one value across nodes despite failures | Needs a majority; an even node count buys no extra safety |
| Coordination service (etcd, ZooKeeper) | Leader election, service discovery and config on top of consensus | Everything depends on it, so its outage becomes your outage |
| Distributed lock | Grants one holder at a time across machines | A partition can split the lock; fencing tokens stop the double write |
| Idempotency key | De-dupes retried writes by a client-supplied id | Only works if you store and check the key before doing the work |
| Saga | Sequences a multi-step transaction with compensating undos | There is no rollback; you write and test the undo path by hand |
Keep the grid in your head as a sketch order rather than a checklist to recite: traffic in at the edge, data at rest, copies of that data, hot copies in cache, work pushed off to the side, and the coordination holding it together. What has shifted in recent loops is that interviewers now ask what each block costs to run and who gets paged when it fails. A candidate who can place these thirty pieces and then talk about the on-call burden of the Kafka cluster they just drew is interviewing at a different level than one who knows only that the boxes have names.
Keep sharpening your system design:
