system design

30 System Design Building Blocks You Can Sketch From Memory

Almost every system design question is assembled from the same thirty-odd parts. Once you can draw them from memory and name where each one breaks, you stop reinventing the architecture on the whiteboard and start arguing about the pieces that decide the grade. This is the grid I keep in my head walking into a design round, grouped the way a request actually moves through a system.

One warning before the tables. Knowing the names is the floor, not the bar. In 2026 loops at the bigger companies, naming a load balancer and three AWS services gets you nowhere; the score comes from picking the right block for the constraint in front of you and being specific about how it fails. The tables give you the block and the gotcha. The interview is you connecting them under pressure.

At the edge, before a request hits your code

The first few hops fix most of the latency you will quote later, and they are where a design quietly becomes a single point of failure. Round-robin load balancing is the default people reach for, but least-connections handles uneven request durations better, and consistent-hash routing is what you want when a client should stick to the same cache node. Say which one and why; that sentence alone separates candidates.

Block What it does Where it bites
DNS Resolves a hostname to an IP, often to the nearest region via GeoDNS TTLs make changes propagate slowly, so it is a poor tool for fast failover
L4 / L7 load balancer Spreads TCP connections (L4) or HTTP requests (L7) across a server pool Health checks decide everything; a bad check black-holes live traffic
Reverse proxy / API gateway Terminates TLS, routes, authenticates and rate-limits in one place Turns into a chokepoint and a graveyard for stale routing config
CDN Caches static assets and video segments at edge POPs near users Global invalidation is slow; version asset URLs instead of purging
Rate limiter Caps requests per client to protect everything downstream One Redis counter per limit is a hot key; shard it or buffer locally

Where the data actually lives

The storage choice is usually the whole interview wearing a trench coat. Pick wrong and every later decision fights you. My default is a relational database until a real number forces me off it: a write rate a single primary cannot take, a blob too big for a row, a query a B-tree cannot serve. Reach for the specialized stores when the access pattern demands it, not because the name sounds impressive.

Store Good at Where it bites
Relational (Postgres, MySQL) ACID, joins, secondary indexes, sane defaults One primary caps write throughput; you will shard eventually
Key-value (Redis, DynamoDB) O(1) get and put by key, trivial to shard No joins; you commit to your access patterns up front
Document (MongoDB) Flexible schema, nested records read in one shot Multi-document transactions are awkward; denormalization rots over time
Wide-column (Cassandra, ScyllaDB) Huge write volume, tunable consistency, no single primary Every query pattern has to be known at table-design time
Search index (Elasticsearch, OpenSearch) Inverted index, full-text, faceted filters Not a source of truth; it lags and sheds writes under load
Object store (S3, GCS) Cheap, durable blobs from kilobytes to terabytes High per-object latency; the wrong choice for small hot reads
Time-series (Prometheus, InfluxDB) Append-heavy metrics, downsampling, retention windows Label cardinality blowups will take it down
Vector DB (pgvector, Pinecone) Approximate nearest-neighbor search over embeddings Recall trades against latency; HNSW index builds are expensive
Graph (Neo4j) Cheap multi-hop traversals over relationships Sharding a graph is genuinely hard; many graph problems fit in Postgres

The vector database is the block that moved from exotic to expected. Retrieval-augmented generation and recommendation questions now assume you can talk about embedding dimensions, an HNSW or IVF index, and the recall-versus-latency dial. You do not always need a dedicated system; pgvector inside the Postgres you already run is the right answer more often than the vendors would like.

Turning one database into many

Scaling reads is replication. Scaling writes is sharding. They are different problems and interviewers watch for candidates who blur them. Replication gives you read capacity and a failover target, at the cost of stale reads from lagging followers. Sharding gives you write capacity, at the cost of cross-shard queries and the misery of rebalancing around a key you chose badly.

Block What it does Where it bites
Leader-follower replication Streams writes from a primary out to read replicas Replication lag serves stale data; read-your-writes needs handling
Sharding / partitioning Splits data across nodes by a partition key Cross-shard joins and resharding are the cost; the key choice is near-permanent
Consistent hashing Maps keys to nodes on a ring so adding one moves few keys Skew still creates hot shards; virtual nodes even it out
Quorum reads/writes Requires R + W > N overlapping nodes for consistency Higher quorum means higher latency; it is a dial, not a default

Caching, and the four ways it goes wrong

Caching is easy to add to a diagram and hard to run in production. The interesting questions are never whether to cache but invalidation, eviction, and what happens the instant a popular key expires. That last one is the cache stampede: a hot key drops out, a thousand requests miss at once, and they all hammer the database together. Request coalescing and jittered TTLs are the standard defenses, and naming them before you are asked reads as production scars rather than book knowledge.

Block What it does Where it bites
In-memory cache (Redis, Memcached) Sub-millisecond reads for hot keys and sessions Memory is finite; the eviction policy quietly sets your hit rate
Cache-aside App checks the cache, falls back to the DB, fills on a miss Data stays stale until the TTL expires; still the right default
Write-through / write-back Writes land in the cache and the DB together, or the DB later Write-back can lose data on a crash; write-through adds write latency
Bloom filter Cheap “definitely not here” check before an expensive lookup False positives only; sizing the bit array is the tradeoff

Doing work later

The moment a request kicks off work the user does not need to wait for, you reach for asynchronous processing. The distinction interviewers probe is a queue versus a log. A queue (SQS, RabbitMQ) hands each message to one consumer and forgets it. A log (Kafka, Pulsar) keeps an ordered, replayable record that many consumers read at their own pace. Picking the log when you need replay or fan-out, and the queue when you only need work taken off the hot path, is most of the answer.

Block What it does Where it bites
Message queue (SQS, RabbitMQ) Buffers work and decouples producer from consumer At-least-once delivery means consumers have to be idempotent
Event log (Kafka, Pulsar) Durable, replayable stream, ordered within a partition Ordering holds only per partition; consumer lag is the metric to watch
Stream processor (Flink, Spark) Windowed aggregation and joins over a live stream Exactly-once needs checkpoints; unbounded state needs TTLs
Worker pool (Celery, Sidekiq) Pulls jobs off a queue and runs them off the request path Retries plus side effects produce duplicates without dedup keys

Getting independent servers to agree

This last group is where senior candidates pull ahead, because most people skip it. The second your system has more than one of anything, you need a story for how the copies agree on who is in charge and what the current truth is. You rarely implement Raft on a whiteboard, but you should know you are standing on it when you say “etcd” or “leader election,” and you should know that a distributed lock without fencing tokens is a split-brain waiting to happen.

Block What it does Where it bites
Consensus (Raft, Paxos) Agrees on one value across nodes despite failures Needs a majority; an even node count buys no extra safety
Coordination service (etcd, ZooKeeper) Leader election, service discovery and config on top of consensus Everything depends on it, so its outage becomes your outage
Distributed lock Grants one holder at a time across machines A partition can split the lock; fencing tokens stop the double write
Idempotency key De-dupes retried writes by a client-supplied id Only works if you store and check the key before doing the work
Saga Sequences a multi-step transaction with compensating undos There is no rollback; you write and test the undo path by hand

Keep the grid in your head as a sketch order rather than a checklist to recite: traffic in at the edge, data at rest, copies of that data, hot copies in cache, work pushed off to the side, and the coordination holding it together. What has shifted in recent loops is that interviewers now ask what each block costs to run and who gets paged when it fails. A candidate who can place these thirty pieces and then talk about the on-call burden of the Kafka cluster they just drew is interviewing at a different level than one who knows only that the boxes have names.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top