system design

How to talk about CAP and consistency without sounding junior

Ask a candidate about the CAP theorem and most will recite “consistency, availability, partition tolerance, pick two.” That answer is memorized, and any interviewer who has run a few system design loops knows it on contact. The follow-up is where the round is scored: which two, under what conditions, and what the choice costs you the rest of the time.

Start by getting one thing straight, because it reframes the whole question. You don’t get to drop partition tolerance. A partition is a network failure that splits your nodes into groups that can’t reach each other, and in any system where components talk over a network, that will happen. A switch reboots, a cross-region link saturates, someone trips over a cable in the wrong rack. You can’t design it away. So the real choice CAP describes is narrower than the slogan: while a partition is happening, do you keep serving requests and risk returning stale or conflicting data, or do you refuse requests on the minority side to keep everyone’s view identical? That is the only fork. The rest of the time there is no partition and CAP has nothing to say.

The C in CAP is not the C in ACID

This is the trap interviewers set on purpose, and it catches people who read the headline and not the mechanism. CAP’s consistency means linearizability: every read sees the most recent write, as if there were a single copy of the data and every operation happened one at a time in a global order. ACID’s consistency means a transaction moves the database from one valid state to another, respecting constraints, foreign keys, and triggers. They share a word and nothing else. A candidate who blends the two tells the interviewer exactly how deep their understanding goes.

Availability has a precise definition too, and it is stricter than “the site loads.” An available system, in CAP terms, guarantees that every request to a non-failing node gets a non-error response, even during a partition. A node that answers “I can’t reach my peers, retry later” is not available in this sense, even though the process is healthy and the CPU is idle. Pin these two definitions down out loud and you have already cleared the bar that most candidates trip on.

PACELC is the framework interviewers actually want

CAP only describes the rare case, the partition. Daniel Abadi’s PACELC extension covers the common case, which is everything else. Read it as: if there is a Partition, choose between Availability and Consistency; Else, in normal operation, choose between Latency and Consistency. The second clause is the one that governs most of your uptime, because partitions are rare and every single read still has to decide whether to wait for a quorum of replicas to agree or answer fast from whatever copy is closest.

Classifying a database into one of the four corners is a strong move, because it forces you to be specific where hand-waving is tempting. Here is where the systems that come up most often land.

Database Behavior during a network partition Behavior in normal operation PACELC class What that means in practice
DynamoDB (default reads) Stays available, may serve stale data Answers from the nearest replica for low latency PA/EL Reads can lag writes; a strongly consistent read costs about 2x the read capacity
Cassandra Stays available, keeps accepting writes Answers at consistency level ONE for low latency PA/EL Consistency is set per query; QUORUM on read and write buys strong reads at higher latency
MongoDB (default replica set) Majority side stays writable; isolated secondaries may serve stale reads Primary reads reflect the latest write PA/EC An isolated primary steps down after an election; tune read concern for weaker or stronger reads
Google Spanner Minority replicas stop serving to preserve one global order Waits out clock uncertainty for consistency PC/EC Globally linearizable using synchronized clocks; you pay commit-wait latency on writes
HBase A region goes unavailable if its region server is cut off Routes each key to a single serving node for consistency PC/EC Strong reads by design; that single region server is the availability chokepoint

The empty corner in that table is PC/EL, a system that sacrifices availability under partition but favors latency the rest of the time. Yahoo’s PNUTS is the textbook example, and the fact that it is obscure tells you something: most designs that care enough to give up availability during a partition also care enough to pay for consistency when the network is fine. If you can name why PC/EL is rare, you have shown the interviewer you understand the framework and not just the acronym.

Quorum math is the mechanism behind “it depends how you configure it”

The reason so many systems now answer questions with “depends on the settings” is quorums. With N replicas, a write acknowledged by W of them and a read that queries R of them will always see the latest write as long as R + W is greater than N. That overlap guarantees at least one replica in your read set holds the newest value. Set N=3, W=2, R=2 and you have strong reads while tolerating one node failure. Drop R to 1 and reads get faster and cheaper but can miss a write that only reached two of the three nodes.

-- Cassandra: strong reads, paid for in latency
CONSISTENCY QUORUM;      -- with N=3, two replicas must respond
INSERT INTO orders (id, total) VALUES (7, 42.00);
SELECT total FROM orders WHERE id = 7;

DynamoDB exposes the same trade-off behind a single flag. A normal read is eventually consistent and cheap; set ConsistentRead to true and you get a linearizable read that consumes roughly twice the read capacity. The interviewer wants to hear that you know the knob is there and what pulling it costs, not that you would blanket-enable strong consistency across every table and eat the bill.

Consistency is a spectrum, not a switch

“Eventually consistent” gets used as if it named one behavior. It doesn’t, and a senior answer names the level. Between linearizable and eventual there are several useful stops. Causal consistency preserves cause and effect: if you post a comment and then reply to it, no reader sees the reply before the comment, though unrelated writes can still land in any order. Read-your-writes guarantees you see your own updates even when other users don’t yet, which is why your own post shows up the instant you hit send while a friend’s takes a beat to propagate. Monotonic reads promise that time never runs backward for you, so a value you already read won’t revert to an older one on your next request.

These levels matter because they map cleanly to product requirements. A bank ledger wants linearizability, no negotiation. A shopping cart is fine with causal or read-your-writes: the user needs to see items they just added, and a few hundred milliseconds of disagreement between replicas hurts nobody. A like count on a video can be eventually consistent and no one files a bug. Naming the weakest model that still satisfies the requirement is the winning move, because weaker models are faster and cheaper, and choosing the right one for each piece of data is the actual skill the question is probing.

What “what happens during a partition?” is really asking

A common way this surfaces: you have sketched a design with replicas in two regions, and the interviewer asks what happens when the link between them drops. The weak answer describes the failure and stops. The strong answer picks a side and defends it with the specific data on the table. Something like this. “The writes here are user session state, so I keep both regions accepting writes through the partition and reconcile with last-writer-wins when the link heals, because a lost session tweak is cheaper than an outage. If this were account balances I would flip it, route every write through one region, and let the other reject writes until the partition clears, because a double-spend is far worse than an error message the user can retry.”

That answer does three things the raw letters can’t. It ties the choice to the specific data, it states plainly what is being given up, and it accounts for the recovery path rather than freezing at the moment of failure. Interviewers are listening for whether you treat consistency as a per-feature decision instead of one global setting, because that is how these systems are actually run in production.

A few phrasings that come up, in rough order of frequency:

  • “Can you build a CA system? Why or why not?”
  • “Your read replica is lagging the primary. Is that a CAP problem or something else?”
  • “Where does DynamoDB sit on CAP, and what changes when you ask for a strongly consistent read?”
  • “Design something where availability beats consistency, then flip the requirement and redesign it.”

The CA question is the one people fumble most, so it is worth having a crisp answer ready. A single-node database is CA in a trivial sense: with no network between replicas, there is no partition to tolerate. The moment you replicate across machines, CA leaves the menu, and any tool that claims all three at once is quietly redefining one of the words, usually availability. Knowing that, and being able to point at which word got bent, is what separates someone who memorized a triangle from someone who has been paged at 3 a.m. because the triangle turned out to be real.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top