system design

Why sagas beat two-phase commit when a write spans services

Almost every backend has the order-placement flow: charge the card, decrement inventory, create a shipment record. Three services, three separate databases. In a system design round the interviewer’s follow-up is nearly always the same one. Payment succeeds, the inventory write fails. Now what?

On a single Postgres box you would wrap all three writes in one transaction and let the database guarantee all-or-nothing. That option disappears the moment the data lives in three services owned by three teams. There is no shared transaction boundary across a payments database, an inventory database, and a shipping database, so you need a way to coordinate commit or rollback across systems that each fail on their own schedule. Two families of answers exist, and knowing where each one falls apart is the point of the question.

Two-phase commit, and why it stalls under load

Two-phase commit (2PC) does what the name says. A coordinator asks every participant to prepare: lock the rows, write the change to a durable log, and promise it can commit if asked. Each participant votes yes or no. If every vote is yes, the coordinator sends commit and the changes become visible. If any vote is no, it sends abort and everyone rolls back. The XA specification is the standard interface for this, and traditional relational databases plus a few message brokers implement it.

You get real atomicity, which is why 2PC still runs in bank ledgers and older ERP systems. The trouble shows up under load and under failure. Between prepare and commit, every participant holds its locks. If the coordinator pauses for two seconds, every row the transaction touched stays locked for two seconds, and contention piles up fast on hot rows like a popular product’s stock count. The blocking case is worse. If the coordinator crashes after some participants voted yes but before it broadcasts the decision, those participants sit holding locks with no idea whether to commit or abort, and they cannot safely pick either. The coordinator is a single point of failure that can freeze the flow.

Then there is the practical wall. Kafka, SQS, most HTTP APIs, and plenty of NoSQL stores do not speak XA at all. The instant one step is “publish an event” or “call a third-party payment API”, 2PC is off the table, because that participant cannot join the protocol. In a microservices stack that describes almost every transaction worth discussing. Reaching for 2PC across services in a 2026 interview reads as a warning sign rather than a strong answer. It is the right tool inside one database cluster and the wrong tool across service boundaries.

Sagas trade atomicity for a sequence you can undo

A saga drops the goal of one atomic commit and replaces it with a chain of local transactions, one per service, each committing on its own. If a later step fails, you run compensating transactions that semantically undo the earlier steps. The order flow becomes reserve payment, then reserve inventory, then create shipment. If shipment creation fails, the compensations run in reverse: release the inventory reservation, then void or refund the payment. Every service only runs a normal local ACID transaction against its own database, so nothing holds a lock across the network.

The part that trips people up: a compensation is not a rollback. Once payment captured the money, you cannot pretend it never happened. You issue a refund, a brand-new transaction that leaves its own audit trail. Some actions cannot be undone at all. There is no way to un-send a shipping-label email. The usual move is to order the steps so the irreversible one runs last, or to mark a pivot point after which the saga only moves forward and never compensates.

Choreography versus orchestration

There are two ways to wire the steps together. With choreography, each service publishes an event after its local commit, and the next service listens for that event and reacts. Payment publishes PaymentCaptured, inventory consumes it and publishes InventoryReserved, shipping consumes that. No central brain. It stays simple for three steps and turns into a mess at eight, because the flow lives implicitly across a dozen event handlers and nobody can point at where the saga actually is. Debugging a stuck order means grepping logs across services.

Orchestration puts one component in charge. The orchestrator sends a command to payment, waits for the reply, sends the next command to inventory, and owns the compensation logic when a step fails. The whole flow is readable in one place. The cost is that the orchestrator is now a service you build, deploy, and keep available. Teams reach for Temporal, AWS Step Functions, or a hand-rolled state machine backed by a database. Past a handful of steps, most interviewers expect you to prefer orchestration and to explain why.

The dual-write bug that sinks naive sagas

Here is the failure that separates people who have shipped a saga from people who have only read about one. A service commits its local transaction, then publishes an event to Kafka to kick off the next step. Those are two systems. If the process dies after the database commit but before the publish, the event is lost and the saga silently stalls. Publish first and let the commit fail, and you have announced something that never happened. Writing to a database and a broker is not atomic, and no clever retry ordering makes it atomic.

The answer interviewers want is the transactional outbox. Instead of publishing directly, the service writes the outgoing event into an outbox table inside the same local transaction as the business change. Now the change and the intent to publish commit or fail together. A separate relay reads new rows from the outbox and publishes them to the broker, marking each one sent. Change-data-capture tools like Debezium tail the database log to do this without polling. Delivery is now at-least-once, which means every downstream consumer has to be idempotent anyway, usually keyed on the event id or a business key so a replayed InventoryReserved does not decrement stock twice.

Picking the model out loud

What the interviewer scores is whether you can name the tradeoff and match it to the scenario, not whether you memorized the definitions.

Coordination approach Consistency guarantee How a mid-flow failure is handled Main weakness Reasonable fit
Two-phase commit (XA) Strong, atomic across all participants Coordinator broadcasts abort and every participant rolls back Holds locks during the prepare window, blocks if the coordinator dies, needs XA support in every participant Several databases inside one trusted cluster, modest throughput, a hard atomicity requirement
Saga with choreography Eventual, each step visible as it commits Each service runs its own compensating transaction Flow is implicit across event handlers and hard to trace past a few steps Short flows of two to four loosely coupled services
Saga with orchestration Eventual, each step visible as it commits A central orchestrator drives compensations in reverse order The orchestrator is another service to build and keep available Longer or frequently changing flows that need visibility and control
Transactional outbox Not a coordinator; it makes event publishing reliable The event survives a crash because it commits with the business write At-least-once delivery forces every consumer to be idempotent Any saga step that has to publish an event or command

The follow-ups that decide the round

Once you land on a saga, the interviewer starts probing the edges. The strongest signal you can send is treating failure and duplication as the normal path, not a rare exception.

  • “Payment captured, inventory reservation failed, and the refund compensation also fails. What now?” Retry the compensation with backoff, and if it keeps failing, park the saga in a dead-letter state for a human to resolve. Money problems get an alert, never a silent drop.
  • “Two OrderPlaced events arrive for the same order. Does the customer get charged twice?” Only if the consumers are not idempotent. Dedupe on the order id before running the local transaction.
  • “How does the customer see an order that is only half committed?” You expose a status of pending, confirmed, or failed, and let the UI reflect eventual consistency instead of pretending the write landed instantly.
  • “When would you actually use 2PC here?” If all three writes lived in the same database cluster and the business could not tolerate any pending window, a single transaction or 2PC inside that cluster beats a saga’s extra machinery.

The mistake that ends these rounds is reaching for a global lock or a distributed transaction because it feels safe. Across services it is neither safe nor available. Sagas are more code, and they force you to reason about compensation and idempotency for every step, which is the reasoning the interviewer is checking for. Open with “this write spans services, so I give up atomicity and design for eventual consistency”, and the rest of the answer builds itself.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top