Low Level Design: System Design Interview Framework

Updated · techinterview.org

System design interviews test your ability to think at scale, communicate trade-offs, and structure ambiguity into actionable plans. A repeatable framework keeps you from forgetting critical steps under pressure. This guide walks through each stage in order, with the reasoning behind it.

Step 1: Clarify Requirements

Never start designing before you understand what you’re building. Spend the first five minutes asking explicit questions. Interviewers reward candidates who clarify rather than assume.

Functional Requirements

Functional requirements define what the system does. Ask: What are the core use cases? Who are the users? What actions can they perform? A URL shortener needs at minimum: create a short URL, redirect a short URL to the original, and optionally track click analytics.

Non-Functional Requirements

Non-functional requirements define how well the system does it. Always ask explicitly about:

  • Scale: Daily Active Users (DAU), Queries Per Second (QPS), and expected data volume.
  • Latency SLO, such as p99 read latency < 100ms or p50 write latency < 500ms.
  • Consistency: whether strong consistency is needed, or eventual consistency is acceptable.
  • Availability SLO, such as 99.9% (8.76h downtime/year) or 99.99% (52min).
  • Read/write ratio, since read-heavy (social feed) vs write-heavy (logging) changes the architecture significantly.
  • Geographic distribution: single region or multi-region, active-active or active-passive.
  • Client types, since mobile clients have different constraints than web clients (bandwidth, battery, offline support).

Step 2: Estimate Scale

Back-of-envelope calculations anchor the design. They determine whether you need a single database or a distributed cluster, a single server or a fleet. Use these standard formulas:

Storage per day  = DAU × avg_data_per_user_per_day
QPS (average)    = DAU × requests_per_user_per_day / 86400
QPS (peak)       = average_QPS × 2 to 3
Storage per year = daily_storage × 365

Example for a Twitter-like system with 100M DAU, 5 tweets/day, 140 bytes/tweet: daily storage = 100M × 5 × 140 = 70GB/day; QPS average = 100M × 50 requests / 86400 ≈ 58K QPS; peak ≈ 175K QPS. These numbers immediately tell you: you need horizontal scaling, sharding, and a caching layer.

Step 3: Define the API

Define the API before drawing components. The API is the contract between clients and the system. It also forces you to think about data shapes early.

  • Choose a protocol: REST for external/browser clients, gRPC for internal service-to-service (lower latency, binary encoding, streaming).
  • Define each endpoint with its HTTP method, path, request parameters, and response schema.
  • Use cursor-based pagination (not offset) for large result sets; offset pagination breaks under concurrent inserts.
  • State the rate limiting strategy: token bucket, leaky bucket, or sliding window counter.

Example for a URL shortener:

POST /urls          { long_url: string } → { short_code: string }
GET  /:short_code   → 301 redirect to long_url
GET  /urls/:id/stats → { clicks: int, created_at: timestamp }

Step 4: High-Level Design

Draw the main components and data flow for the two or three core use cases. Do not go deep yet. Breadth first. Standard components to consider:

  • Clients such as web, mobile, and third-party callers.
  • A load balancer or API gateway for TLS termination, routing, auth, and rate limiting.
  • API servers running stateless application logic, easy to scale horizontally.
  • Message queues like Kafka or SQS for async processing, decoupling producers from consumers.
  • Databases for primary storage, SQL or NoSQL depending on access patterns.
  • A cache such as Redis or Memcached in front of the database for hot data.
  • A CDN for static assets, large media files, and geographically distributed reads.
  • Object storage such as S3 for blobs like photos, videos, and logs.

Walk through the data flow: "A write request comes in → hits the load balancer → API server validates and writes to primary DB → publishes event to Kafka → consumer processes async work." This narrative shows you understand how components interact.

Step 5: Storage Design

Storage is where most candidates spend too little time. Go deep here; the interviewer is evaluating whether you can make principled choices.

SQL vs NoSQL

Choose SQL (PostgreSQL, MySQL) when: you need ACID transactions, relationships between entities are complex, or the query patterns are unpredictable. Choose NoSQL (Cassandra, DynamoDB, MongoDB) when: you need to scale writes horizontally, the access pattern is known and narrow (key-value or time-series), or the schema will evolve rapidly.

Schema Design

Design the schema to match the primary access pattern. Denormalize for read-heavy workloads. Add indexes on columns you filter or sort by. State which indexes you’d add and why: "I’d add a composite index on (user_id, created_at DESC) to support paginated feed queries efficiently."

Sharding Strategy

If the data volume or QPS exceeds what a single node can handle, discuss sharding. Common strategies: shard by user_id (keeps a user’s data co-located), shard by hash (even distribution but cross-shard queries are expensive), shard by geography (reduces latency for regional data). Always discuss the hotspot problem: a celebrity user on one shard creates imbalance.

Caching Layer

Cache what is read frequently and changes infrequently. Use Redis for structured data (sorted sets for leaderboards, hashes for user sessions). Define the cache invalidation strategy: TTL-based expiry for tolerating stale data, write-through for strong consistency, cache-aside (lazy loading) for flexibility.

Step 6: Deep Dive

The interviewer will direct you to go deep on one component. This is where the interview is won or lost. Common deep dive areas:

  • Scaling the write path: to handle 100K writes/second, use write buffering, batching, WAL, and async replication.
  • Feed generation algorithm: choose a push model (fan-out on write, pre-compute timelines), a pull model (fan-out on read, compute at request time), or a hybrid.
  • Handling failures: what happens if a database node goes down? Discuss replication (primary-replica), automatic failover, circuit breakers, and retry with exponential backoff.
  • Consistency guarantees: how do you ensure two users don’t book the same hotel room? Discuss optimistic locking, pessimistic locking, distributed transactions, and the saga pattern.

Step 7: Discuss Trade-offs

Every architectural decision is a trade-off. Articulating them explicitly demonstrates senior-level thinking. Key trade-offs to be ready to discuss:

  • Pull vs push (feed delivery): push is fast at read time but expensive for celebrity users, while pull is cheap at write time but slow at read time.
  • Consistency vs availability (CAP theorem): during a network partition you must choose whether to return stale data or reject the request.
  • Latency vs throughput: batching increases throughput but adds latency, while processing one-at-a-time minimizes latency but reduces throughput.
  • Simplicity vs scale: a monolith is simpler to operate but harder to scale individual components, while microservices enable independent scaling but add operational complexity.
  • Normalization vs denormalization: normalized data is consistent but requires expensive joins, while denormalized data is fast to read but harder to keep consistent.

Common Pitfalls

These mistakes consistently lose candidates points:

  • Jumping to implementation: drawing database schemas before clarifying scale. The right schema for 1K users is wrong for 1B users.
  • Over-engineering: proposing a distributed system for a problem that a single PostgreSQL instance handles fine. Match the solution to the stated scale.
  • Ignoring failure modes: every component fails. If you don’t mention what happens when the cache goes down or the message queue falls behind, the interviewer notices.
  • Forgetting caching: almost every high-scale system needs a caching layer. Omitting it suggests inexperience with production systems.
  • Not discussing trade-offs: stating decisions without justification ("I’ll use NoSQL") is weak. Always follow with "because X, at the cost of Y."

Time Allocation

A 45-minute system design interview should be paced as follows:

PhaseTime
Requirements clarification5 minutes
Scale estimation5 minutes
High-level design10 minutes
Deep dive (interviewer-directed)25 minutes
Trade-offs and wrap-up5 minutes

If you’re still doing estimation at minute 15, you’ve lost time on the deep dive where the real evaluation happens. Practice pacing with a timer.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack
Scroll to Top