system design

What interviewers actually grade in a RAG design round

The prompt is almost always mundane. “Design a system that answers employee questions from our internal wiki.” You get 45 minutes, a shared doc, and an interviewer who has watched a hundred candidates reach for embeddings and stop there. The ones who pass treat retrieval as the hard part and generation as the easy part. That ordering is the whole game.

Retrieval-augmented generation fetches relevant text at query time and puts it in the model’s context, so the answer comes from your data instead of the model’s training. The interviewer already knows that definition. What they’re grading is whether you can point to where each stage loses quality and state the specific tradeoff behind every box you draw.

What the round actually looks like

In 2026 this shows up in two shapes. For an applied-AI or ML-platform role it’s a full 45-to-60-minute design round. For a general backend role it’s often a 20-minute segment inside a technical screen, where the interviewer just wants proof you’ve built one. The frontier labs and their applied teams run the long version; smaller startups tend to fold it into a take-home against a sample corpus, which is harder to fake because you have to make real retrieval numbers move.

The opening they want is scoping, not architecture. Spend the first few minutes on the corpus. How many documents, how big each one is, how often they change, and whether they’re clean text or PDFs full of tables and scanned pages. Ten thousand markdown pages and four million legal contracts are different systems. Drawing the same diagram for both is the fastest way to signal you’re pattern-matching.

Some questions that tend to land mid-round:

  • “The corpus is 50 million chunks. Does your index still fit in memory, and what happens when it doesn’t?”
  • “A user asks something whose answer spans three documents. What does your retriever return?”
  • “The model just cited a policy we deleted last week. Why did that happen, and how do you stop it?”

Chunking is where most answers quietly fail

Everyone draws the embedding box. Far fewer have an opinion about what goes into it. Chunk too small and a 200-token window holds a sentence with no surrounding context, so the retrieved passage is technically relevant and practically useless. Chunk too large and you dilute the embedding: a 2,000-token block covering five subtopics gets a fuzzy vector that matches everything weakly and nothing well. The sweet spot for most prose sits somewhere around 256 to 512 tokens with a little overlap, but the real answer is that it depends on the document, and saying so with reasons is what earns the point.

Approach How it splits Where it wins Where it hurts
Fixed-size Every N tokens with a small overlap Cheap, predictable, fast to build Cuts sentences and tables mid-thought
Structure-aware On headings, then paragraphs, then sentences Keeps whole ideas together Needs clean structure in the source
Semantic Groups sentences by embedding similarity Chunks follow topic shifts Slow, and costs extra embeddings to build
Parent-document Index small pieces, return the larger section they belong to Precise match plus full context More plumbing, bigger stored index

If you remember one thing here, make it the parent-document trick: index small chunks for precise matching, then feed the model the larger section they came from. It fixes the “relevant but context-free” failure without wrecking retrieval precision, and interviewers perk up when you bring it unprompted.

Dense retrieval, sparse retrieval, and why you want both

Pure vector search has a known blind spot. Ask about error code “TX-4092” and embedding similarity may rank a paragraph about general transaction failures above the one page that names the exact code, because the model never saw that token in training and its vector sits close to noise. Keyword search (BM25) nails exact matches and proper nouns but misses paraphrase. Running both and merging the results with reciprocal rank fusion covers each one’s weakness, and hybrid retrieval is close to table stakes for a strong answer now.

For the vector index itself, HNSW is the default worth naming. It gives roughly logarithmic search over millions of vectors and dominates production setups; the cost is memory, since the graph lives in RAM. When the interviewer pushes on 50 million or 500 million vectors, that’s your cue to talk about IVF or product quantization, which trade a little recall for a much smaller footprint, and about sharding the index across nodes. Bring up metadata filtering too, because half of real queries are scoped (“only docs from this team, published after January”) and a design that can’t filter on metadata falls apart on the first follow-up.

Reranking is the cheap win people forget

The retriever is built for speed, not judgment. A bi-encoder embeds the query and the documents separately, so it never compares them directly; it just measures vector distance. A cross-encoder reads the query and a candidate together and scores actual relevance, which is far more accurate and far too slow to run over a whole corpus. So you use both: pull the top 50 with the fast retriever, then rerank down to the top 5 with a cross-encoder like Cohere Rerank or bge-reranker-v2-m3. This two-stage shape adds maybe 50 to 200 milliseconds and reliably lifts answer quality. It’s the single biggest gain you can bolt onto a mediocre pipeline, and skipping it when the interviewer asks “how would you improve retrieval quality” is a missed layup.

Generation, and the parts that bite in production

Now you assemble a prompt: the retrieved chunks, the user’s question, and an instruction to answer only from the provided context and say “I don’t know” when the context doesn’t cover it. That last rule is what stops the system from confidently inventing a refund policy. Budget the context window carefully, because five chunks plus a long question plus system instructions can crowd out room for the answer, and cramming twenty chunks in makes the model worse, not better, as the one that matters gets buried in the middle where attention is weakest.

Citations matter more than candidates expect. If the answer links back to the source chunk, a wrong answer becomes debuggable and a right answer becomes trustworthy. Teams building enterprise assistants care about this a lot, because “the bot made something up and a customer acted on it” is the failure that gets a project shut down.

How you prove it works

The strongest signal you can send is that you know how to measure this thing, because most people can’t. Split the evaluation in two. Retrieval either found the right chunks or it didn’t, and you can check that with no model in the loop. Generation either used those chunks faithfully or it drifted. A framework like RAGAS packages the common metrics, but you should be able to say what each one catches on its own.

Metric What it catches
Context recall Whether retrieval pulled the chunks that hold the answer
Context precision How much of what you retrieved was actually relevant
Faithfulness Whether the answer is grounded in the chunks or invented
Answer relevancy Whether the answer addresses the question that was asked
MRR / hit@k How high the first correct chunk ranks
p95 latency, cost per query The numbers the product team actually tracks

Build a golden set of a few hundred real questions with known-good answers and known source chunks, then run it on every change so you can tell whether a new embedding model or a different chunk size helped or just felt better. In production you watch the same faithfulness and latency numbers live, sample low-confidence answers for review, and treat every thumbs-down as a fresh eval case.

Which brings back the deleted-policy question, the one that separates people who’ve run RAG from people who’ve read about it. The model cited a dead document because your vector index still had its embedding; nothing ever told the index the source was gone. The fix is boring, and that’s the point. Your ingestion pipeline needs to process deletes and updates, not inserts alone, and you need a clear story for how fast a wiki edit reaches the index. Answer that one well and the interviewer stops wondering whether you’ve actually built one of these.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top