The prompt is almost always mundane. “Design a system that answers employee questions from our internal wiki.” You get 45 minutes, a shared doc, and an interviewer who has watched a hundred candidates reach for embeddings and stop there. The ones who pass treat retrieval as the hard part and generation as the easy part. That ordering is the whole game.
Retrieval-augmented generation fetches relevant text at query time and puts it in the model’s context, so the answer comes from your data instead of the model’s training. The interviewer already knows that definition. What they’re grading is whether you can point to where each stage loses quality and state the specific tradeoff behind every box you draw.
What the round actually looks like
In 2026 this shows up in two shapes. For an applied-AI or ML-platform role it’s a full 45-to-60-minute design round. For a general backend role it’s often a 20-minute segment inside a technical screen, where the interviewer just wants proof you’ve built one. The frontier labs and their applied teams run the long version; smaller startups tend to fold it into a take-home against a sample corpus, which is harder to fake because you have to make real retrieval numbers move.
The opening they want is scoping, not architecture. Spend the first few minutes on the corpus. How many documents, how big each one is, how often they change, and whether they’re clean text or PDFs full of tables and scanned pages. Ten thousand markdown pages and four million legal contracts are different systems. Drawing the same diagram for both is the fastest way to signal you’re pattern-matching.
Some questions that tend to land mid-round:
- “The corpus is 50 million chunks. Does your index still fit in memory, and what happens when it doesn’t?”
- “A user asks something whose answer spans three documents. What does your retriever return?”
- “The model just cited a policy we deleted last week. Why did that happen, and how do you stop it?”
Chunking is where most answers quietly fail
Everyone draws the embedding box. Far fewer have an opinion about what goes into it. Chunk too small and a 200-token window holds a sentence with no surrounding context, so the retrieved passage is technically relevant and practically useless. Chunk too large and you dilute the embedding: a 2,000-token block covering five subtopics gets a fuzzy vector that matches everything weakly and nothing well. The sweet spot for most prose sits somewhere around 256 to 512 tokens with a little overlap, but the real answer is that it depends on the document, and saying so with reasons is what earns the point.
| Approach | How it splits | Where it wins | Where it hurts |
|---|---|---|---|
| Fixed-size | Every N tokens with a small overlap | Cheap, predictable, fast to build | Cuts sentences and tables mid-thought |
| Structure-aware | On headings, then paragraphs, then sentences | Keeps whole ideas together | Needs clean structure in the source |
| Semantic | Groups sentences by embedding similarity | Chunks follow topic shifts | Slow, and costs extra embeddings to build |
| Parent-document | Index small pieces, return the larger section they belong to | Precise match plus full context | More plumbing, bigger stored index |
If you remember one thing here, make it the parent-document trick: index small chunks for precise matching, then feed the model the larger section they came from. It fixes the “relevant but context-free” failure without wrecking retrieval precision, and interviewers perk up when you bring it unprompted.
Dense retrieval, sparse retrieval, and why you want both
Pure vector search has a known blind spot. Ask about error code “TX-4092” and embedding similarity may rank a paragraph about general transaction failures above the one page that names the exact code, because the model never saw that token in training and its vector sits close to noise. Keyword search (BM25) nails exact matches and proper nouns but misses paraphrase. Running both and merging the results with reciprocal rank fusion covers each one’s weakness, and hybrid retrieval is close to table stakes for a strong answer now.
For the vector index itself, HNSW is the default worth naming. It gives roughly logarithmic search over millions of vectors and dominates production setups; the cost is memory, since the graph lives in RAM. When the interviewer pushes on 50 million or 500 million vectors, that’s your cue to talk about IVF or product quantization, which trade a little recall for a much smaller footprint, and about sharding the index across nodes. Bring up metadata filtering too, because half of real queries are scoped (“only docs from this team, published after January”) and a design that can’t filter on metadata falls apart on the first follow-up.
Reranking is the cheap win people forget
The retriever is built for speed, not judgment. A bi-encoder embeds the query and the documents separately, so it never compares them directly; it just measures vector distance. A cross-encoder reads the query and a candidate together and scores actual relevance, which is far more accurate and far too slow to run over a whole corpus. So you use both: pull the top 50 with the fast retriever, then rerank down to the top 5 with a cross-encoder like Cohere Rerank or bge-reranker-v2-m3. This two-stage shape adds maybe 50 to 200 milliseconds and reliably lifts answer quality. It’s the single biggest gain you can bolt onto a mediocre pipeline, and skipping it when the interviewer asks “how would you improve retrieval quality” is a missed layup.
Generation, and the parts that bite in production
Now you assemble a prompt: the retrieved chunks, the user’s question, and an instruction to answer only from the provided context and say “I don’t know” when the context doesn’t cover it. That last rule is what stops the system from confidently inventing a refund policy. Budget the context window carefully, because five chunks plus a long question plus system instructions can crowd out room for the answer, and cramming twenty chunks in makes the model worse, not better, as the one that matters gets buried in the middle where attention is weakest.
Citations matter more than candidates expect. If the answer links back to the source chunk, a wrong answer becomes debuggable and a right answer becomes trustworthy. Teams building enterprise assistants care about this a lot, because “the bot made something up and a customer acted on it” is the failure that gets a project shut down.
How you prove it works
The strongest signal you can send is that you know how to measure this thing, because most people can’t. Split the evaluation in two. Retrieval either found the right chunks or it didn’t, and you can check that with no model in the loop. Generation either used those chunks faithfully or it drifted. A framework like RAGAS packages the common metrics, but you should be able to say what each one catches on its own.
| Metric | What it catches |
|---|---|
| Context recall | Whether retrieval pulled the chunks that hold the answer |
| Context precision | How much of what you retrieved was actually relevant |
| Faithfulness | Whether the answer is grounded in the chunks or invented |
| Answer relevancy | Whether the answer addresses the question that was asked |
| MRR / hit@k | How high the first correct chunk ranks |
| p95 latency, cost per query | The numbers the product team actually tracks |
Build a golden set of a few hundred real questions with known-good answers and known source chunks, then run it on every change so you can tell whether a new embedding model or a different chunk size helped or just felt better. In production you watch the same faithfulness and latency numbers live, sample low-confidence answers for review, and treat every thumbs-down as a fresh eval case.
Which brings back the deleted-policy question, the one that separates people who’ve run RAG from people who’ve read about it. The model cited a dead document because your vector index still had its embedding; nothing ever told the index the source was gone. The fix is boring, and that’s the point. Your ingestion pipeline needs to process deletes and updates, not inserts alone, and you need a clear story for how fast a wiki edit reaches the index. Answer that one well and the interviewer stops wondering whether you’ve actually built one of these.
Keep sharpening your system design:
