AI/ML interview questions

What RAG evaluation interviews test beyond recall@k

A RAG system returns a wrong answer, and the interviewer asks what you’d check first. The weak response reaches straight for the prompt. The strong one asks a narrower question: did the right chunk even reach the model, or did the model have what it needed and still get it wrong? Most retrieval-augmented generation interviews turn on that split, because a retrieval miss and a generation miss need completely different fixes.

One accuracy number hides two different bugs

The instinct on a first pass is to score the whole pipeline end to end: feed in a question, compare the final answer to a reference, count how often it’s right. That number is real, and it’s also close to useless for debugging. A 62% pass rate tells you the system is mediocre. It doesn’t tell you whether the retriever is dropping the relevant passage half the time or whether the generator is inventing details the context never mentioned.

Those two problems point in opposite directions. If retrieval is the bottleneck, you tune chunk size, embedding model, or the reranker, and no amount of prompt work will save you. If generation is the bottleneck, the context already contains the answer and the model is ignoring it, padding it, or contradicting it, and swapping your vector index does nothing. Interviewers want to see you separate the two before you touch anything, because candidates who skip that step tend to spend a week optimizing the component that was already fine.

The retrieval metrics you’re expected to name

Retrieval is an information-retrieval problem wearing a new hat, so the metrics come straight from that world. Recall@k is the one that matters most for RAG: of all the chunks that were actually relevant to the question, what fraction showed up in the top k you retrieved. A reasonable bar on a broad corpus is around 0.8 at k=20. Recall dominates precision here for a specific reason. You can afford to hand the model a few irrelevant chunks, but a relevant chunk that never gets retrieved is gone, and the generator can’t cite what it never saw.

Precision@k still matters, just less, and mostly because stuffing the context with junk creates its own failure. Models lose track of information buried in the middle of a long context, so retrieving forty chunks to be safe can score worse than retrieving eight good ones. When there’s exactly one correct passage per question, MRR is the cleaner metric: the reciprocal of the rank where the first relevant chunk appears, averaged over your queries, so a right answer at position one scores 1.0 and one at position five scores 0.2. When relevance is graded rather than yes-or-no, and rank order matters, NDCG@k is the standard, rewarding a system that puts the best chunk first. Aim for roughly 0.8 at k=10.

A common follow-up: why not eyeball the retrieved chunks yourself? Because you can’t eyeball a regression. Change your chunking from 512 tokens to 256 and your recall@k might climb on short factual questions while it craters on ones that need a full paragraph of context. You only catch that with numbers on a fixed query set.

Faithfulness, relevance, and the RAG triad

Once the right context is in hand, the generation side gets its own scores. The pattern most teams borrow comes from Ragas, which popularized four metrics: faithfulness, answer relevance, context precision, and context recall. Faithfulness, sometimes called groundedness, is the one interviewers probe hardest. The judge breaks the generated answer into individual claims, then checks each claim against the retrieved context. Faithfulness is the share of claims the context actually supports, so an answer with three grounded statements and one invented figure scores 0.75. That single metric is your hallucination detector.

Answer relevance asks a different question: does the response address what was asked, or does it wander into correct-but-unrequested territory? Context precision, and the broader idea of context relevance, closes the loop back to retrieval by scoring whether the retrieved passages were on-topic in the first place. TruLens packages the three that carry the most diagnostic weight as the RAG triad: context relevance, groundedness, and answer relevance. If all three are high and the answer is still wrong, you’ve found a genuinely hard case worth a human’s attention.

RAG evaluation metric What it measures Needs a labeled reference? Rough passing bar
Recall@k Fraction of the relevant chunks that land in the top k retrieved Yes, relevance labels ~0.8 at k=20
MRR Reciprocal rank of the first relevant chunk, averaged over queries Yes Higher is better; 0 when the relevant chunk is never retrieved
NDCG@k Rank-weighted retrieval quality that rewards putting the best chunk first Yes, graded labels ~0.8 at k=10
Context precision Share of retrieved passages that were actually on-topic No, LLM judge ~0.7
Faithfulness / groundedness Share of answer claims supported by the retrieved context No, LLM judge ~0.75
Answer relevance Whether the answer addresses the question that was asked No, LLM judge ~0.8
Context recall Share of ground-truth answer claims covered by the retrieved context Yes, reference answer ~0.8

Context recall is the odd one out because it needs a reference answer to score against, while faithfulness and answer relevance can run without labels. That distinction comes up constantly, because labeled data is the expensive part, and knowing which metrics you can run label-free signals that you’ve actually shipped one of these.

Why LLM-as-judge invites the follow-up questions

All those generation metrics get computed by another model reading the answer and scoring it. A capable judge like GPT-4o clears 80% agreement with humans at telling a genuinely relevant passage from a hard negative built to look relevant, which is good enough to steer development and not good enough to trust blindly. Interviewers know this and will push on where it breaks.

The judge inherits the blind spots of the model family it belongs to, so a judge from the same family as your generator will forgive the exact mistakes that generator makes. It also carries biases worth naming. It tends to prefer the first option in a pairwise comparison, it rewards longer answers, and it scores its own family’s outputs higher. In a specialized domain like medicine or securities law, it fails quietly, rating a fluent but functionally wrong answer as correct because it can’t tell the difference either. The mitigations are the interesting part of the answer: swap the order of options and average the two runs to cancel position bias, calibrate the judge against a few hundred human labels before you trust its numbers, and pick a judge from a different model family than the one you’re grading.

Building an eval set without hand-labeling everything

The practical objection is that all of this needs a test set, and nobody wants to label ten thousand question-answer pairs by hand. The standard move is to generate synthetic questions from your own corpus. Point a strong model at each document and have it write questions that document answers, which hands you a query with a known correct source chunk for free. You clean the obvious junk, keep a few hundred that read like real user questions, and freeze that as your golden set. Ragas and DeepEval both ship generators for this, and Arize Phoenix exposes retrieval evaluators for MRR and NDCG if you want the classic IR numbers next to the LLM-judged ones.

The freeze matters more than the size. A fixed set turns every change into a measurable before-and-after. Switch embedding models, adjust the reranker, rewrite the system prompt, then rerun the same queries and watch the metrics move. Without it, you’re back to eyeballing outputs and trusting a gut feeling that this version seems better, which is the exact habit the interview is checking you’ve outgrown.

What the questions actually sound like

Interviews on this rarely ask you to recite definitions. They hand you a broken system and watch you reason. A typical prompt: “Users say the answers are wrong maybe a third of the time. You have logs of the questions, the retrieved chunks, and the answers. What do you measure first?” The answer they want starts by splitting retrieval from generation, names recall@k and faithfulness as the two dials, and only then talks about fixes.

  • “How would you tell whether a wrong answer is a retrieval failure or a generation failure, using only the logs?”
  • “Why is recall usually more important than precision when you retrieve for a language model?”
  • “Your faithfulness score is 0.9 but users still complain. What’s going on?”
  • “How do you build an evaluation set without manually labeling thousands of examples?”
  • “Your generator and your judge are the same base model. Why is that a problem?”

The last one separates people who’ve read about RAG from people who’ve run it in production. A faithfulness score of 0.9 with unhappy users usually means the retrieved context was faithful but incomplete or subtly off-topic, so the model grounded its answer perfectly in the wrong passage. That sends you back to context relevance and recall, not to the generator. Tracing that chain out loud, from the metric to the likely cause to the component you’d touch, is the whole point of the round.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top