coding interview questions

The debugging interview and how not to freeze in it

You clone a repo, run the test suite, and three tests come back red. There’s no array to sort and no clever data structure to reach for, just a few thousand lines of someone else’s code and a failing assertion insisting the invoice total is wrong. Forty-five minutes on the clock. This is the debugging round, and over the past two years it has gone from a Stripe quirk to a standard filter at Amazon, Retool, Meta, and Google.

It tests a skill that pure algorithm practice never touched: walking into unfamiliar code, forming a theory about what’s broken, and proving it. Plenty of engineers who can write Dijkstra from memory come apart the moment the task is “here is a working-ish system, find the one line that lies.”

The formats you’ll actually run into

The shape changes by company, and the differences decide how you should prepare.

Amazon bolts a debugging section onto the online assessment. You get five to seven short code blocks, each carrying a single planted logic error, and about twenty minutes for the whole set. The bugs are small on purpose: a flipped comparison operator, an off-by-one in a loop bound, a return statement sitting one line too high so the function bails before it finishes. Speed matters as much as accuracy, because the section is auto-graded and timed tight.

Stripe’s version, which engineers there call the Bug Squash, drops you into a real open-source library with a known bug and a failing test. The codebase is large and mostly uncommented, so the difficulty is navigational: the hard part is finding the three files that matter out of two hundred. The work is locating the fault, not inventing an algorithm.

Retool runs a similar failing-test repo as a first-round screen, except the repo is small and custom, and an interviewer watches the whole time. Sometimes more than one thing is broken. They care less about whether you close every failing test and more about how you move through the code with someone looking over your shoulder.

Google’s newer code comprehension round hands you an existing system to read, debug, and optimize, with Gemini sitting right there as an assistant. The broken thing is usually a production-flavored AI system, a retrieval pipeline or an agent loop or an eval harness, rather than a textbook data-structures problem. Meta’s AI-enabled loop often opens with bug-fixing before it lets you write anything new, on the theory that reading code is a stronger signal than writing it.

Company What you get Where the difficulty hides Watched live
Amazon Five to seven short snippets, one bug each, inside the OA Speed under a tight clock No, auto-graded
Stripe Real library, known bug, failing test Finding the right files in a big codebase Usually
Retool Small custom repo, one or more bugs Working cleanly while observed Yes
Google Existing AI system, with an AI assistant on hand Reading fast and not trusting the model blindly Yes
Meta Bug-fix opener before you build anything Comprehension before construction Yes

Why the industry swung this way

Two things happened at once. Autocomplete got good enough that a blank-editor LeetCode problem stopped separating candidates, since a model can emit a correct two-pointer solution faster than most people can read the prompt. And hiring managers stopped pretending: the job is mostly reading code you didn’t write and working out why it misbehaves in production. A round that mirrors that is harder to fake and closer to the actual work.

Debugging also resists full outsourcing to an assistant. To get useful help from a model you have to know which function to point it at and how to check its answer, and that judgment is exactly what the round is measuring.

How your process gets scored

The interviewer is watching for one habit above all: whether you change a single variable at a time and predict what will happen before you run anything. Strong candidates reproduce the failure first, read the actual stack trace instead of scrolling past it, and state a specific hypothesis before touching a key. Weaker candidates start editing on a hunch, flip a sign here, add a null check there, rerun, and hope the number of red tests goes down.

The method that reads as senior is dull and repeatable. Reproduce the bug so you can see it fail on demand. Read what the error actually says. Narrow the location by bisecting, dropping a breakpoint, or printing state at the boundary between what looks right and what looks wrong. Then make the smallest change your hypothesis predicts will fix it, and confirm both that the test passes and that you can explain why.

Consider a function whose test says the tax line comes out a few cents low:

def invoice_total(items):
    subtotal = sum(item.price for item in items)
    if subtotal > 100:
        subtotal *= 0.9  # 10% off orders over $100
    tax = round(subtotal * 0.0875)
    return subtotal + tax

Nothing here looks wrong at a glance, and that’s the point. The bug is that round() with no second argument snaps the tax to the nearest whole dollar, so an $88 subtotal gets $8 of tax instead of $7.70. You find it by reading the failing assertion and the tax line together, not by rewriting the function. The fix is round(subtotal * 0.0875, 2), and being able to say why is the part that scores.

What to practice

More LeetCode barely moves this. The practice that transfers is reading code you’ve never seen and driving a test suite through it. Clone a mid-size open-source project, pick a real open issue, reproduce it, and use git bisect to walk back to the commit that introduced it. Do that a dozen times and the round stops feeling foreign.

Learn the bug classes that show up over and over: off-by-one errors on loop bounds and slice indices, a comparison operator pointed the wrong way, mutating a collection while you iterate over it, an early return that skips the rest of the work, integer division where you wanted float, a variable shadowed in an inner scope, an awaited result that resolves in the wrong order, and rounding or timezone mistakes in money and date math. Know your debugger and your language’s stack traces well enough that you’re not reading them for the first time under pressure.

It helps to rehearse against prompts phrased the way interviewers phrase them:

  • “Here’s a failing test for our rate limiter. It lets one extra request through. Find out why.”
  • “This pagination endpoint drops the last row on the final page. Fix it.”
  • “Our retry wrapper retries even on success. Make it stop.”

When they let you use AI

When Gemini or an internal model is available, the assistant tends to expose weak candidates rather than rescue them. Paste the whole file and ask “what’s the bug” and you’ll often get a confident wrong answer, and the interviewer watches you accept it. The people who look strong use the model the way a good engineer uses a sharp junior: ask it to explain one specific function, have it propose a hypothesis, then check that hypothesis against the failing test yourself. Spotting when the model is bluffing has become part of the signal.

The through-line across every version of this round is the same. The team wants to see what you do on a Tuesday when a payment fails in production and nobody knows why. Reproduce it, read what the system is telling you, change one thing, and be ready to say why the fix is right instead of just watching the test turn green.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top