A candidate presents a long-short equity strategy: Sharpe 4.2 from 2015 through 2023, worst drawdown under 6 percent. The interviewer doesn’t ask about the alpha. The first question is how many versions got tried before this one landed. If the real answer is a few hundred, the Sharpe on the slide is measuring persistence more than edge.
The distance between a backtest and live trading is the most tested idea in a quant research loop, and it surfaces in every round that touches methodology. A paper Sharpe of 4 that lands at 0.5 in production isn’t bad luck. It’s the predictable result of several biases stacking, each one shaving the number, a couple of them able to erase it outright.
The number that doesn’t survive contact with production
Start with why the collapse happens at all. A gross backtest sees a cleaner world than the one you trade in. It often peeks at information that wasn’t public yet. It gets selected as the best of many attempts. And it pays costs that no real desk pays. Any one of those shaves a point or two off the Sharpe. Together they routinely take a 4 down to something between 0 and 1, and they do it every time, which is why experienced researchers assume the haircut before they see it.
Interviewers know this, so they rarely reward a high number on its own. They probe whether you know where yours leaks. A Sharpe of 1.5 that you can defend line by line beats a Sharpe of 4 you can’t explain.
Lookahead bias hides in the timestamps
The most common way a backtest lies is by using data that wasn’t available at the moment it claims to trade. The clean version everyone catches: computing a signal from the day’s closing price and then trading at that same close. Nobody can do that. The subtle versions are what separate people.
Fundamental data is the classic trap. A company’s quarter ends March 31, but the 10-Q might not be public until early May, and the figure you pull from Compustat today may be a restated number that didn’t exist in that form back then. Use it on April 1 and you’re trading on information nobody had. Point-in-time databases exist precisely because the naive pull gives you the latest known value, not the value known on the date you’re simulating. Lag every fundamental to its actual release, not its period end.
Index membership has the same problem. Build a universe from today’s S&P 500 and run it back to 2005, and you’ve silently written out every company that went bankrupt, got acquired, or fell out of the index. Corporate actions, dividends applied before the ex-date, prices forward-filled over a trading halt, all of these leak the future in small, hard-to-spot ways.
Survivorship and the universe you forgot to include
Survivorship is lookahead’s quieter cousin. Many vendor datasets, if you’re not careful, contain only the securities that still exist. Delisted, bankrupt, and merged names drop out. A value or momentum strategy tested only on survivors looks far stronger than it was, because the companies that would have blown up your book aren’t in the sample. Adding delisted names back, with their real returns through the delisting, typically knocks half a point to a point and a half off a long equity Sharpe.
The reverse mistake matters too. Including names that were never actually tradable, illiquid microcaps that print enormous paper returns you could never capture at size, inflates the backtest in a way live execution will never reproduce.
The multiple-testing problem is the one that actually gets you
This is the result every serious quant interviewer wants you to have internalized. Given enough trials on the same data, you can manufacture a high Sharpe from pure noise. Bailey, Borwein, López de Prado, and Zhu made it explicit in Pseudo-Mathematics and Financial Charlatanism (Notices of the AMS, 2014): the expected maximum Sharpe from N independent strategies with zero true edge grows roughly with the square root of 2 ln N. Test a thousand random signals and the best one shows a Sharpe near 3.7 by chance alone.
That’s why “how many things did you try?” is the sharpest question in the room. A researcher who tested one pre-registered hypothesis and got Sharpe 2 has found something. A researcher who grid-searched ten thousand parameter sets and reports the maximum has found noise with good marketing. The Deflated Sharpe Ratio and the Probability of Backtest Overfitting (PBO) are the tools that haircut a result for the number of trials behind it. You don’t need to derive them on the whiteboard, but you should be able to say why a reported Sharpe means little without the trial count attached.
Costs, slippage, and the capacity nobody modeled
Gross returns are a fantasy. Every fill pays the half-spread, moves the price against you by an amount that grows with your participation (market impact scales roughly with the square root of size), and shorts pay borrow and financing on top. A mid-frequency equity strategy at Sharpe 3 gross can sit at 0.5 net once you charge realistic costs, and a high-turnover version can go negative.
Capacity is the part junior candidates miss. A signal that works on $10 million can be flat at $500 million, because at size you become the market you were trying to trade against. Your own orders move the price, your alpha decays, and the pretty backtest was run at a size you’ll never actually deploy. Pod shops like Millennium, Citadel, and Point72 press on capacity hard, since a signal that can’t hold real money isn’t a book.
Validating without fooling yourself
The defense is structural, not clever. Hold out a test set you look at once. Prefer walk-forward evaluation, where you fit on a window and test on the next, rolling forward, so the model never sees its own future. Plain k-fold cross-validation leaks in time series because neighboring samples share information, so the standard tool is purged, embargoed cross-validation: drop training samples whose label windows overlap the test period, then embargo a gap right after it. López de Prado lays this out in Advances in Financial Machine Learning, and it comes up by name in research rounds.
The cheapest defense is discipline. Write down the hypothesis before you test it. Count every configuration you run. Freeze the out-of-sample data until the strategy is final. None of it is glamorous, and all of it is what stops you from shipping a backtest that dies in week one.
What the interview actually looks like
A quant researcher loop usually opens with a phone screen on probability and statistics, then a research or methodology round where an interviewer takes your project apart, a coding round in vectorized NumPy or pandas (sometimes a deliberately buggy backtest to fix), and a presentation where you defend a piece of real work against people who do this for a living. The methodology round is where everything above gets tested, often through your own project rather than an abstract prompt.
The phrasing tends to be direct:
- “Your backtest shows a Sharpe of 3. Give me three reasons I shouldn’t believe it.”
- “You tried 500 signals and kept the best. What’s its real expected Sharpe?”
- “How would you avoid lookahead when your signal uses quarterly earnings?”
- “Set up cross-validation for a strategy on daily equity data. What breaks if you use plain k-fold?”
- “How do you model transaction costs for something that trades a few times a day?”
Keeping the failure modes in one place, with the check that catches each, makes the pattern obvious:
| Backtest bias | What goes wrong in the simulation | Rough effect on Sharpe | The check that catches it |
|---|---|---|---|
| Lookahead / point-in-time | Signal uses data not yet public: restated fundamentals, same-bar close, today’s index membership | Can flip a losing strategy into a winning one | Point-in-time data; lag every fundamental to its release date, not its period end |
| Survivorship | Delisted, bankrupt, and merged names dropped from the universe | Overstates a long-equity Sharpe by roughly 0.5 to 1.5 | Full historical constituents including delistings, with real returns through delisting |
| Multiple testing / selection | Best of many trials reported as if it were the only hypothesis | Sharpe of 2 to 4 reachable from pure noise past ~1,000 trials | Deflated Sharpe / PBO; count trials; keep an untouched holdout |
| Costs ignored | Gross returns with no spread, market impact, borrow, or financing | Halves or erases a mid-frequency Sharpe | Charge half-spread plus square-root impact, borrow, and financing |
| Capacity blindness | Signal decays as assets grow because your own orders move the price | Live Sharpe far below paper once deployed at size | Stress the strategy at target size; model participation and impact |
| Cross-validation leakage | Overlapping labels leak training data into the test window | Inflated out-of-sample metrics that still overfit | Purged, embargoed cross-validation instead of plain k-fold |
The candidates who do well here aren’t the ones with the highest backtest Sharpe. They’re the ones who walk in already knowing their number is inflated and can tell you by how much and why. Present a smaller, defensible Sharpe with the trial count and the cost model attached, and you read like someone who has watched a live strategy underperform its backtest and gone and found the leak. That’s who the desk wants, because the alternative is the person who still trusts the 4.
