quant finance

A quant probability cheat sheet that holds up live

Most probability questions in a trading interview collapse to a single move: write down the expected value, then ask whether the conditioning, the variance, or the stopping rule changes the number you just wrote. Do that calmly out loud and you clear the first probability round at SIG, Optiver, IMC, and the early screens at Citadel or Jane Street. The math rarely runs past a second-year course. Staying organized while someone watches the clock is the actual test.

Here is the reference I keep in my head before a quant loop: the formulas that recur, the distributions worth knowing cold, and the setups interviewers reuse because they separate people who memorized an answer from people who can rebuild it.

Expected value is the spine of every question

For a discrete variable, E[X] = Σ x·P(X = x). Say that first, then reach for linearity of expectation: E[X + Y] = E[X] + E[Y], which holds even when X and Y are dependent. That second fact is where most of your speed comes from, because it lets you skip joint distributions entirely.

Roll a die until a six appears. Expected rolls? Geometric with p = 1/6, so 1/p = 6. Now the version they actually ask: expected rolls to see all six faces, the coupon collector. You wait a geometric time for each new face, and the waits add: 6·(1 + 1/2 + 1/3 + 1/4 + 1/5 + 1/6) ≈ 14.7. No joint distribution, just six expectations summed. Reach for a messy combinatorial sum here and you have missed the point of the question.

The trade framing underneath all of this: take the bet when its expected value is positive after fees and slippage, size it by how much variance you can stomach. Interviewers dress EV questions up as games with a die, a deck, or a coin, but they are watching whether you instinctively price the game.

Conditional probability and the base-rate trap

P(A | B) = P(A ∩ B) / P(B), and Bayes rewrites it as P(A | B) = P(B | A)·P(A) / P(B). The setup that catches people: a test is 99% accurate, a disease hits 1 in 10,000, you test positive, what is the chance you have it? The reflex answer is 99%. The real answer sits near 1%, because the false positives from the huge healthy population swamp the true positives. Plug in: the true-positive mass is 0.99 × 0.0001 ≈ 0.0001, the false-positive mass is roughly 0.01 × 1 ≈ 0.01, so the posterior is about 0.0001 / 0.0101, near one in a hundred. Naming the base rate before you compute is the signal they want.

Monty Hall lives in the same family: switching wins 2/3 of the time because the host’s choice is not independent of where the car is. Explain why the conditioning matters instead of reciting “switch,” and you are done with that question.

Linearity does the heavy lifting

Indicator variables turn scary counting problems into one-line answers. Throw n letters into n addressed envelopes at random; expected number that land correctly? Each letter matches with probability 1/n, there are n of them, so the expected count is exactly 1, regardless of n. Expected number of people who get their own hat back from a random pile: also 1. You never touched the distribution of derangements, you summed n indicators.

This pattern handles a surprising share of “expected number of ___” questions. Define an indicator for each position, find its probability, add them up.

The distributions worth knowing cold

You should be able to write the mean and variance of these without thinking, and recognize which one a word problem is describing.

Distribution Shows up as Mean Variance
Bernoulli(p) a single yes/no trial p p(1 − p)
Binomial(n, p) successes in n independent trials np np(1 − p)
Geometric(p) trials until the first success 1/p (1 − p)/p²
Poisson(λ) rare events in a fixed window λ λ
Uniform(a, b) flat odds across an interval (a + b)/2 (b − a)²/12
Exponential(λ) waiting time, memoryless 1/λ 1/λ²
Normal(μ, σ²) sums and averages via the CLT μ σ²

Two properties earn their own mention. The geometric and exponential are memoryless: P(X > s + t | X > s) = P(X > t), so a component that has already survived s units is as good as new. And the normal carries the 68 / 95 / 99.7 rule for one, two, and three standard deviations, which is enough to eyeball a tail probability when an interviewer asks for a rough number instead of an exact one.

Why the central limit theorem keeps coming up

Average enough independent draws from almost any distribution and the average looks normal, with the spread shrinking like 1/√n. This is why a market maker quoting thousands of small edges worries about variance per trade but sleeps fine on the aggregate. Expect something like: you win $1 with probability 0.51 and lose $1 otherwise, play 10,000 times, roughly what does your P&L look like? The mean is 10,000 × 0.02 = $200, the standard deviation is about √10,000 = 100, so you are very likely positive but not guaranteed. Showing that the edge is real and the noise is bounded is the whole answer.

Markov chains and absorbing walks

Gambler’s ruin is the template. On a symmetric random walk between 0 and N starting at k, the probability of reaching N before 0 is k/N, and the expected number of steps until you hit a boundary is k(N − k). When the steps are biased, the hitting probability turns into a ratio of powers of (q/p), worth deriving once so it never surprises you.

The method generalizes to anything you can describe as states. Write one equation per state for the expected time to absorption, E_i = 1 + Σ_j p_ij · E_j, then solve the linear system. A clean example: how many fair coin flips on average until you first see HH? Set up a state for “no progress” and one for “just saw an H,” solve, and you get 6. Until you first see HT? Only 4. People find that gap baffling, and explaining it (an HH attempt that fails wastes the H you were holding, while an HT attempt never does) shows you understand the state machine rather than a memorized formula.

Variance, for when the question turns to risk

Var(X) = E[X²] − (E[X])². Scaling: Var(aX + b) = a²·Var(X), since shifting by a constant moves the mean but not the spread. For independent variables, variances add: Var(X + Y) = Var(X) + Var(Y). When an interviewer pivots from “what is the expected payout” to “how would you size this,” they are asking about variance, and the jump from EV to risk-adjusted sizing is the conversation they actually want.

The phrasings you will hear

A few that recur almost verbatim across firms:

  • “You flip a fair coin until the first heads. What is the expected number of flips, and what would you pay to play if the nth flip pays $2 to the n?” (the St. Petersburg setup, where the naive EV is infinite)
  • “Two people each pick a uniform random number in [0, 1]. What is the probability they land within 0.1 of each other?”
  • “I roll two dice. Given that at least one shows a 6, what is the probability both are 6?” (the answer is 1/11, not 1/6)
  • “Over 100 fair flips, what is the expected number of times three heads appear in a row?”

None of these need exotic math. They need you to set up the right random variable, name the distribution or write the recurrence, and land on a number you can defend.

How to actually answer one

State your assumptions before computing, because half the credit is in the framing. Name the distribution or define the states out loud. Do the arithmetic where the interviewer can follow it, and sanity-check the magnitude at the end, since a probability above 1 or an expected count larger than the number of trials means you should start over. In the early rounds a structured wrong answer beats a silent correct one, because the desk is hiring for how you think when the clock is running, and a clean recovery from a wrong turn reads as exactly the skill they need on a live book.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top