Harvey pays software engineers a median total package north of $340K a year, per Levels.fyi data from 2026, and it screens for something most AI startups treat as a bonus: whether you can reason about what an M&A associate or a litigator actually needs from the tool you’re building. The company sells to more than a thousand law firms and corporate legal teams, and its engineers spend real hours sitting with lawyers. Walk into the loop treating legal work as a black box and the project round will find out.
Why the domain changes this interview
Gabriel Pereyra and Winston Weinberg started Harvey in 2022. Pereyra came out of AI research at Meta and DeepMind; Weinberg was a first-year litigation associate at O’Melveny & Myers. That split runs through the whole process. About half of what they test is ordinary engineering, and the other half is whether you can take a messy legal task and shape it into something a model does reliably enough that a lawyer will stake their name on the output.
Know the product surface before you show up. Assistant covers day-to-day drafting and research. Vault runs analysis across large document sets, up to roughly 100,000 files in a single project, which is the scale a due-diligence review or a regulatory second request actually reaches. Knowledge grounds answers in a firm’s own precedent, templates, and prior deals. Sitting on top are workflow agents that chain the steps an associate or paralegal would otherwise grind through by hand. When an interviewer asks you to design a feature, it is almost always one of these, so be clear on the difference between retrieval over a firm’s private corpus and generation from a general model.
The rounds, and what each one is really scoring
The loop runs about three to five weeks: a recruiter screen, a hiring manager screen, a technical phone screen, then a virtual onsite of four or five rounds. Nothing exotic in the shape of it. The signal lives in the emphasis.
The recruiter call is logistics and fit: your background, target level, location, and comp expectations. The hiring manager screen, usually 45 minutes, is where “why legal AI, why now” carries weight. A generic answer about wanting to work on frontier models lands flat. They want a reason you care that this particular domain gets solved, because the work involves a lot of unglamorous grinding on document formats and evaluation, and people without that pull tend to drift.
The technical phone screen is 60 minutes with one medium coding problem in a shared editor. Talk through your approach before you type, then test it on a small example. Expect arrays, strings, and hashmaps, sometimes with a parsing flavor since so much of the real work is chewing through documents. It is not an obscure dynamic-programming gauntlet.
The onsite adds another coding round (often a medium with a follow-up that changes a constraint and makes you adapt your solution live), a system design round, a behavioral round, and a project walkthrough. That walkthrough is the round candidates underrate. Bring something you actually built and be ready to go three or four decisions deep on why you chose what you chose. They probe tradeoffs, not the narrative.
A few problems candidates have reported, phrased close to the way they came up:
- Given a long document and a phrase, return every position where the phrase appears.
- Merge a set of overlapping intervals and return the consolidated ranges.
- Design retrieval over a law firm’s private documents so each answer cites the exact source clause.
- Walk through a project where a non-engineer depended on what you shipped, and describe what broke.
| Interview round | Format and length | What Harvey is scoring |
|---|---|---|
| Recruiter screen | ~30 min call | Background, target level, location, comp expectations |
| Hiring manager screen | ~45 min call | Motivation for legal AI, role fit, past work |
| Technical phone screen | 60 min, shared editor | One medium coding problem, reasoning out loud |
| Onsite coding | 45-60 min | Medium with a live follow-up constraint |
| Onsite system design | 45-60 min | LLM-application design: retrieval, evals, tenancy |
| Onsite behavioral | ~45 min | Working with lawyers and non-engineers, handling ambiguity |
| Project walkthrough | 45-60 min | Depth on real tradeoffs you owned |
What a Harvey system design answer needs
Harvey’s design rounds skew toward LLM-application design rather than distributed-systems trivia. A common prompt sounds like “design a contract review feature over 50,000 uploaded PDFs” or “build retrieval over a firm’s document library so every answer cites the exact clause it came from.” The interviewer is listening for a chunking strategy, embeddings, a vector index, a reranking step, and how you ground generation so the model quotes the source instead of paraphrasing from memory.
Two constraints separate a strong answer from a generic retrieval diagram. The first is confidentiality. Each customer is a firm with privilege obligations, so one client’s documents can never bleed into another client’s context, and matter-level access control has to hold even inside a single firm. The second is faithfulness. In Mata v. Avianca, lawyers filed a brief with citations a chatbot invented and drew sanctions for it. A fabricated citation is not a cosmetic defect in this product, it is the failure that ends a customer relationship. So evaluation, citation checking, and a path for the system to say “I don’t know” belong in your design from the start, not bolted on at the end. It also helps to know Harvey routes across model providers rather than betting on one, so treating the model as a swappable component reads as informed rather than naive.
How to prep without burning three weeks
On coding, drill mediums across arrays, strings, hashmaps, and graphs, and get comfortable narrating while you write. Skip the exotic material. On the domain, spend an hour understanding how contract review and due diligence actually work and why a confident wrong answer is far more dangerous in law than in a consumer chatbot. That reading pays off in the behavioral and product rounds more than another LeetCode hard would.
Come in with an opinion on one Harvey feature: what you would change, what you would measure, where it might fail a real user. That single piece of preparation separates candidates who want an AI job from candidates who want this one. For system design, practice retrieval over a private corpus end to end, including evaluation and multi-tenant isolation, until you can sketch it without stalling.
On money, treat public numbers as a starting range, not gospel. Levels.fyi and 6figr put Harvey software-engineering total comp somewhere from the low $300Ks to around $490K depending on level and equity assumptions, and startup equity at a company repricing this fast is the part most worth pressing the recruiter on. Ask how they value the equity, what the latest 409A implies, and what a realistic four-year outcome looks like across a couple of scenarios instead of accepting one headline figure.
Practice the behavioral round:
