Cognition runs one of the smaller engineering teams you’ll interview with relative to how much code ships out of it. This is the company behind Devin, the autonomous coding agent that went viral in March 2024, and the team that bought what was left of Windsurf in July 2025 after Google hired away its founders. A few dozen engineers, a product that Goldman Sachs now describes as an “AI employee,” and a hiring loop tuned to find people who can build agent infrastructure that holds up while a language model in the loop keeps trying to do something dumb.
If you’re prepping for this one, drop the assumption that it’s a standard big-tech loop with the logo swapped. The questions overlap, but the weighting is different, and the difference is the whole game.
What the loop looks like
Expect a recruiter or hiring-manager call, a technical phone screen with live coding, then an onsite that runs three to five sessions across coding, systems, and a deeper conversation about how you work. The onsite is usually in San Francisco and in person, because the culture is aggressively in-office. For most IC roles they want three or more years of real engineering behind you. New-grad pipelines exist but they’re thin, and the bar in the room doesn’t drop much when you walk in green.
The phone screen is closer to pair programming than to a quiz. You get a problem that resembles something Devin’s infra team would actually hit: parse a stream of tool-call outputs, reconcile state after a step fails, write the retry logic that doesn’t double-apply a side effect. They care more about how you reason through the messy middle than whether you nail the optimal Big-O on the first pass.
The coding round isn’t LeetCode hazing
You’ll still see data-structure work, but Cognition leans toward problems with real engineering texture rather than a single clever trick. A common shape is “here’s a simplified version of something we run in production, extend it.” Build a small in-memory file system an agent can read and write against. Implement a token-budget tracker that evicts the least useful context when you blow past a limit. Write a function that diffs two versions of a directory tree and produces the minimal set of edits to get from one to the other.
The tell that you’re doing well is that you treat the model as an adversary. Real agent code has to assume the thing calling it will pass garbage, loop forever, or confidently invoke a tool with the wrong arguments. Write the happy path and stop, and you’ve missed what the role is. Reach for idempotency, timeouts, and a way to bound the blast radius before anyone asks you to, and you’re speaking their language.
Designing the sandbox
The systems round is where Cognition diverges hardest from a generic interview. Instead of “design Twitter,” you get something like “design the execution environment Devin runs inside.” That one prompt has a dozen real sub-problems hiding in it, and the interviewer will pull on whichever thread you look weakest on.
The threads they tend to pull:
- How do you isolate untrusted code the agent generates so it can’t reach the host or other tenants? Containers, microVMs like Firecracker, gVisor, and the tradeoffs between them.
- How does a task get a fresh, warm environment in under a second when it starts? Cold-starting a container per run is too slow.
- How do you snapshot and restore state so a long task can pause, branch, or roll back a bad step?
- What happens when the agent runs for six hours? How do you stream logs, checkpoint progress, and recover if the orchestrator dies mid-task?
- How do you give the agent network access for package installs without opening a hole someone can pull secrets out through?
You don’t need to have built this exact system. You need to decompose it, reason about the security boundary without hand-waving, and put numbers on the tradeoffs. A candidate who says “I’d use a container” and stops is below the bar. A candidate who says “containers share the host kernel, so for arbitrary model-generated code I’d want a microVM like Firecracker for the isolation, eat the few-hundred-millisecond start cost, and keep a warm pool to hide it” is having the conversation Cognition wants to have.
Tool use, context, and the parts most people haven’t touched
A chunk of the loop probes territory new enough that few candidates have production scars from it. How do you design a tool interface an LLM can call reliably? Structured schemas, validation, and what you do when the model returns malformed JSON anyway. How do you manage a context window across a task that produces far more output than fits in it? How would you build evals so you can tell whether a change to the agent made it better or just different from last week?
The Model Context Protocol comes up because Cognition builds against it. You won’t get quizzed on the spec line by line, but knowing why a standard tool-calling protocol matters, and where it leaks in practice, signals you’ve been paying attention to the part of the field they live in. The same goes for the difference between an agent that plans once and an agent that re-plans after every step, and why the second is harder to make cheap.
Where the weight sits, versus a typical big-tech loop
| Signal | Generic FAANG loop | Cognition |
|---|---|---|
| Algorithmic puzzles | Heavy | Light, and always grounded in a real task |
| Distributed-systems design | Standard “design X” prompts | Agent execution, sandboxing, long-horizon orchestration |
| LLM and agent fluency | Rarely tested | Central, including tool use and evals |
| Shipping speed and ownership | Implied | Probed directly; they expect you to move fast |
| In-person commitment | Often hybrid | Mostly in-office in SF, stated up front |
What they’re actually scoring
Two things sit underneath every round. The first is judgment under ambiguity. Agent products fail in weird, non-deterministic ways, and the people who do well there are comfortable making a call with incomplete information and then instrumenting the system so they learn fast. Interviewers will deliberately under-specify a problem to see whether you ask the right clarifying question or charge ahead building the wrong thing.
The second is communication, and Cognition is unusually explicit that they weight it. Not presentation polish. The ability to take a tangled technical idea and make it legible to someone in thirty seconds. On a small team where one person owns a whole surface, an engineer who builds the right thing but can’t explain it costs the team a lot. Expect at least one session that’s mostly conversation: walk me through the hardest system you’ve built, where it broke, what you’d change now.
Pace shows up too. Cognition is candid about being an intense place: in person, long hours, fast iteration. The interview is partly a filter for people who want that and partly a preview so you can opt out. If a team that ships daily and expects you in the office five days a week sounds draining rather than energizing, the comp won’t fix it, and they would rather both sides learn that in the loop than three months in.
How to prep without overfitting
Build something agentic before you walk in. Wire up a small tool-calling loop against an LLM API, give it two or three tools, and watch it fall over. You’ll learn more about retries, timeouts, and malformed output in a weekend of that than in a month of reading. If you’ve ever written a sandbox, a job runner, a CI executor, or anything that runs untrusted code, pull those stories to the front of your memory, because they map almost directly onto the systems round.
Refresh the security-boundary material: containers versus VMs versus microVMs, what gVisor and Firecracker each buy you, how seccomp and namespaces actually work. Skim the Model Context Protocol docs, and read a couple of postmortems on agent failures, including Cognition’s own writing on what eighteen months of Devin in production taught them. Keep your coding sharp on medium-difficulty problems, because the live round still expects clean, working code at conversation speed.
The candidates who struggle are usually strong generalists who treated this like any other onsite and never engaged with the agent layer. The ones who get offers have opinions about why agents fail and a few hard-won ideas about how to make them fail less often. Show up with the second kind of preparation and the rest of the loop gets a lot more forgiving.
Practice the behavioral round:
