# What Decagon’s engineering interview loop really tests

Source: https://www.techinterview.org/post/3233475471/decagon-engineering-interview-loop/
Updated: 2026-07-02 · techinterview.org

Decagon builds AI agents that close customer support tickets end to end, with no human quietly taking over halfway through the conversation. Its engineering interview is shaped around that one hard problem: an agent that talks to real customers, takes real actions against real systems, and has to be right often enough that a company will let it run without a babysitter. Walk in treating this as a generic startup loop and you'll feel the gap inside the first system design round.

A little context helps you read the questions correctly. Decagon was founded in 2023 by Jesse Zhang and Ashwin Sreenivas, whose backgrounds run through Palantir, Citadel Securities, and a pair of earlier startups (Lowkey and Helia). The product resolves support for companies like Notion, Eventbrite, and Bilt, and the pitch to those customers is a deflection rate north of 80%, meaning most tickets never reach a person. By early 2026 the company had raised a Series D that put its valuation around $4.5 billion, and it had shipped a voice product built with ElevenLabs. None of that is trivia. It tells you the interview cares about agents that act, retrieval over messy enterprise knowledge, and measuring whether the agent was actually correct.

## The shape of the loop

Reported loops drift as a company grows and splits roles, so treat the table below as a shape rather than a fixed script. Decagon hires across software engineering, research engineering, and more agent-focused roles, and the emphasis shifts depending on which one you're in. Research-leaning candidates report a faster, more concentrated process; product engineering loops tend to run longer and lean harder on system design and past work.

| Stage | Rough length | What they're reading |
| --- | --- | --- |
| Recruiter screen | 30 min | Why agentic support, why now, comp and timeline fit |
| Coding pair | 60 min | Working code in TypeScript or Python, often with concurrency or an LLM-tooling flavor |
| System design | 60 min | How you'd build an agent that resolves tickets, and where it breaks |
| Project walkthrough | 60 min | Something hard you actually shipped, pressure-tested for depth |
| Behavioral / values | 45 min | Ownership, speed, how you handle being wrong in production |

Start to finish is usually three to four weeks. The bar is high in the way frontier-AI startups tend to be: fewer rounds than a big company, but each one goes deeper and the interviewers are mostly engineers who built the thing you're being asked about.

## The coding round is about building agents, not balancing trees

The pair-programming round leans practical. You're more likely to get a small, open-ended build than a tidy LeetCode hard. Think parsing a stream of tool-call results, writing a retry wrapper around a flaky API, deduping events that arrive out of order, or wiring up a tiny agent loop that has to decide between calling a tool and replying to the user. They want readable code, sensible handling of the unhappy path, and someone who narrates tradeoffs instead of going silent for ten minutes.

The "LLM-tooling flavor" trips people up because it rewards a mental model many candidates haven't built yet. A useful way to think about it: an agent that takes actions is a loop with a model in the middle proposing tool calls, plus a lot of code making sure those calls are safe and bounded. If you can write something like this from memory and talk about every line, you're in good shape.


```
async function resolve(ticket) {
  let state = init(ticket);
  for (let step = 0; step < MAX_STEPS; step++) {
    const action = await model.next(state);       // propose a tool call or a reply
    if (action.type === "reply") return action;   // hand the conversation back
    if (!allowed(action.tool, ticket.user))       // authz before any side effect
      return escalate(ticket, "tool not permitted");
    const result = await tools[action.tool](action.args);
    state = append(state, action, result);
  }
  return escalate(ticket, "step budget exceeded");
}
```


The interesting parts are the parts that aren't the model. The step budget stops a confused agent from looping forever and burning tokens. The authorization check runs before the side effect, because issuing a refund or canceling an order is not something you let a language model do on vibes. The escalation path means failure has a defined exit instead of a hang. If the interviewer pokes at your code, it'll usually be here: what happens when a tool times out, when the model asks for the same action twice, when two agents act on one account at once.

## System design: build an agent that resolves a support ticket

This is the round that separates people. The prompt is some version of "design a system that autonomously resolves customer support tickets for an enterprise," and the strong answers spend their time on the parts that are genuinely hard, not on drawing a load balancer in front of a web server.

The retrieval layer is where most candidates get loose. A support agent needs the right company-specific context for a question that might be phrased in a hundred ways, pulled from help docs, past tickets, internal policy, and live order data. Talk about how you chunk and index that knowledge, how you keep it fresh when a refund policy changes on Tuesday, and how you stop the agent from confidently citing a doc that was deprecated months ago. Saying "throw it in a vector database" is a starting point, not an answer. Be ready to discuss hybrid retrieval, why you'd re-rank, and how you'd handle a customer whose question depends on their own account state rather than any document.

Then there are actions. A read-only agent that answers questions is a different risk profile than one that can cancel a subscription or move a shipment. Good candidates separate those tiers explicitly, gate the dangerous actions behind permission checks and sometimes a confirmation, and design for the case where an action half-succeeds. Idempotency keys come up naturally here. So does the question of what the agent does when it isn't sure: a clean handoff to a human, with the full conversation and the agent's reasoning attached, beats a wrong action every time.

Expect the interviewer to push on failure and scale. What happens when the model is down? When a customer tries to talk the agent into giving a refund it shouldn't? How do you keep latency low enough that the conversation feels live, especially once voice is involved and a two-second pause sounds like a dropped call? You don't need a perfect answer to all of it. You need to show you see the sharp edges and have an opinion about each one.

## The eval question they care about more than you'd expect

Because Decagon sells deflection, it lives or dies on knowing whether the agent was actually right. So somewhere in the loop, usually inside system design, you'll get pulled toward evaluation: how would you measure whether an autonomous resolution was correct?

This is harder than it sounds and interviewers know it. Customer satisfaction surveys are sparse and biased. A ticket marked "resolved" might mean the customer gave up. The thoughtful path covers a few angles at once: offline evals on a labeled set of past conversations, LLM-as-judge graders with their own well-known failure modes, human review on a sampled slice, and production signals like reopen rate and escalation rate. Bonus points for talking about regression testing the agent the way you'd test code, so a prompt change that fixes one case doesn't silently break thirty others. If you've ever shipped something where you couldn't fully trust your own metrics, that story lands well here.

## What gets people cut, and how to prep

The most common failure isn't a missed algorithm. It's a candidate who designs the happy path beautifully and never reckons with the agent being wrong. Decagon's whole business is the gap between an agent that's right 80% of the time and one a Fortune 500 will trust on its brand. If your design has no answer for the bad 20%, no escalation, no guardrails, no way to know it happened, you've shown them the wrong instinct.

To get ready, build a tiny agent yourself before you interview. Wire a model up to two or three tools, give it a real task, and watch it fail. You'll learn more about loops, retries, authorization, and evaluation from one afternoon of that than from any amount of reading, and you'll have a concrete story for the project round. Brush up on retrieval beyond the cartoon version, since that's where strong candidates pull ahead. And come with a genuine opinion about where autonomous support should and shouldn't be trusted, because the people interviewing you have very specific opinions and they're testing whether you've thought about it at all.

Decagon is hiring fast against a problem that's still half-unsolved, which is exactly why the interview rewards engineers who treat reliability as the interesting part rather than the boring part. Show up as someone who finds the failure cases fun, and the loop tends to go your way.
