Getting hired at Project Prometheus, Bezos’s physical-AI lab

Updated · techinterview.org

Project Prometheus has hired more than 120 people since it came out of stealth in November 2025, and almost none of them found the job on a careers page. Jeff Bezos and Vik Bajaj built the recruiting machine before they built much of a public face, pulling researchers out of OpenAI, Google DeepMind, xAI, Meta, Anthropic, and Nvidia at a reported pace of roughly 15 to 20 senior hires a month. If you came here looking for the standard loop with a recruiter screen and a timed coding round, reset your expectations. This company hires more like a fund poaching principals than like a big tech org clearing a req.

What they are building shapes every conversation you will have with them. Prometheus is betting on world models for the physical economy: AI that simulates how metal deforms under a press, how a semiconductor fab behaves under load, how a drug candidate moves through a manufacturing process, so those decisions can be tested in simulation before anyone spends money in the real world. The founding advisors include Ashish Vaswani and Jakob Uszkoreit, two of the authors on the original transformer paper, and the November acquisition of General Agents brought in a team working on video-language-action models. The technical bar is a frontier-lab bar. The problem framing is aerospace, chips, cars, and pharma.

What “physical AI” changes about the questions

The roles fall into a few tracks, and knowing which one you are in tells you what the hard round will be. Frontier model researchers, the people with NeurIPS, ICML, or ICLR papers, get pressed on modeling choices and on whether their prior work transfers to physical dynamics. ML infrastructure and systems engineers get a training-at-scale problem with a simulation loop bolted into it. Robotics and embodied-AI specialists get sim-to-real. Data and evaluation engineers get asked how you measure a model that is supposed to predict the physical world when ground truth is expensive and slow to collect. There is also a product-minded engineering track for people who can turn a research artifact into something a manufacturing customer will actually run.

The through-line is that a chatbot answer does not help you here. You can generate a plausible sentence for free and check it instantly. You cannot generate a titanium press run for free, and you find out whether the model was right days later, if at all. Nearly every question they ask circles back to that asymmetry: how do you learn a good model of physics when the data is scarce, noisy, and costs real dollars per sample.

There is no application form, and that is the first filter

Most people enter through a warm introduction or direct outreach from someone already inside. Kyle Kosic, a co-founder of xAI and a former OpenAI infrastructure engineer, is the kind of hire who then pulls in a cluster of people he has worked with. If a founding-team member or a strong former colleague vouches for you, you skip most of the noise. If you are cold, your published work has to do the vouching. Either way the first real conversation is usually with a researcher or engineer, not a recruiter, and it is technical from the first ten minutes.

That first call is a calibration. They want to know what you have actually built with your own hands versus what you supervised, and they will follow any claim you make down two or three levels. Vague ownership is where cold candidates lose it. Say “we improved throughput” and the next question is which part you wrote, what the bottleneck was, and what you tried that did not work.

The research round

If you are on the research track, expect to present a result you are proud of and then defend it under pressure. The interesting part is never the happy-path result. It is what happens when they perturb the setup. A strong candidate has already thought about the failure modes and can reason about them out loud instead of freezing.

A few questions phrased the way they tend to come out:

  • “Walk me through a result you are proud of. Now suppose the labels are 10 percent wrong and you cannot tell which 10 percent. What breaks first, and how would you know?”
  • “We need a model that predicts how a sheet of titanium springs back after a press. Each real run costs thousands of dollars. How do you set up the training signal at all?”
  • “Where does the transformer stop being the right inductive bias for physical dynamics, and what would you reach for instead?”

The titanium question is the tell. There is no clean answer. They are watching whether you reach for simulation-in-the-loop, whether you think about active learning to spend your expensive real runs where they matter, whether you know the difference between a model that interpolates inside its data and one that extrapolates to a new alloy. Confidently reciting a paper you half remember is worse than saying “here is how I would find out.”

The systems round for ML infrastructure

The infra interview looks like a frontier-scale training problem with one twist: the pipeline is doing more than reading tokens off disk. It is calling a physics engine or a renderer in the training loop, and that changes where the bottlenecks live. A GPU cluster that stalls waiting on a slow simulator is a different debugging problem than one that stalls on the network.

  • “You have a few thousand GPUs and a training job with simulation in the loop. GPU usage sits at 40 percent. Where do you look first?”
  • “Design the data path for video-language-action training when one episode is half an hour of 4K footage. What do you precompute, what do you stream?”
  • “Your loss curve looks healthy but eval is flat for three days. Debug it out loud.”

What they reward is a mental model of the whole pipeline and a habit of measuring before guessing. Reach for a profiler, name the specific thing you would measure, and separate the compute-bound from the IO-bound from the sim-bound. Candidates who jump straight to “add more GPUs” without finding the stall tend to get filtered here.

The world-models round

This is the one that surprises people coming straight from language modeling. The domain interview probes whether you can think about physical dynamics as a learning problem. Sim-to-real gap, distribution shift when the model meets an alloy or a part geometry it never trained on, how you validate a predictor when a wrong prediction in production means a scrapped batch on a real factory floor. You do not need a physics PhD, but you need to take the physics seriously rather than treating the environment as an abstract token stream.

A good answer here shows you know why the naive approach fails. Train a model on simulated deformation, deploy it on real presses, and the tiny mismatches between your simulator and reality compound into predictions that are confidently wrong. Talking through how you would detect that drift, calibrate uncertainty, and decide when the model should abstain rather than guess is the signal they want.

Hiring track at Project Prometheus What the hard round centers on A question phrased the way they ask it
Frontier model researcher Defending your own published work, then extending it to physical dynamics under noisy, expensive data “How do you build a predictor when each real training example costs thousands of dollars?”
ML infrastructure and systems Training at scale with a simulator or renderer in the loop; finding the real bottleneck “A few thousand GPUs, simulation in the loop, 40 percent GPU usage. Where do you look?”
Robotics and embodied AI Sim-to-real transfer, control, video-language-action models from the General Agents team “Your policy works in sim and fails on hardware. Walk me through narrowing it down.”
Data and evaluation Measuring a model of the physical world when ground truth is slow and costly to collect “Design an eval where a wrong answer means a scrapped manufacturing batch.”
Product-minded engineering Turning a research artifact into something an aerospace or semiconductor customer will run “A customer wants this model in their fab next quarter. What has to be true?”

What they pay, with the caveat attached

Compensation is the reason people take these calls even when they are happy where they are. Reporting from early 2026 put senior researcher and ML infrastructure packages in the 5 to 10 million dollar range in total compensation, with a handful of marquee hires said to clear 20 million. Treat those as the top of the market, not the median offer, and remember they are press figures rather than a published band. Prometheus is private and does not post levels or numbers, so the only reliable data point is whatever lands in your own inbox.

The structure matters more than the headline. At a private company at roughly a 41 billion dollar valuation after the June 2026 Series B, a large share of any offer is equity that is worth what a future round or exit says it is worth. If you get to the offer stage, push on the split between cash and equity, the strike price and valuation the equity is priced against, and the vesting schedule. A 10 million dollar number built almost entirely on paper at a 41 billion dollar mark is a very different bet than the same number in liquid stock.

What actually gets you through

The people clearing this loop are not the ones with the cleanest LeetCode history. They are researchers and engineers who have shipped something real, can defend it down to the details, and stay curious when a question has no tidy answer. The titanium press question and the stalled-GPU question are the same test wearing different clothes: can you reason about a hard, underspecified problem in the physical world without pretending you already know the answer. Show up with a body of work you can talk about in real detail, a feel for why physical AI is harder than text, and the willingness to say “I would measure that” instead of bluffing. That combination is rarer than it sounds, which is exactly why they are paying what they are paying to find it.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack
Scroll to Top