system design

Why ELT won, and when interviewers still expect ETL

The question usually arrives dressed up as a design prompt. “You’re pulling data from a Postgres database and three SaaS APIs into a warehouse for the analytics team. Walk me through it.” The interviewer is listening for one decision. If you say you’ll clean and reshape the data before it lands in the warehouse, you’re describing ETL. If you say you’ll dump it in raw and shape it with SQL afterward, that’s ELT. Same three letters. The order is the whole conversation.

ETL means extract, transform, load: pull data out of the source, run it through a transformation step (often a Python job, a Spark cluster, or an old Informatica server), then write finished tables into the destination. ELT swaps the last two steps. You extract, load the raw data straight into the warehouse, and transform it there with the warehouse’s own compute. For most of the past decade that swap has been the default answer, and an interviewer wants to hear that you understand why it happened, not simply that it did.

Why ELT became the default

The reason is unglamorous and it’s about money. Old warehouses coupled storage and compute, so landing terabytes of raw data you might never query was slow and expensive. Snowflake and BigQuery split those two apart. Storage got cheap enough to stop thinking about, and compute became elastic: spin up a big warehouse for a heavy transform, then let it idle. Once that was true, there was little reason to transform on the way in. Land everything, keep the raw copy, and reshape it whenever a new question shows up.

That reordering dragged a specific set of tools along with it. Extraction and loading got commoditized by connector vendors. Fivetran ships more than 500 prebuilt connectors; Airbyte has 600-plus on the open-source side and is aiming at a thousand by the end of 2026. You don’t write a Salesforce or Stripe extractor anymore, you configure one. Transformation moved into dbt, which is really SQL models plus a dependency graph, running inside the warehouse. So the standard stack today is a connector tool for the E and L, a cloud warehouse in the middle, and dbt for the T. A candidate who can name that shape and say what each piece does has cleared the bar for most data-engineering screens.

One current detail worth carrying into 2026 interviews: Fivetran and dbt Labs merged, closing the deal on June 1. The two halves of the reference stack now ship from a single vendor. It doesn’t change the concepts, but if you talk about “the Fivetran and dbt stack” as two separate companies, a sharp interviewer may quietly note that you haven’t been keeping up.

The follow-up that separates people

“ELT is just better” is a losing answer, because the next question is “when wouldn’t you use it?” and now you’ve boxed yourself in. ETL is not legacy. It’s the correct call in a few specific situations, and naming them is how you show you understand the tradeoff instead of reciting a vendor blog. The prompts tend to sound like this:

  • “Why did the industry move from ETL to ELT, and what changed to make that possible?”
  • “You’re loading customer records that include Social Security numbers. Would you still use ELT?”
  • “In an ELT setup, where does your compute cost actually land, and how do you keep it from exploding?”
  • “What is reverse ETL, and when would you reach for it?”

The sensitive-data case is the big one. If you’re pulling records with PII, health data, or anything under a compliance regime, loading it raw means the unmasked data now sits in the warehouse where every analyst with access can query it. Transforming first lets you hash, tokenize, or drop the sensitive columns before they ever touch storage. Stripe, Uber, and plenty of healthcare companies run genuine ETL on their regulated pipelines for exactly this reason.

The other cases come up less but still count. A legacy on-prem warehouse without elastic compute can’t absorb heavy in-database transforms, so you do the work outside it. Transformations that aren’t expressible in SQL, image processing, ML feature generation, parsing odd binary formats, want a real compute engine and not a SELECT statement. And edge or IoT deployments often transform at the source, because shipping raw sensor data over the wire is wasteful.

The move that signals real experience is ETLT: extract, apply a light transform (mask the PII, drop obvious garbage, dedupe), load into the warehouse, then run the heavy analytical transforms in dbt. You get the compliance benefit of handling sensitive fields early and the flexibility of doing analytics work in SQL later. Bringing up ETLT unprompted tends to shift an interviewer’s read of you from “studied the flashcards” to “has actually built this.”

ETL and ELT, side by side

Dimension ETL (transform before load) ELT (transform inside the warehouse)
Where the transform runs A separate engine (Spark, Python, Informatica) before the warehouse The warehouse itself (Snowflake, BigQuery) via SQL and dbt
Raw data retained? Usually not; only the transformed result lands Yes; raw lands first and is kept for reprocessing and audit
Cost model Compute on a dedicated ETL cluster; warehouse stores clean data only Warehouse compute for every transform; storage for raw is cheap
Schema handling Schema-on-write, defined up front Schema-on-read, structure applied at model or query time
Best fit Regulated and PII data, legacy warehouses, non-SQL transforms Cloud warehouse with elastic compute, SQL-expressible logic, SQL-fluent team
Typical 2026 tools Informatica, Spark, custom Python with masking before load Fivetran or Airbyte to load, dbt to transform, Census or Hightouch for reverse ETL
Main risk Reprocessing means re-extracting from the source Warehouse bill balloons if you rebuild everything on every run

Where the compute really lands, and who pays

A question that trips people: “In ELT, what cost are you signing up for?” The reflexive answer is that ELT is cheaper. Often it is, but the sharper answer is that you’ve moved the compute bill into the warehouse. Every dbt model that rebuilds a large table is warehouse compute, metered by the second on Snowflake or by bytes scanned on BigQuery. Teams that land everything and rebuild everything nightly get a startling invoice. The mature response talks about incremental models that only process new rows, materialization choices (view versus table versus incremental), and not rebuilding a hundred-million-row table because ten thousand rows changed.

This is also where idempotency shows up. A pipeline worth trusting can rerun the same load without duplicating data or corrupting downstream tables. In ELT that usually means merge or upsert logic keyed on a primary key, plus a clean boundary between the raw landing tables, a cleaned staging layer, and the business-facing marts. If someone asks you to sketch the warehouse layout, “raw, staging, marts” is the layering they’re fishing for, with raw kept as an append-only record of whatever the source actually sent.

Reverse ETL, and why it exists

Once the warehouse holds clean, modeled data, teams want it back in the tools where work happens: the CRM, the ad platform, the support desk. Pushing warehouse tables back out to those operational systems is reverse ETL, handled by tools like Census or Hightouch. It closes the loop. Raw data flows in through Fivetran, gets modeled by dbt, and the useful result syncs back to Salesforce so a sales rep sees a churn score without opening a dashboard. Interviewers raise it to check whether you picture the warehouse as a dead-end reporting store or as the center of the data flow. The second framing is the one they want.

If you want to sound current without overreaching, the plain version is that the T keeps moving. It started before the load, moved after it into the warehouse, and now lives both early (light compliance transforms) and late (dbt models, then reverse-ETL logic on the way back out). Get the ordering straight, know why cheap elastic compute triggered the shift, and keep one real reason ETL still exists in your back pocket. That will land better than the answer most candidates give, which is memorizing which letter comes first.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top