coding interview questions

Why your Spark job is slow, and what interviewers ask about it

The fastest way to fail a Spark interview is to talk about groupByKey as if it were free. Teams that run real pipelines, the streaming groups at Netflix, ad-tech shops, anyone with a warehouse measured in petabytes, mostly want to know one thing. Do you understand what moves across the network when your code runs, and can you stop it from moving. The rest is detail.

Spark interviews sprawl, which is part of the problem. A data engineering loop might have one round on SQL, one on pipeline design, and one that is pure Spark internals, and the internals round is where people who use Spark every day tend to fall apart. They know the API. They have never watched the Spark UI during a slow job and asked why stage 4 has one task still running after the other 199 finished.

The questions that actually get asked in that round sound like:

  • “Walk me through what happens when you call join on two large DataFrames.”
  • “A stage has one task still running after all the others finished. What’s going on, and how do you fix it?”
  • “Why is reduceByKey better than groupByKey?”
  • “What is spark.sql.shuffle.partitions set to by default, and when would you change it?”

RDDs still come up, and here is why

You will get asked to explain RDDs even though you write DataFrames all day. Interviewers ask because the RDD model is where the execution concepts live: partitions, the lineage graph, and narrow versus wide dependencies. A DataFrame is a higher-level API that compiles through the Catalyst optimizer down to RDD operations anyway, so the person who understands RDDs can reason about why a query plan looks the way it does.

The distinction that earns points is narrow versus wide transformations. A narrow transformation like map, filter, or union keeps each output partition dependent on a single input partition, so it runs in place with no data movement. A wide transformation like groupByKey, join, or repartition needs records from many input partitions to land together, which forces a shuffle. Stage boundaries in Spark fall exactly on the wide transformations. If you can read code and predict where the stage boundaries are, you can predict where the job spends its time.

The shuffle is the whole game

A shuffle writes intermediate data to local disk on every executor, then reads it back across the network so that all records with the same key sit in the same partition. Disk writes, network transfer, serialization on the way out and deserialization on the way in. That is the expensive part of almost every slow Spark job, and interviewers probe it hard.

The default number of partitions after a shuffle is 200, set by spark.sql.shuffle.partitions. This default is a classic trap. Two hundred partitions for a 50 GB join means each partition is around 250 MB, which is fine. Two hundred partitions for a 5 TB join means 25 GB per partition, which will spill to disk and can blow up the executor. The number was reasonable in 2015 and is often wrong now. A strong answer notes that Adaptive Query Execution coalesces this count down at runtime, so the common failure mode has shifted from too many tiny partitions to one giant skewed one.

reduceByKey beats groupByKey, and you should be able to say why

This is the most common single Spark question, and the answer is map-side combine. groupByKey ships every value across the network and groups them on the reducer side, so if you are summing a billion rows down to a thousand keys, you still move a billion rows. reduceByKey applies the reduce function locally on each partition first, so each partition sends at most one partial result per key. Same output, a fraction of the shuffle.

# moves every value across the network
rdd.groupByKey().mapValues(sum)

# combines locally first, then shuffles partials
rdd.reduceByKey(lambda a, b: a + b)

In the DataFrame API you rarely call these directly, but the same idea shows up: a groupBy().agg() gets a partial aggregation pushed below the shuffle by Catalyst, which is the map-side combine happening for you. If an interviewer asks why the DataFrame version beats a hand-rolled RDD groupByKey, that is the answer.

Skew: the one task that never finishes

Data skew is the problem senior interviewers actually care about, because it is the one that pages you at 2 a.m. You join on user_id, one user is a bot with 40 million events, and that single key lands in one partition. The stage shows 199 tasks done in two minutes and the 200th running for an hour. AQE handles the common case now by splitting the skewed partition, but you should know the manual fix, because AQE does not catch everything, especially on the write side or with a user-defined partitioner.

The manual technique is salting. You append a random suffix to the hot key so its rows spread across many partitions, join against a dimension table you have exploded to match every salt value, then drop the salt. It is ugly and it works.

from pyspark.sql import functions as F

salted = big.withColumn("k", F.concat(
    F.col("user_id"), F.lit("_"), (F.rand() * 16).cast("int")))
# dim table replicated across all 16 salt buckets, then join on k

A cheaper move, when it fits, is to filter the handful of hot keys out, join them separately with a broadcast, and union the result back. Interviewers like this answer because it shows you would rather avoid the shuffle than tune it.

Broadcast joins and the 10 MB line

When one side of a join is small, Spark can send a full copy to every executor and skip the shuffle entirely. That is a broadcast hash join, and it is the single biggest win available on a two-table join where one side is a dimension lookup. The automatic threshold is spark.sql.autoBroadcastJoinThreshold, default 10 MB. A frequent production fix is raising that to 50 or 100 MB so a mid-size lookup table broadcasts instead of triggering a sort-merge join that shuffles both sides.

The gotcha the interviewer is waiting for: broadcasting is bounded by driver and executor memory, and Spark estimates table size from statistics that can be stale, so it sometimes tries to broadcast something far larger than it thinks and throws an out-of-memory error on the driver. Knowing that failure mode is worth more than reciting the happy path.

The settings you should know cold

These come up by name. If you can state the default and the reason you would change it, you sound like someone who has tuned a job rather than read a blog post about one.

Spark setting Default What it controls Why you would change it
spark.sql.shuffle.partitions 200 Partition count after a shuffle (joins, aggregations) Raise for multi-terabyte data so partitions do not spill; AQE now coalesces it down at runtime
spark.sql.autoBroadcastJoinThreshold 10 MB Largest table Spark will broadcast instead of shuffle-joining Raise to 50-100 MB to broadcast a mid-size dimension table; set to -1 to turn off
spark.sql.adaptive.enabled true (since Spark 3.2) Adaptive Query Execution, runtime re-planning from real statistics Almost always left on; a few legacy jobs pin a fixed plan
spark.sql.adaptive.skewJoin.enabled true Splits skewed partitions in a sort-merge join at runtime Left on; depends on AQE being enabled
spark.executor.memory 1g Heap memory per executor process Raise when the UI shows out-of-memory errors or heavy disk spill

What Adaptive Query Execution changed

AQE has been on by default since Spark 3.2, and Spark 4.0 in 2025 tightened it further. It re-plans the query at runtime using real shuffle statistics instead of trusting the estimates Catalyst made before any data moved. It coalesces the 200 shuffle partitions down to a sensible count when the data turns out small, flips a sort-merge join to a broadcast join once it sees a side is actually tiny, and splits skewed partitions in a sort-merge join.

The interview value is knowing its limits. AQE reacts to what it observes at a shuffle boundary, so it cannot help a job whose cost is somewhere else, and it will not rescue a broadcast that fails on bad statistics. When someone asks why the job is still slow after you enabled AQE, the answer is usually that the bottleneck sits upstream of any decision AQE gets to make: a bad file layout, too few input partitions, or skew on the write path.

coalesce, repartition, cache, persist

The quick ones that trip people up. repartition(n) does a full shuffle and can raise or lower the partition count while rebalancing data evenly. coalesce(n) only reduces the count and avoids a shuffle by merging existing partitions, which is what you want before writing output so you do not produce a thousand tiny files, though it can create skew if the merge is uneven. Reaching for repartition when coalesce would do adds a shuffle you did not need.

cache() is just persist() with the default storage level, memory first and spilling to disk. The question behind the question is when caching hurts. If you read a dataset once, caching it wastes memory and adds overhead, and Spark’s lazy evaluation means cache() does nothing until an action forces it, so people who call it and see no change are usually confused about that timing.

Read the Spark UI on a job that is failing

The best preparation is to run a deliberately bad job and read the Spark UI while it struggles. Write a groupByKey over a skewed dataset, watch one task hang while the rest finish, then fix it and watch the stage flatten out. An interviewer can tell within two questions whether you have done that or whether you are reciting definitions. The person who has stared at a stuck stage talks about tasks and partitions and spill; the person who has not talks about big data and distributed processing in the abstract. One of them gets the offer.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

newsletter

What's actually being asked right now

Interview patterns & comp trends, straight to your inbox.

No spam. Unsubscribe anytime.

1972 Soviet postage stamp commemorating the Mars 2 probe

worth a read

Mars For The Rest of Us — a weekly-or-more deep dive on the technical side of Mars exploration: rocket propulsion, microbiology, mission architecture, and everything in between. Written by Maciej Ceglowski.

Read it on Substack →
Scroll to Top