top of page


Diagnosing Data Skew in PySpark for Azure Databricks AutoBroadcast Joins Salting and AQE
A PySpark job can look healthy for 95% of its runtime, then sit frozen on a handful of tasks while one executor keeps spilling, retrying, and finally dying with an out-of-memory error. That pattern is often not a cluster sizing problem. It is usually data skew. In Azure Databricks, skew shows up most clearly during shuffle-heavy stages: joins, aggregations, window functions, `dropDuplicates`, and wide transformations. One or a few keys carry far more rows than the rest, so Sp
Ramesh D
7 hours ago9 min read
Â
Â
Â
bottom of page
