top of page


Diagnosing Data Skew in PySpark for Azure Databricks AutoBroadcast Joins Salting and AQE
A PySpark job can look healthy for 95% of its runtime, then sit frozen on a handful of tasks while one executor keeps spilling, retrying, and finally dying with an out-of-memory error. That pattern is often not a cluster sizing problem. It is usually data skew. In Azure Databricks, skew shows up most clearly during shuffle-heavy stages: joins, aggregations, window functions, `dropDuplicates`, and wide transformations. One or a few keys carry far more rows than the rest, so Sp
Ramesh D
6 hours ago9 min read
Â
Â
Â


Architecting Enterprise Medallion Architecture with Microsoft Fabric and Azure Databricks
Enterprise data platforms fail less often because of bad dashboards and more often because the layers underneath are unclear. Raw files get overwritten. Schemas drift quietly. Duplicate events inflate metrics. Bad records disappear into logs. A production Medallion Architecture solves these problems by making data quality, lineage, and ownership explicit from ingestion to business consumption. A strong design with Microsoft Fabric and Azure Databricks gives teams the best of
Ramesh D
7 hours ago8 min read
Â
Â
Â
bottom of page
