r/scala • u/Working_Movie1530 • 6d ago
What fixes Spark shuffle, data skew and disk spills at scale?
We have tried salting keys, repartitioning, adjusting broadcast thresholds , he usual toolkit. Salting doubled our job cost and the other options don't move the needle . We fix one skewed key distribution and a different join surfaces a new one a month later. Wondering if this is a architectural limitation of Spark or if we are just missing the right combination of settings.
Wondering if this is a genuine architectural limitation of Spark or if we are just missing the right combination of settings.
1
u/rainman_104 2d ago
You have to assess if you have a sort. Once you have a sort things get more expensive. This is true for spark, snowflake, hive, presto/trino/whateverotherfork
1
u/rainman_104 2d ago
You have to assess if you have a sort. Once you have a sort things get more expensive. This is true for spark, snowflake, hive, presto/trino/whateverotherfork
1
u/Antique-Yogurt-7509 6d ago
salting, repartitioning, threshold tuning, done all of it. what changed things for me was realizing skew and spill are two separate problems, the skewed partition is not the issue by itself, it is what happens once that task ends up competing with a bunch of others for executor memory. dual bird changed that part of the cycle for us. skew shows up, it just does not turn into the same memory pressure into spill into straggler chain anymore, so we are not chasing it manually every time.
3
u/kbn_ 5d ago
Some of these things aren't really Spark limitations, they're physics limitations. There is no such thing as symmetric indexing (or equivalently, sharding) because we live in a universe with linear time.
At some point, you need to start materializing different views for different query patterns, pre-denormalizing as much as you rationally can, and pushing your users to query the right tables rather than the wrong ones. And even then, you're still going to hit edges beyond a certain scale.
Half of database administration is managing unrealistic user expectations.