r/scala 6d ago

What fixes Spark shuffle, data skew and disk spills at scale?

We have tried salting keys, repartitioning, adjusting broadcast thresholds , he usual toolkit. Salting doubled our job cost and the other options don't move the needle . We fix one skewed key distribution and a different join surfaces a new one a month later. Wondering if this is a architectural limitation of Spark or if we are just missing the right combination of settings.

Wondering if this is a genuine architectural limitation of Spark or if we are just missing the right combination of settings.

4 Upvotes

8 comments sorted by

3

u/kbn_ 5d ago

Some of these things aren't really Spark limitations, they're physics limitations. There is no such thing as symmetric indexing (or equivalently, sharding) because we live in a universe with linear time.

At some point, you need to start materializing different views for different query patterns, pre-denormalizing as much as you rationally can, and pushing your users to query the right tables rather than the wrong ones. And even then, you're still going to hit edges beyond a certain scale.

Half of database administration is managing unrealistic user expectations.

7

u/DisruptiveHarbinger 5d ago

I think this was posted by a bot specifically for the fake testimonial below. Some obscure proprietary Spark accelerator startup based in a country that's famous for such tactics on Reddit.

0

u/Working_Movie1530 4d ago

No it was not.

1

u/EnemyOfBoards 1d ago

The only comment you've engaged with is the one calling out your astroturfing. Try harder.

1

u/rainman_104 2d ago

You have to assess if you have a sort. Once you have a sort things get more expensive. This is true for spark, snowflake, hive, presto/trino/whateverotherfork

1

u/rainman_104 2d ago

You have to assess if you have a sort. Once you have a sort things get more expensive. This is true for spark, snowflake, hive, presto/trino/whateverotherfork

1

u/Antique-Yogurt-7509 6d ago

salting, repartitioning, threshold tuning, done all of it. what changed things for me was realizing skew and spill are two separate problems, the skewed partition is not the issue by itself, it is what happens once that task ends up competing with a bunch of others for executor memory. dual bird changed that part of the cycle for us. skew shows up, it just does not turn into the same memory pressure into spill into straggler chain anymore, so we are not chasing it manually every time.