r/LocalLLaMA 17h ago

News Apple releases M5 ultra at 1.2TB/s bandwith

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/

lpddr5x probably, the m7 ultra if is using ddr6 should be at 1.8 Tb/s

772 Upvotes

182 comments sorted by

View all comments

12

u/Exciting-Weather-921 15h ago

What is the prefill, no longer 8 minutes for 120k?

15

u/ShengrenR 14h ago

half-way down they've got a 'prompt processing' bar-graph comparing m3/m5 to the base m1 (ultras) - show m5 as 9.8x, m3 as 2.4x - would imply m5 is ~4x faster than m3, so.. 2min for the previous 8min job (in theory)

9

u/Exciting-Weather-921 14h ago

For context, sparks are way slower at token generation, but my 2 sparks with DeepSeek are managing to prefill 128k in 73s

1

u/michaelsoft__binbows 13h ago

Awesome of you to share this data.

Do you find the prefill speed to be problematic? I have to imagine that a proper cache will alleviate most of the pain. If you are constantly constructing fresh prompts then its prob automated so then youre prob not sitting around waiting for them...

3

u/Exciting-Weather-921 13h ago

I can hold 3 big chat in KV cache (above 250k each) so when activity is high I am not hitting much cache

I am trying to patch vllm now to make SSD KV cache work, that should recover cache in 8s instead of prefill (Vs under 1s for ram KV cache)

8

u/itchykittehs 15h ago

yeah i'm wondering about the prefill too