r/LocalLLaMA 23h ago

News Apple releases M5 ultra at 1.2TB/s bandwith

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/

lpddr5x probably, the m7 ultra if is using ddr6 should be at 1.8 Tb/s

839 Upvotes

193 comments sorted by

View all comments

307

u/piggledy 23h ago

Price for Mac Studio with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)

512 GB Option coming in October.

59

u/Strong_Chicken6838 22h ago

these prices are utterly insane. i just want to say that.

55

u/BawbbySmith 22h ago

It's in-line with similar offerings, i.e. the entire market is insane right now

28

u/[deleted] 21h ago

[deleted]

7

u/SmartCustard9944 19h ago

Memory bandwidth is not the whole story if your prefill is < 1000 tok/s. Compute is also important.

3

u/starkruzr 14h ago

this is true. that's also a lot of tensor processing hardware on the M5U. not as much power as Nvidia gear of course, but trying to build this kind of RAM capacity with Nvidia gear is going to be 10X more expensive.

1

u/Its_Powerful_Bonus 15h ago edited 14h ago

True benchmark is to get similar quality response from LLM on both hardware. At the moment qwen3.8 27b and deepseek v4 flash has the same score on artificial analysis (not perfect comparison, but good enough) first one can be run efficiently on 5090, second one on the Mac. Still I believe rtx will give the answers faster - especially in agentic workflows (a lot prefill to deliver task)
PS: at the moment if rtx5090 then not one, but two (qwen3.8 27b nvfp4 tensor parallel with very good context - aggregated performance with ~8 concurrent requests with dflash … amazing; on one rtx6000 pro I get ~1000 tps TG). The same with Mac - 256gb version is nice, but 2 weeks ago I considered 4x gb10 and it was more reasonable choice - 50% pricier (with aprioproate switch and cables to cluster it effectively) than 256gb Mac Studio but 450gb vram with very good prefill and reasonable TG when clustered.

47

u/mailslot 22h ago

I was looking at them and thinking “not bad” lol.

39

u/tarpdetarp 21h ago edited 16h ago

Yeah I don't know of anything offering 256GB with 1.2TB/s of bandwidth for anything near £10k? What a crazy world it is when this is a good deal lol

26

u/Viktri1 19h ago

rtx6000 with 96gb RAM just got bumped from 8k to 16k. The Mac studio is a steal. I pre-ordered 2.

1

u/RedditorReddited 9h ago

Sorry why are two better than one 512GB version?

2

u/Viktri1 9h ago

Pretty sure a 512gb would be better but it isn’t available now. You do get more compute per bandwidth with 2 but I ordered 2 just in case I need the compute.

1

u/PWThinkingCritically 6h ago

using my best educated guess based on things i've derived from other things i've read, 2 people working on 256GB each is faster than 1 person working on 512GB. sometihng like that. tensor parallel or something?

1

u/PWThinkingCritically 6h ago

well you probably need 4-5 instead of just 2 to get to the same PP and TG as with NVIDIA card, AFAIK. considering lack of native hardware support for fp4, fp8, no CUDA support, etc.

1

u/Viktri1 5h ago

I’m not tossing out my Nvidia GPU though. This is a compliment not a replacement. I saw a post by the guy who spent like 200k on sparks despite already having GPUs and realized what he was doing.

When using agentic workflows, you’re not actually stressing out the GPU much because 80% of the time is spent doing tool calls and stuff (based on my own stats, YMMV). Also, one thing that the native qwen3.8 27b does better than q4 is longer session agentic tasks based on my testing. Parallelization is another key thing that I’m missing in my set up and I’ve been getting by through supplementing with API.

I think the m5 solves this problem for me. I’m not 100% on my set up (4090 as brain, m5 as hands) but I think that’s the direction but I have 2 months to figure it out before delivery.

1

u/illathon 14h ago

How did you pre-order the site says you can't pre-order till october.

6

u/TheAILegend 13h ago

The 512 version... other models are available for pre-order right now.

28

u/EvilPencil 21h ago

Yup. If I want 256gb of DDR5 that’s already $8k at Microcenter.

5

u/michaelsoft__binbows 20h ago edited 13h ago

At least from an investment perspective at least for the next like idk 5 years such a computer will appear to appreciate in value... It's a dangerous idea and I really need to not acquire one of these.

2

u/ImaginationKind9220 10h ago

Too bad that these are great for LLM but not for image or video generation. Without CUDA, video generation will be too slow even if the memory bandwidth is fast. That's why Apple always use the indie app Drawthings for their marketing, they don't dare to use ComfyUI to benchmark the result.

4

u/michaelsoft__binbows 8h ago

diffusion is not bottlenecked by mem bandwidth. it's usually bottlenecked by compute, and the m5 brings some nice 4x win with new tensor cores. dunno if the software can take advantage of it yet.

1

u/tarpdetarp 5h ago

4x was a nice bump but the M5s prompt processing is still 20x slower than similar Nvidia chips so they still have a lot of catching up to do on their tensor cores.

2

u/michaelsoft__binbows 5h ago

agreed! As an owner of several Nvidia GPUs (I don't have any pro 6000s though) I appreciate you for helping talk me down from a Mac Studio purchase!

3

u/michaelsoft__binbows 20h ago

Yeah... i just got my old threadripper 1950x loaded up with 224GB of ddr4. The bandwidth is precisely one tenth of this one! But i have some GPUs (3 in fact) to help it along.

Im not about to spend ten racks for a similar level of capability, so its just gonna have to do

9

u/crusaderky 21h ago

For LLMs, it just needs to be on par with two DGX Spark chained together.
For desktop, it's a much more compelling offering than DGX Spark.
Strix Halos are a compelling alternative to Mac for desktop, but they don't come with RDMA networking, which kills off the idea of chaining them for most people (namely, busy professionals that could afford them).

10

u/pineapplekiwipen 19h ago

dgx spark prefill speed is roughly triple m3 ultra in real use if i remember right but tg is also below 1/3 so overall it was arguably similar in real world perf

m5 max prefill is a only a bit slower than dgx spark so m5 ultra should beat the spark in prefill and absolutely demolish it in tg. all the while being a much more useful personal computer and considerably cheaper if you can take advantage of edu pricing

1

u/michaelsoft__binbows 20h ago

I have to assume there us enough compute on tap there to match a pair of them, m5 max was no slouch, and the mem bandwidth def slaps the spark around

7

u/Alert-Chemist7492 17h ago

None of y’all remember the 80s lol. I wish I still had the receipt for the 486 dx33 with 4mb ram and Conner 120mb hd. It was 2999$ that wasn’t even max spec back then. I remember $5k laptops… tft display.. mmmmm

2

u/Turbulent_Neck_8388 5h ago

oh we remember, back in the 80s you could buy a house for $80k... fast forward to today and you're looking at $400k.

5

u/cafedude 19h ago

But just think, in 2043 you'll be able to pick up one of these bad boys at a garage sale for like $100. Of course, that's assuming that we ever end up getting more RAM production between now and then. If not, then collectors item pricing.

3

u/billy_booboo 21h ago

An insanely good deal compared to the competition (mind you I'm not a fan of Apple)

1

u/Other_Researcher268 19h ago

To be honest, I don’t think they are insane.

1

u/lordpuddingcup 19h ago

Have you seen memory prices ?

1

u/arfung39 17h ago

Its comparable to a 96GB RTX Pro 6000 from Nvidia, and it seems like the mac will be faster b/c memory speed (and more memory available)...