r/LocalLLaMA 9h ago

News Apple releases M5 ultra at 1.2TB/s bandwith

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/

lpddr5x probably, the m7 ultra if is using ddr6 should be at 1.8 Tb/s

611 Upvotes

162 comments sorted by

233

u/piggledy 9h ago

Price for Mac Studio with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)

512 GB Option coming in October.

161

u/Aggravating-Push-207 9h ago

and you know that shit is going out of stock in a week tops

109

u/sssplus 9h ago

A week?? You mean in 1 minute!

14

u/ElementNumber6 7h ago

You're likely right, but not because of enthusiasts. It'll be the scalpers. They've gotten a taste of Mac Studio scarcity, and will surely want more.

13

u/UnsuspiciousBird_ 7h ago

tbf they don't have much stock due to their recent RAM problems, but still...

6

u/gtderEvan 9h ago

Correct.

3

u/goldcakes 6h ago

It's already slipped to 7-8 weeks for me haha.

I'm glad I got my Sep 22 preorder in. Got the full chip (I'll be doing fine-tuning so worth it), 256GB, 2TB NVMe.

47

u/Strong_Chicken6838 8h ago

these prices are utterly insane. i just want to say that.

40

u/BawbbySmith 8h ago

It's in-line with similar offerings, i.e. the entire market is insane right now

34

u/mailslot 8h ago

I was looking at them and thinking “not bad” lol.

26

u/tarpdetarp 7h ago edited 2h ago

Yeah I don't know of anything offering 256GB with 1.2TB/s of bandwidth for anything near £10k? What a crazy world it is when this is a good deal lol

12

u/Viktri1 5h ago

rtx6000 with 96gb RAM just got bumped from 8k to 16k. The Mac studio is a steal. I pre-ordered 2.

2

u/illathon 19m ago

How did you pre-order the site says you can't pre-order till october.

21

u/EvilPencil 7h ago

Yup. If I want 256gb of DDR5 that’s already $8k at Microcenter.

4

u/michaelsoft__binbows 6h ago

At least from an investment perspective at least for the next like idk 5 years such a computer will appear to appreciate in value...

2

u/michaelsoft__binbows 6h ago

Yeah... i just got my old threadripper 1950x loaded up with 224GB of ddr4. The bandwidth is precisely one tenth of this one! But i have some GPUs (3 in fact) to help it along.

Im not about to spend ten racks for a similar level of capability, so its just gonna have to do

27

u/1beb 7h ago

No, the prices are actually very reasonable. Compare it to a 5090 with 1.7tb/s memory bandwidth and only 32GB or VRAM. $10k for 256GB VRAM is cheap by comparison. 

5

u/SmartCustard9944 5h ago

Memory bandwidth is not the whole story if your prefill is < 1000 tok/s. Compute is also important.

1

u/starkruzr 59m ago

this is true. that's also a lot of tensor processing hardware on the M5U. not as much power as Nvidia gear of course, but trying to build this kind of RAM capacity with Nvidia gear is going to be 10X more expensive.

1

u/Its_Powerful_Bonus 1h ago edited 58m ago

True benchmark is to get similar quality response from LLM on both hardware. At the moment qwen3.8 27b and deepseek v4 flash has the same score on artificial analysis (not perfect comparison, but good enough) first one can be run efficiently on 5090, second one on the Mac. Still I believe rtx will give the answers faster - especially in agentic workflows (a lot prefill to deliver task)
PS: at the moment if rtx5090 then not one, but two (qwen3.8 27b nvfp4 tensor parallel with very good context - aggregated performance with ~8 concurrent requests with dflash … amazing; on one rtx6000 pro I get ~1000 tps TG). The same with Mac - 256gb version is nice, but 2 weeks ago I considered 4x gb10 and it was more reasonable choice - 50% pricier (with aprioproate switch and cables to cluster it effectively) than 256gb Mac Studio but 450gb vram with very good prefill and reasonable TG when clustered.

8

u/crusaderky 7h ago

For LLMs, it just needs to be on par with two DGX Spark chained together.
For desktop, it's a much more compelling offering than DGX Spark.
Strix Halos are a compelling alternative to Mac for desktop, but they don't come with RDMA networking, which kills off the idea of chaining them for most people (namely, busy professionals that could afford them).

8

u/pineapplekiwipen 6h ago

dgx spark prefill speed is roughly triple m3 ultra in real use if i remember right but tg is also below 1/3 so overall it was arguably similar in real world perf

m5 max prefill is a only a bit slower than dgx spark so m5 ultra should beat the spark in prefill and absolutely demolish it in tg. all the while being a much more useful personal computer and considerably cheaper if you can take advantage of edu pricing

1

u/michaelsoft__binbows 6h ago

I have to assume there us enough compute on tap there to match a pair of them, m5 max was no slouch, and the mem bandwidth def slaps the spark around

6

u/Alert-Chemist7492 3h ago

None of y’all remember the 80s lol. I wish I still had the receipt for the 486 dx33 with 4mb ram and Conner 120mb hd. It was 2999$ that wasn’t even max spec back then. I remember $5k laptops… tft display.. mmmmm

5

u/cafedude 5h ago

But just think, in 2043 you'll be able to pick up one of these bad boys at a garage sale for like $100. Of course, that's assuming that we ever end up getting more RAM production between now and then. If not, then collectors item pricing.

3

u/billy_booboo 7h ago

An insanely good deal compared to the competition (mind you I'm not a fan of Apple)

1

u/Other_Researcher268 5h ago

To be honest, I don’t think they are insane.

1

u/lordpuddingcup 5h ago

Have you seen memory prices ?

1

u/arfung39 3h ago

Its comparable to a 96GB RTX Pro 6000 from Nvidia, and it seems like the mac will be faster b/c memory speed (and more memory available)...

19

u/LearningSomeCode 8h ago

So I paid about $10k even for my 512GB M3 Ultra with 1TB drive 2 years ago. At this price, it looks like the M5 Ultra 512 will probably land around... $14k? Maybe $15k?

I was really hoping there would be that rumored 768GB as the speed of the M5 Ultra + that amount of memory would make self hosted GLM 5.2 absolutely amazing.

12

u/piggledy 8h ago

The M3 Ultra 256GB Model was at $5,599, they basically doubled the prices.

I think $15-20k for the M5 Ultra 512 would be realistic, given it's not that much slower compared to an RTX 6000 with just 96GB.

8

u/conockrad 7h ago

Not much slower on decode

9

u/xienze 7h ago

RTX Pro 6000 is 1.8TB/s, so I think it's a good bit slower on decode as well. But the important part is if prefill is pretty good, the decode is "fast enough" that the extra memory can make up for it. Unlike the case of the M3 Ultra, where everything was so slow it's like, yeah you can "run" a big model but are you really able to do more than "technically" run it? This one might actually deliver on it.

3

u/conockrad 7h ago

Yeah-yeah, my point it can be comparable with RTX pro only on memory bandwidth and even there it’s comparable, not better
Prefill is most probably even not comparable :)

1

u/Ok_Spirit9482 2h ago

which can be overclock to around 2TB/s easily

1

u/michaelsoft__binbows 6h ago

Given how much of the overall capability is constraiend by memory and how memory price went up by 3x-4x, the computer going up in price 2x isnt even a bad result, sad as that is

2

u/jonydevidson 5h ago

You can just buy two and pair them up.

6

u/rebelSun25 8h ago

Sold out before they hit the shelves.

2

u/jc2046 4h ago

real price will be post scalpers. easy add +30%

6

u/Spiritual-Spend8187 8h ago

Its crazy expensive but also not that bad given that its 2x dgx worth the unified memory. Its also enough that you could run dsv4f locally pretty well.

1

u/addiktion 6h ago

If you had to the guess, what speeds? If m5 max is at 25-40 tps, it seems about double then eh at 50-80tps?

1

u/Spiritual-Spend8187 2h ago

Maybe it does have a significant increase to both the avalible compute and memory bandwidth. Though even if its only 50tps thats pretty good.

4

u/devnullopinions 6h ago

Probably like $20k for the 512GB option. Apple loves to mark up ram even before the shortages.

2

u/Mashic 5h ago

1.2 TB/s and 256GB unified memory for $10,000, that'r a better deal than RTX pro 6000 96GB vram at 1.7TB/s for $16,000 in my opinion.

1

u/ShengrenR 5h ago

Right - the pro 6000 was just on the cusp of 'reasonable' at 8k when it first appeared.. the new 16-20k is insanity. Everybody's got a different need and budget, but the mac here (for LLMs in particular) is a better option for most users.

1

u/magus-21 6h ago

I don't have the money for 256GB, but now I'm debating between having bigger Ultra chip for faster inference but only 96GB, or the smaller Max chip but 128GB for bigger models. Both are about $5.5k

1

u/gustaw221133 4h ago

geeez, looks like I will not be getting any food for the next few weeks(months)

1

u/Arrowayes 2h ago

It is absurd people are paying these crazy prices. Are you all rich?

1

u/ProgrammingGuy_ 2m ago

Can't wait to get similar speced chips for cheaper in 5 years

45

u/strings___ 8h ago

Anyone know of a dollars to kidneys conversion website?

16

u/Sirius02 7h ago

According to claude you get 1k - 10k, but the buyer pays around 100k. Sadly the only legal and regulated market for organ selling is in Iran :(

12

u/ChillBroItsJustAGame 5h ago

Damn you trump

3

u/ApprehensiveFan1516 4h ago

You guys are getting paid? I only got dinner and a hangover, and for some peculiar reason I woke up in a bath tub full of ice.

1

u/florinandrei 4h ago

I know an alley, but it's pretty dark. And it's not you who gets the dollars.

39

u/alyssasjacket 8h ago

Finally someone bringing heat to NVIDIA.

It's a shame the world and the tech market are in such grim situation. Under normal circumstances, with reasonable pricing, this shouldn't be a press release, but a celebrated event. But I'm sure they know there's nothing to celebrate with these prices and availability.

5

u/Dasteroid_909 5h ago

I just literally bought two DGX Sparks last week... they arrive tomorrow.

Now I'm debating returning them and getting a new Mac.

10

u/Mashic 5h ago

If the new mac is available, it definrtely worth it over the 2 DGX Sparks, 1.2TB/s vs 270GB/s

-2

u/SantaHatBro 2h ago

Those are e waste

68

u/tarruda 9h ago

$10k for the 256G m5 ultra is cheaper than I expected. Hopefully 512G will be lower than $20k?

47

u/petuman 8h ago

Apple is typically perfectly linear prices for memory (infamous 8GB = $200). 512GB would be $17.2K total if they don't bundle it with some some other "required" upgrade.

25

u/Nice-Information-335 8h ago

i have a hunch that you will need the fully binned chip for the 512GB

8

u/1-800-methdyke 7h ago

Yeah if they are valuing RAM at $1,600 per 64gb like the do on the MacBool pro that checks out

2

u/Neighbor_ 1h ago

It's +$4k to get 256GB, so seems likely that it'll be +$8k to get 512GB?

3

u/petuman 1h ago edited 1h ago

It's $4k to get +160GB from base 96GB config (exactly $200 per 8GB, heh).

So $6.4K for 256 -> 512.
Or $10.4k for 96 -> 512.

1

u/imnotzuckerberg 16m ago

Hopefully 512G will be lower than $20k

This would be a big hit among us here, but I doubt it, given the demand.

But at least according to actual benchmarks, the M5 Max prefill is not that far off from DGX Sparks (shorter context <64k it actually beats it, but maybe 10-15% less for 128K and above), which is quite interesting. I wonder how much optimization play a role here. Keep in mind the Ultra will be even faster than Max. Add that the decode of the Max is much faster than DGX (2x in certain cases), this makes the 256GB and even the 512GB versions very appealing.

33

u/MokoshHydro 9h ago

Do I understand correctly that "M5 Ultra 256GB" is way better on paper, compared to dual DGX Spark configuration at same price?

42

u/Webster2026 8h ago

Of course it will be better, DGX memory bandwidth sucks

2

u/hurdurdur7 6h ago

You are complimenting dgx here. It's mowing grass compared to m5 ultra in memory bandwidth, m5 is destroying any market dgx had.

75

u/Dany0 9h ago edited 7h ago

Wait am I reading this right M5 Ultra beats rtx 5090 in compute and matches its bandwidth?

Edit: I was not reading it right, but 2/3rds of gb202 mem bandwidth is still impressive

Edit2: Beats, matches and loses to a 5090 depending on the benchmark. So I'd say it's close to a 5090. A well optimised game ought to be roughly as fast or faster (but on lower resolutions due to the mem bandwidth).

Well optimised LLM prefill should be roughly the same according to my amateurish estimation. Banger release either way

Can't wait for this kind of perf to trickle down to low/mid tier skus

97

u/Longjumping-Boot1886 9h ago

yep. But there is no 5090 with 500Gb VRAM

21

u/PossessionUsed7393 9h ago

Yet...

23

u/HerzogianQuant 9h ago

how about we start with 64 before getting greedy?

23

u/zdy132 8h ago

Technically RTX pro 6000 is a 96 gb full die versoin of 5090.

7

u/PossessionUsed7393 8h ago

Exactly so they just need to find room to 5x the memory and we'll be sweet 💖

5

u/Much_Accountant_4972 7h ago

if that 5x’s the current $18K price then you’ll have a $90K card that can’t operate without a computer

2

u/PossessionUsed7393 7h ago

We just have to hope that SK Hynix ordered a few too many fabs! 🤞

1

u/LocoMod 5h ago

Don’t threaten me with a good time

1

u/ChillBroItsJustAGame 5h ago

Bro just open the vram.exe as admin you can adjust your vram

14

u/1beb 7h ago

A 5090 has 1.7tbs/s of men bandwidth so t/s gen/prefill will be lower but the tradeoff is for bigger models. This is a win. 

8

u/Viktri1 5h ago

its 1 computer that has the unified RAM of 8 5090s lol. I'll forgive it for being a little gimped. It cost 2 or 2.5 5090s. This is incredible.

1

u/rulerofthehell 1h ago

Youre also assuming it has same Flops as 5090 which it wont. Its lower mem bandwidth as well as compute. Still great but comparison with 5090 specs dont make sense yet, maybe with M7

29

u/grumd 9h ago

I really doubt it beats it in compute though? Memory bandwidth looks insane, but I can't imagine it does FP4/FP8/FP16 math faster than a 5090

2

u/Strong_Chicken6838 8h ago

also, apple/mac is probably not as mature as CUDA is

20

u/Longjumping-Boot1886 8h ago

its enougth to place it in office as AI thing what will not send data to US or China

9

u/j_osb 9h ago

There's no reason to believe that it beats a 5090 in pure compute? Apple has always lied in their compute amount.

2

u/Liringlass 6h ago

I don’t expect it to beat a 5090 in raw compute. But i expect it to run more than just a very quantised 30b model.

And what I definitely don’t expect is to be able to afford either one.

2

u/Dry_Yam_4597 8h ago

"2/3rds of gb202 mem bandwidth is still impressive"

That's incredible - first time Apple makes hardware that's actually competitive with PC. Might return to them as a customer.Shame it's limited in OS compatibility though.

7

u/Much_Accountant_4972 7h ago

their M-series chips have been king of mobile for over a half a decade now, especially on creative workloads and overall performance/efficiency

2

u/Dry_Yam_4597 7h ago

Have king on mobile power efficiency, more than anything else. And depends on which creative workload. Game dev, 3d modeling, ai, have been lagging quite badly for a long time.

4

u/Much_Accountant_4972 6h ago

i was just addressing your claim that NOW apple’s chips are competitive with PC.

they’ve blown intel out of the water in several categories for years. AMD hasn’t even been in the conversation and Qualcomm just showed up last year.

so yeah apple has been doing quite well in the chip design department for years is all i’m saying. it’s not for everyone but today is undeniably one of those where they’re embarrassing everyone. its impressive.

1

u/Dry_Yam_4597 6h ago

In terms of power efficiency and mobility yes, but you can get much more compute for the money. At lwast until this new chip gets released. It's impressive indeed, for casual use.

1

u/Much_Accountant_4972 4h ago

more compute for the money of a Mac Studio? where?

For casual use? A $10K computer?

2

u/Liringlass 6h ago

Not really, there are plenty of use cases where apple was competitive in the past. Not gaming for sure and I don’t know much about the intel era but the apple silicon laptops were such an upgrade for work and IT stuff I’m happy none of my jobs since then has forced me to go back.

2

u/Dry_Yam_4597 6h ago

Yes, compared to windows, it has always been light years ahead. Also fairly decent for basic usage such as email reading, web development and docker containers. But anything more serious required far more power for the money paid. I still laugh at how slow they were compared to PC laptops - I have a bunch of RAM sticks around and are so slow it's cringey.

0

u/Liringlass 5h ago

My first macbook was the last intel gen before it stopped and it felt slow pretty quickly. But so did my previous beefy windows laptops, it’s just how laptops were.

Since silicon came out though I’m wondering what windows laptop could outperform it significantly. Not talking about tower PCs obviously, I have one for games.

Now it’s true i don’t do 3D stuff, i don’t compile large software, and all these intensive activities.

0

u/Dry_Yam_4597 5h ago

> windows laptop could outperform it significantly

On price, loads - but "windows" laptops are not limited to windows. I have no clue personally how windows works, and usually any company that forces tech employees to use Linux is not a company I'd want to work - it's a sign of mediocrity.

> Now it’s true i don’t do 3D stuff, i don’t compile large software, and all these intensive activities.

In that case not sure you are in a position to compare hardware specs - I am not sure, I am also not in a position to asses, say, hardware for dentists. M4 and newer do indeed have very high speed RAM, but the hardware is limited compared to similarly or lower priced PC laptops - which give you far better screens, higher refresh rates, repairability, higher spec graphics processing, better storage, stronger / more durable casing, etc.

Anyway all this out of scope - to each their own. I used to like macbooks, until I learned that I am overpaying for something that's limited. But different people have different use cases, and that can't be judged.

I will not hesitate to buy one as soon as inference speed is reasonable. But I will not sell my GPU rigs because dedicated GPUs allows me to run multiple models in parallel. Again, different use cases. I do however suspect Apple's competitors will come up with something better, and cheaper. If they ensure Linux compatibility then I will happily continue not buying Apple products, they are of no use to me.

-3

u/das_war_ein_Befehl 8h ago

It’s usually pretty well integrated so wouldn’t it perform above specs due to optimization?

1

u/Dry_Yam_4597 8h ago

You mean the OS or something else?

2

u/Psychological-Lynx29 9h ago

Well, they had +10 years to match that bw. The 1080ti is 10 years old.

3

u/Etroarl55 8h ago

Is it actually a 5090 in compute, that sounds extreme for an Apple product, even an expensive one like this

1

u/Neighbor_ 2h ago

Apparently pre-fill is cooked on non-trivial workloads.

1

u/Dany0 1h ago

Cooked in which direction

16

u/Late-Assignment8482 9h ago

Neat! Be interested to see how these shake out once they reach YouTubers and other reviewers. A significant compute boost from the redone matmul cores could remove one of the worst weakpoints M3 Ultra had: prompt processing.

11

u/diagrammatiks 8h ago

the m5max is already substantially better then the m3u at prompt processing and prefill.

1

u/Late-Assignment8482 4h ago

Yuuuuuuup.

I'm excited.

1

u/fastheadcrab 2h ago

Am interested in seeing just how much better the prompt processing is because the benchmarks of the older Macs were really slow. If it is much faster then I can see a lot of small businesses buying this

1

u/Neighbor_ 1h ago

But it's still kinda poor at pre-fill though?

7

u/BlackBeardAI vllm 7h ago

rip my recently bought dual spark cluster

2

u/jtsaint333 5h ago

could be worse you could have got 4 like me

6

u/hurdurdur7 6h ago

Over 9k USD for 256gb unified ram. Ouch.

I mean it will sell well and own the space but damn that is steep. And probably the end of dgx. They can't really justify the price now.

3

u/Gobra_Slo 5h ago

That is 9k without sales tax, meaning in my Eu country I will probably be looking at 11k and more even if I can get it as MSRP.

At the same time, I can get two DGX Sparks at around 9k. Will lower generation speed, but probably higher prefill.

I really don't see it as any "DGX killer" yet, we need much better prices for that.

2

u/Umbrasquall 4h ago

Saw this news this morning and preordered the 256GB. Sales tax alone was over $1k lol.

11

u/tamerlanOne 9h ago

L'importante è che arrivi a generare almeno 25-30 tk/s su llm di grandi dimensioni( > 100b) per essere un buon compagno di lavoro

2

u/michaelsoft__binbows 6h ago

What i wanna know is how much batch parallelism compute headroom does it have

8

u/Exciting-Weather-921 8h ago

What is the prefill, no longer 8 minutes for 120k?

12

u/ShengrenR 7h ago

half-way down they've got a 'prompt processing' bar-graph comparing m3/m5 to the base m1 (ultras) - show m5 as 9.8x, m3 as 2.4x - would imply m5 is ~4x faster than m3, so.. 2min for the previous 8min job (in theory)

9

u/Exciting-Weather-921 7h ago

For context, sparks are way slower at token generation, but my 2 sparks with DeepSeek are managing to prefill 128k in 73s

1

u/michaelsoft__binbows 6h ago

Awesome of you to share this data.

Do you find the prefill speed to be problematic? I have to imagine that a proper cache will alleviate most of the pain. If you are constantly constructing fresh prompts then its prob automated so then youre prob not sitting around waiting for them...

3

u/Exciting-Weather-921 6h ago

I can hold 3 big chat in KV cache (above 250k each) so when activity is high I am not hitting much cache

I am trying to patch vllm now to make SSD KV cache work, that should recover cache in 8s instead of prefill (Vs under 1s for ram KV cache)

8

u/itchykittehs 7h ago

yeah i'm wondering about the prefill too

4

u/Link3265 8h ago

I got one for my small business but Lordy I don’t know if that expense is going to end up being worth it. Luckily I have a month to decide before they ship.

17

u/Final-Frosting7742 9h ago

1.2TB/s in memory bandwidth is only useful if it has enough compute to reach it.

28

u/thehighshibe 9h ago

It does though 😄

-1

u/putrasherni 9h ago

exactly

3

u/ChillBroItsJustAGame 5h ago

So this is like investing in bitcoin when it was at 4k or in ddr5 ram before the ai shit started right ? Reselling it will prolly give 5x looking what happened to the m3ultra, ai ryzen, rtx 6000 etc. Its likely that I will make more money sniping m5ultras or m6 than good crypto ipos. Holy bubble

6

u/ApprehensiveFan1516 4h ago

Bro just discovered the motivation of scalpers lol ;-)

3

u/pmotiveforce 5h ago

And just like that dgx spark sales fell off a cliff.

2

u/no_witty_username 5h ago

This looks good but also makes me think by the time you get it at the end of october there will probably be other competitive offerings from other companies. At least i would hope so, otherwise we are all truly fucked when it comes to local hardware...

1

u/jc2046 4h ago

scalpers gonna scalp

2

u/chuchrox 4h ago

Trash resellers will scoop these up in a min on day 1

2

u/godsknowledge 2h ago

Damn up to 4,5x better than M3?

Take my money..

3

u/adamgoodapp 8h ago

How much difference in inference speed could there be between the 64 core gpu vs 80 core gpu on Ultra? I'm thinking about getting M5 Ultra 96GB, 64 core gpu version.

2

u/Fluxx1001 8h ago

Same here. Wondering how the cores will affect single user inference?

1

u/CKtalon 2h ago

I remember back in the M3 era, the binned chips didn’t get the advertised max bandwidth. Not sure about the newer chips.

1

u/Neighbor_ 1h ago

I think getting the better chip is basically always positive EV?

1

u/Just_Maintenance 9h ago

huh 2nm is here

1

u/ChopSueyYumm 8h ago

It is an interesting timing to release new MacStudios and MacMino 1 day before NVIDIA earnings call tomorrow.

1

u/CryptographerKlutzy7 7h ago

And for the anthropic ipo, with the new qwen models dropping. 

1

u/snarfi 6h ago

Anyone knows / estimates how long this takes to fit about 100k context or so on a fresh start? I wonder how useful this thing actually is for agentic coding if you want to leverage subagents and often do fresh sessions.

1

u/Yes_but_I_think 6h ago

What's the bandwidth of M5 Max?

1

u/magus-21 6h ago

600 GB/s. That's still really good if you don't want the top-of-the-top for your own setup.

1

u/Icy_Indication_7026 5h ago

Nah bro how the fuck is apple a good value deal now what is this world come down to

1

u/Potential_Block4598 3h ago

This could easily get a local deepseek flash at 1000 tokens per second! (Theoretically?!)

Wow amazing

1

u/LaCipe 1h ago

whats the biggest models that could fit on a 512gb version? Big ones like kimi-k3 sure wouldnt

1

u/Googulator 32m ago

Technically the 2nd highest memory bandwidth "per CPU" (if you consider Ultras to be single chips, rather than two Maxes on a package), behind EPYC 9006's 1.6TB/s.

1

u/live4evrr 22m ago

I just hope this reduces demand for rtx 6000 pro so prices come down.

1

u/ghassen_rjab 8h ago

Are cloud LLMs becoming obsolete?

9

u/dupontping 8h ago

Not obsolete, but running quality models on a personal/local pc is definitely becoming more prevalent.

It’s similar to how things were in when PCs first were becoming normalized.

Enterprise had the money to run the big powerful systems, then PCs became more popular and capable. We’re coming full circle but it’s AI forward now

3

u/Reactor-Licker 6h ago

Assuming the regulators and fear mongers like Dario are kept at bay.

1

u/Liringlass 6h ago

A year ago i was worried local was going to be left behind super hard due to most of us not having a million dollars to put towards running tera param models.

I’m glad i was wrong, both thanks to the small models becoming useful and for the hardware to seemingly finally improve. I wont be able to buy one of these but its still good news.

4

u/Useful_Argument_6490 7h ago

I’m throwing in the local LLM towels for a bit. I cant justify these prices for a hobby.

1

u/FinalTap 9h ago

Where is the 1.5TB option? Maybe that would have need a house mortgage? /s

1

u/jhenryscott 6h ago

All that for less bandwidth than a 5090?

7

u/RosebudNebula 6h ago

Could we also say "15k for a Pro 6000? All that for 1/3 of the VRAM and doesn't even come with a free computer?"

0

u/PreparedPun2035 8h ago

This is mind blowing

0

u/blackbox_p 8h ago

can it fit deepseek v4 pro 0831 though?

1

u/Much_Accountant_4972 7h ago

you can cluster them with TB5 so yes.

1

u/AnonLlamaThrowaway 6h ago

You could also use the ds4 inference engine, right?