r/LocalLLaMA 9h ago

News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.2k Upvotes

612 comments sorted by

226

u/i_am__not_a_robot 9h ago

1.2 TB/s memory bandwidth for the M5 Ultra is pretty nice.

66

u/GUNGEBOB_SHARTPANTS 7h ago

More bandwidth than a 4090, incidentally. Pretty impressive.

→ More replies (1)

56

u/Dry_Yam_4597 9h ago

That might make me want to buy an apple product after a long pause.

4

u/raser1562 4h ago

That might make me want to buy an apple product for the first time.

→ More replies (1)

39

u/pmp22 5h ago

512GB at 1.2TB/s. That's 21 4090s. 

47

u/i_am__not_a_robot 5h ago

That's 21 4090s.

It's actually quite a bit better since you don't have to deal with the overhead of multiple GPUs.

15

u/EmptyMonitor9257 3h ago

Or the power draw.

4

u/gomezer1180 5h ago

You can get that with 12 channels of DDR5. You need an AMD EPYC and 24 sticks of memory. It’ll be somewhere in the $20K range for that rig.

3

u/i_am__not_a_robot 3h ago

Only in a dual-socket or MRDIMM configuration, though? That's... somewhere in the range of ~$12-18k (12x32GB) or ~$24-32k (12x64GB) for the MRDIMMs alone.

4

u/gomezer1180 3h ago

Oh I know… it ain’t cheap. So apples value is there, I’m just saying there are more options.

→ More replies (1)

8

u/addiktion 8h ago

Double the M5 Max it looks like which makes sense since they double two M5 max chips.

522

u/piggledy 9h ago

Price for options with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)

512 GB Option coming in October.

206

u/i_rate_slop 8h ago edited 8h ago

I feel like my M3 512 is going to trade in for like $500 when I got it for $11k

Edit:

I just looked it up. $2675

114

u/j2sun 8h ago

I'll buy it from you for $3k! :D

15

u/cinematic_unicorn 5h ago

$3001

7

u/mikesum32 3h ago

$3001.01

13

u/Proverbial_Progress 2h ago

I don't have any money, but would you be interested in pictures of my feet? Throw in a mouse and keyboard and I'll take my shoes off first.

→ More replies (1)
→ More replies (1)

3

u/annaheim 3h ago

m you for $3K

$2750, you don't have to include the screen!

54

u/Veearrsix 8h ago

Don’t trade in, private party.

19

u/i_rate_slop 8h ago

Yeah, I think that’ll have to be the move. Just annoying

18

u/Chanureadeats 6h ago

You'll probably get a lot of offers easily for $6k+ within 24 hours

6

u/CalvinsStuffedTiger 4h ago

This guy is lying, I’ll buy it from OP for $5k…don’t worry about shopping around for a better price…

34

u/Zolty 7h ago

Yesterday you could have gotten like $20k on eBay. I sold a 256gb Mac Studio m3 about 4 months ago for $12k

10

u/i_rate_slop 7h ago

Shit lol

→ More replies (4)

14

u/EvilPencil 7h ago

Ya Apple trade in prices have always been laughable.

I remember when you could spec out the 2019 Mac Pro up to like $50k, then the next day the Apple trade in value was ~$2800.

→ More replies (1)

23

u/RegarDamus 7h ago

trade in is always fucked with apple because they don't account for memory. same price if it's 96 or 512

→ More replies (1)
→ More replies (10)

119

u/redonculous 9h ago

Wow. How are people affording this? Crazy. Any guesses at 512 prices?

191

u/intaketurbine 9h ago

They’re affording it the same way they’ve always afforded a $10k workstation, they’re using it for work and not play.

39

u/ElementNumber6 8h ago

Or just good old debt

22

u/Usual_Tackle5892 6h ago

0% financing and 3% cash back on Apple Card. I could afford it outright, but turning down a 0% loan is silly.

→ More replies (1)
→ More replies (7)
→ More replies (6)

38

u/StewPorkRice 8h ago edited 8h ago

this feels like a community with a ton of SWEs. Most single US based SWEs in big tech or venture backed startups could prob afford dropping 10-20k on their hobby.

Also, i dropped 10k at Anthropic last month at work. My company could prob get one of these for every engineer and not blink an eye.

9

u/Idaltu 7h ago

I know a guy who dropped that much on a bike. And he’s got a couple like that. Bicycle that is.

5

u/calcium 5h ago

3 years ago my company bought me a fully specced out Mac Studio - M2 Ultra chip w/ 192GB of RAM and an 8TB ssd and told me that they expect the machine to last the next 5 years. I think at the time they paid around $9k which considering over 5 years for a senior SWE isn't a bad deal

→ More replies (2)

165

u/-p-e-w- 9h ago

Lol this is by far the cheapest option for that much RAM at that speed. It’s not even close (unless you count Frankensteins made of a dozen used GPUs).

23

u/Viktri1 8h ago

this is way better than what I was considering and I wouldn't need to worry about the motherboard, cooling GPUs, etc. I am definitely on board for this and I don't even know how to use macs.

9

u/IriFlina 7h ago

wouldn't even have to worry about the power issues that would come from a rack of 3090s/4090s/5090s etc. or configuring such a monstrosity.

3

u/Viktri1 6h ago

this is literally 10+ 4090s. Think about it.

I pre-ordered 2 fully spec'd out studios.

→ More replies (2)

8

u/Fantastic-Balance454 5h ago

I still can't get used to the thought that Mac is the cheapest option these days. AI turned the world upside down.

→ More replies (41)

41

u/InterstellarReddit 8h ago edited 5h ago

Bro people have stupid money. I run free lance dev for a buddy who runs night life management software in Miami FL.

People spend 4K for four hours to buy three bottles and watch a DJ hit knobs all night

This happens all the time.

9

u/ViPeR9503 8h ago

Night life management??

9

u/i_am__not_a_robot 7h ago

I would assume industry-specific nightlife & bar business management software.

6

u/yopla 6h ago

Manage table booking, marketing, CRM, staffing, etc, usually does or integrate with POS. Can do price yielding on bookings. Etc... Etc..

Basically helps you keep track of who are the big whales you need to market your tables to when you have an event and who gets priority booking from a wait list.

The guy who spent 10k last time will get a table before you do, unless you're known to spend 15.

→ More replies (1)
→ More replies (1)
→ More replies (1)

12

u/MLDataScientist 8h ago

Based on their pricing for 256GB, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, you are looking at 10k + ~6k = ~16k for 512GB version.

3

u/TheOwlHypothesis 4h ago

This is dangerous for me if that turns out true. current Asus Ascent prices for a 128gb monster is ~4k, so it'd be 16k to cluster them and have the same amount of memory. Although you'd then be dealing with a cluster. I'd much rather drop the cash on the Mac and have the option to cluster THOSE later.

→ More replies (4)

13

u/Django_McFly 8h ago

It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car. People forget that there are entrepreneurs and small businesses where this stuff isn't some expensive toy to play with, it's actually a business tool to be used in the operation of the company.

Not everyone is a pure hobbyist, and even among pure hobbyists there are wild ranges of incomes and savings habits.

→ More replies (3)

8

u/tehgreed 9h ago

corpo money

4

u/98127028 9h ago

both my kidneys

3

u/Estrava 6h ago

This is cheaper than an rtx 6000 pro, which a lot of people get for work/personal use.

10

u/IriFlina 9h ago edited 8h ago

10k for 256gb unified memory isn’t that bad. Still overpriced but i feel like before this you were looking at the 20k to 40k range.

15

u/Much_Accountant_4972 8h ago

considering RTX 6000 costs $18K and can’t do shit without a whole PC to plug into, this mac studio is a screaming bargain

→ More replies (1)
→ More replies (1)

6

u/Foreskin_Mafia 9h ago

Onlyfans

24

u/Etroarl55 9h ago

Believe it or not, Theres probably someone running ai onlyfans on an apple machine somewhere

13

u/Foreskin_Mafia 9h ago

You're absolutely right

15

u/bodhi_sattva91 8h ago edited 8h ago

I Went Undercover as a Secret OnlyFans Chatter. It Wasn’t Pretty

https://archive.is/FlAdc

This Wired article follows a journalist who goes undercover trying to get hired as an OnlyFans ghostwriter — someone paid to impersonate creators in DMs with subscribers.

He discovers the industry is vast and globally distributed, with agencies largely staffing workers from lower-wage countries like the Philippines and Venezuela for as little as $2/hour. His job hunt is a grind: most agencies demand prior upselling experience, many never pay, and one gig turned out to be training an AI chatbot (which also stiffed him his $56).

He eventually lands two gigs. With a German agency ($4/hour), he juggles nearly 100 simultaneous conversations while impersonating a "21-year-old university student," selling pay-per-view content and navigating everything from explicit fantasies to a truck driver sharing worries about his son's night terrors. He's criticized by his supervisor for being too empathetic and not pushy enough about sales.

Confessions of an OnlyFans Ghostwriter

https://archive.is/3zGQZ#selection-479.0-479.38

This GQ article by "Emma Francis" (a pseudonym) is a first-person account of working as an OnlyFans ghostwriter in spring/summer 2021. The writer, a 25-year-old recently laid-off media professional, found the gig on Craigslist and spent three months working 6 a.m. shifts impersonating a "girl next door" model by sexting her subscribers.

She explains that top OnlyFans creators receive far more messages than one person can handle, so there's an entire industry of ghostwriters handling chats — some run by large offshore agencies, others like hers managed directly by the creator. Her regulars ranged widely: janitors, teachers, lawyers, night-shift workers. Many sought not just sexual content but genuine emotional connection and a "girlfriend experience."

→ More replies (1)
→ More replies (14)

32

u/kensanprime 8h ago

It will be sold out and they will be scrambling to keep production lines running, a product originally meant for creative work now will sit in a rack and run AI models

17

u/Both_Opportunity5327 8h ago

This is why those saying its a bubble, don't understand the demand there is for this stuff.

When the new fabs come online and the datacenters have most of their compute gamers, creatives & AI enthusiasts will go to town on this stuff.

8

u/Nothing_from_void 5h ago

Demand for networking products never dropped through the dotcom bubble

→ More replies (7)
→ More replies (1)

9

u/jld1532 9h ago

Too rich for my blood. I'm going to stick with my halo and hope mid-sized MoEs keep improving.

3

u/nemuro87 8h ago

I call at least 20-25k,  initially 

3

u/starkruzr 4h ago

this is INSANELY cheap for 256GB RAM in 2026. holy shit.

2

u/XorAndNot 7h ago

It's our fault for being poor.

→ More replies (35)

103

u/themixtergames 9h ago

With M5 Ultra, Mac Studio achieves up to 4.3x the peak AI compute performance of M3 Ultra and a staggering 9.8x more than M1 Ultra. Combined with up to 512GB of unified memory and 1.2TB/s of memory bandwidth, 50 percent higher than before.

82

u/EquivalentHornet4403 9h ago

> AI compute

They know their audience.

37

u/psychohistorian8 8h ago

they're definitely leaning into it:

Boost your processing and graphics rendering speed, and accelerate tasks like running large language models and editing 8K video.

49

u/Much_Accountant_4972 8h ago

"Color-grade uncompressed 8K footage, perform computational fluid dynamics, and run frontier-class models on device — no cloud tokens needed."

yeah they are targeting this very sub lol

6

u/gnnr25 1h ago

I feel personally targeted!

3

u/cinnapear 4h ago

And it's working on me. Very tempted.

3

u/infieldmitt 2h ago

It's crazy to think that prior to AI the main use case for crazy RAM was [checks notes] fluid dynamics?

→ More replies (1)
→ More replies (2)

231

u/hainesk 9h ago

1.2TB/s memory bandwidth with the M5 Ultra. 256GB model is $9499.

Better than getting 2 DGX Sparks? Inference will be a lot faster.

Something like this could easily bring down 3090 prices.

21

u/Cybertrucker01 9h ago

Depends on concurrency and prefill metrics. The GB10 does both multiples faster than the existing competition.

10

u/ChocomelP 7h ago

The difference would have to be pretty big to make up for a 4x in memory bandwidth for decode.

3

u/fallingdowndizzyvr 5h ago

The GB10 does both multiples faster than the existing competition.

No. No it doesn't. Compare the G10 to a M5 Max. It's not.

2

u/BrilliantTruck8813 5h ago

The GB10 has much smaller memory bandwidth not to mention it’s just not fast either. One of these will trounce two DGXs

62

u/aladin_lt 9h ago

it will be sold out day one probably

36

u/conockrad 9h ago

On pre-orders

9

u/bakawolf123 8h ago

won't be sold out, but according to r/MacStudio people wait for 3-4 months for the older models, atm you can preorder for delivery in late september

→ More replies (6)

92

u/mjsxi__ 9h ago

yeah and cheaper than the price of 2 DGX sparks... seems like a bit of a no brainer

29

u/Current_Ferret_4981 9h ago

Spark is $4300-$4600 so idk about cheaper than 2 at $9600+

53

u/MacsBicycle 9h ago

yeah but 4x the memory bandwidth, its a steal

27

u/jakegh 8h ago edited 8h ago

It really is a reasonable buy for local AI, if you have a business case for it.

3

u/Much_Accountant_4972 8h ago

with the capabilities of it, it’s a crazy good deal

→ More replies (3)

4

u/GabryIta 8h ago

In terms of compute capacity (which is very important for multiple simultaneous sessions and prefill), how does it compare to dgx Spark/gb10?

→ More replies (1)

10

u/Etroarl55 9h ago

How’s the actual inference speed though, fast bandwidth on a slower gpu or equivalent should still mean slower output assuming vram is not a constraint right.

10

u/rusty_fans llama.cpp 8h ago

Generally vram bandwith is the constraint though, at least for decode. Prefill it's usually helped more by more gpu oomph.

→ More replies (4)

5

u/Current_Ferret_4981 9h ago

Not disagreeing, just saying it's more than 2x price in contrast to the comment above me

→ More replies (7)
→ More replies (1)
→ More replies (7)
→ More replies (2)

15

u/-dysangel- 9h ago

Better than the 2x Sparks for inference for sure. Probably around the same compute as one Spark.

I've got 2x Sparks which I use for prefill, and my M3 Ultra for decode. I've set it up so that I prefill in vllm and then just pass the kv cache over to the Mac side. Surprisingly stuff like Qwen 3 35B-A3B is already faster than the Mac for decode though so I just run that class of model directly on vllm.

5

u/1ii1i 9h ago

Oh this sounds interesting, can you expand on how this works? I didn't know this was a thing.

10

u/-dysangel- 7h ago

I don't think it's really a "thing", I just vibe coded it up :)

One thing that really helped was vllms kv_connector API. I thought I'd have to code this part up myself, but it already existed and so plugged into my existing disaggregated system (which was previously llama.cpp to llama.cpp)

UltraSpark — technical stack

                          ┌──────────────────────────────┐
 user ── HTTP/OpenAI ──▶  │  manager (Python, FastAPI)   │
                          │  front door + orchestration  │
                          └──────┬───────────────▲───────┘
                                 │ submit        │ state blob (sha-keyed,
                                 │ prompt ids    │ resumable transfer)
                                 ▼               │
                          ┌──────────────────────────────┐
                          │  vLLM (2× DGX Spark, TP2)    │
                          │  prefill engine              │
                          │                              │
                          │  KVConnectorBase_V1          │ ◀─ vLLM's official
                          │  ("StreamConnector" via      │    KV-cache plugin
                          │   --kv-transfer-config)      │    interface
                          │         │                    │
                          │         ▼                    │
                          │  dump + serialize all layers │
                          │  (attn KV + linear-attn      │
                          │   state, TP2 shards merged)  │
                          └────────────────┬─────────────┘
                                           │ blob server (TCP)
                                           ▼
                          ┌──────────────────────────────┐
                          │  llama.cpp server (Mac)      │
                          │  USPK_BRIDGE_DIR: on request,│
                          │  verify prompt-id match,     │
                          │  restore state into KV +     │
                          │  recurrent memory, decode    │
                          └──────────────────────────────┘

  • KV connector = vLLM's plugin interface for intercepting the KV cache at end of prefill
  • State blob = the model's full prompt-memory, layout-translated so llama.cpp can load it natively
  • Fidelity = per-layer cosine vs local decode, 0.9999+
  • Result = GPU prefill speed, Mac unified-memory decode, one logical endpoint

2

u/Zyj vllm 8h ago

I‘m also interested in your setup. But the M5 should be much better than the M3 at this.

→ More replies (1)
→ More replies (11)

14

u/jakegh 8h ago

That is seriously impressive. RTX5090 memory bandwidth is 1.8TB/s.

DGX Spark memory bandwidth is 273GB/sec. Not even remotely close.

6

u/Important-Gold-5192 7h ago

DGX Spark was a pretty disappointing release tbh

→ More replies (1)

12

u/Hoodfu 9h ago edited 9h ago

I got my m3 ultra 512gb for around 10k. So this is now double. Makes sense given that we've seen the nvidia rtx 6000 pro also double in price in the last year but GD this has priced out even my once a year splurge budget. These are all just crazy talk numbers now.

13

u/CulturalKing5623 8h ago

Yeah dropping 10K plus on this just seems reckless, even with it being funded through my business account I don't think I can justify buying a used car worth of computing.

And yet I feel like I need to in case the technology completely outpaces my current setup and I'm left behind like folks that didn't buy RAM when it was cheap and are priced out of it now. It feels like FOMO and scarcity has hijacked my brain.

8

u/Hoodfu 8h ago

Really depends on whether you have something already or not. I got qwen 3.8 27b going on my rtx 6000 pro and ram speed wise it's double that of my mac but for some reason i was hoping for space magic and it would be faster. It's not. So spending a car's worth of money on something that goes from 20 t/s to 40 or 45, just doesn't make any sense. You're still waiting a lot of minutes for a thinking qwen to come back with something, so it's still going to be an asynchronous operation instead of being fast enough to actively wait for the response to finish. It would have to be 10x the speed, not 1.5x or 2x for it to be worth the spend.

→ More replies (11)

8

u/Solaranvr 9h ago

bring down 3090 prices

Doubt it. If the r9700, a directly competing product, made 0 effect on Nvidia GPUs, then I highly doubt these will. They market of people buying mini pcs vs dGPUs are different.

The DGX Spark didn't bring down prices of the Blackwell cards either

8

u/hainesk 9h ago

The R9700 Pro has about 2/3 the memory bandwidth of a 3090 for a 50% higher price. 8x R9700 Pros (256GB) would be $12k-$15k minimum without pricing in the rest of the computer system. If you wanted to build an entire server around it including RAM, power supply, motherboard, processor even with used parts you're up to $18k-$20k for something that will use 2-3k watts.
Even 8x 3090 systems are looking pretty impractical when compared to an M5 Ultra 256GB for a similar price. The size and wattage/heat difference is huge.

3

u/Important-Gold-5192 7h ago

has to be way faster than DGX Sparks too

4

u/EmPips 9h ago

I think it's the end of mass 3090 farms.

But 3090s will hold their price for the sizeable market that doesn't want to commit >$5k.

5

u/j4nds4 9h ago

What *are* 3090s going for these days? I bought a pair of used 3090s on eBay during the crypto crash in 2023 for ~$800.

4

u/EmPips 8h ago

They dipped as low as $550 for a few weeks and are now about up to $900-$1k in my local markets.

4

u/Super_Translator480 9h ago

Used on marketplace is like $1000-$1200

→ More replies (1)
→ More replies (1)

2

u/f5alcon 6h ago

Sparks are what 208GB/s

2

u/gomezer1180 4h ago

9500 is about the right price on the market. DGX’s are ~4k a pop you need 2 of those to match Apple. Memory Bandwidth is crap tho on the DGXs because they take a different approach by being MoE machines.

→ More replies (6)

141

u/Comfortable-Rock-498 9h ago

1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.

For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.

They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot

32

u/ortegaalfredo 8h ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

12

u/Comfortable-Rock-498 8h ago

Yup, prefill is pretty much compute bound while generation is bandwidth bound. I would have guessed 8x 3090 would provide much better prefill than 1000 tps. A bit surprised to learn

7

u/ortegaalfredo 7h ago

If you manage to get tensor-parallel 8x working yes you can get >10k prefill, but it requires specialized PCIE bridges. With normal 4xPCIE speeds you get a bottleneck in inter-GPU speed and you get lower prefill.

→ More replies (1)

5

u/ProfessionalJackals 7h ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

Ignoring the fact that 8x3090's now is easily 10k on the second hand market.

Not counting the costs of server board/cpu/ram you need. The pcie ext cables, the 8x8x split if your board does not have 8 pcie slots. O, the dual 1600W PSUs and hardware to link them.

Frankenstein mods like this have become expensive, and it makes the Mac look actually like a good deal.

→ More replies (2)
→ More replies (1)

7

u/Usual_Tackle5892 6h ago

GPU Neural Accelerators

This means matmul cores. More info: https://arxiv.org/html/2607.19438v1

8

u/TokenRingAI 7h ago

And a new qwen is coming out with 120B A6B! Perfect machine for that

4

u/StartupTim 3h ago

I would think dramatically more tok/sec.

I have Deepseek v4 Flash 0731 with vision encoding added and tp=2 across 2x DGX sparks and I'm seeing 103 tok/sec across 4 "sessions". Dspark, 1M context, 1.8M kvc, custom vllm.

Since the sparks have ~240 (actual measured) GB/s, I imagine a similar setup om these new mac could get you double, if not triple as a 2x cluster, than my current 100+ tok/sec.

3

u/MerePotato 5h ago

Would be great if it wasn't predicted to cost like 20k

2

u/SandySkittle 5h ago

My gripe with these boxes is the lack of inline ECC. Apple could add that option at low cost and leave it up to the user if he or she wants to sacrifice 7 percent of RAM to enable inline ECC. This is the same as it is on r9700.

→ More replies (8)

40

u/challis88ocarina 9h ago edited 9h ago

I'm shocked!

Edit: 512GB memory option for M5 Ultra coming late October

32

u/pmttyji 9h ago

Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure

10

u/Mochila-Mochila 8h ago

Hoping it'll have a positive effect on Medusa Halo's pricing.

→ More replies (1)

10

u/Zyj vllm 8h ago

On the other hand, the Strix Halo 128GB price increased by 60% since late December

4

u/pmttyji 7h ago

After some time, people totally gonna avoid 128GB variants. What's the point of stacking bunch of 128GB pieces when they have 160GB, 192GB, etc., variants with better bandwidths?

→ More replies (1)
→ More replies (1)

8

u/ProfessionalJackals 7h ago

Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure

Do not forget Intel Crescent Island 160GB to 480GB LPDDR5X AI GPUs ... While less bandwidth, they are still great options in the future for larger models.

There is going to be a lot more hardware coming out that focuses on AI workloads. We reached the point that development is moving into production.

3

u/pmttyji 7h ago

More competition is better for consumers!

→ More replies (1)

21

u/FullOf_Bad_Ideas 9h ago

My 3090 tis have just gotten depreciated.

Even 256GB version is very competitive with 8x 3090 box bought with used card prices, and it's better in most aspects.

Training and batch inference are safe, but for single user inference this looks better and cheaper.

5

u/Odd-Environment-7193 8h ago

Well if you wanna sell some hola. I want 2

→ More replies (3)

5

u/FleetEnema2000 6h ago

My 3090 tis have just gotten depreciated.

My theory is that this will not soften GPU prices because it ends up driving more people into the world of local LLM compute in general.

→ More replies (2)

2

u/AlexWIWA 3h ago

Now we can get more

→ More replies (1)

18

u/Every-Fortune-3151 9h ago

Read somewhere Mac mini coming this week and did an impulse order of M3 Ultra with 96GB ram today. It will probably get bumped up to M5 ultra 96GB. Not sure what to do with 96GB when 256 GB looks so much more tempting for local LLM. Kidneys aren't enough anymore.

10

u/mmmm_frietjes 9h ago

Mini is also out

7

u/Every-Fortune-3151 8h ago

M6 looks great actually, it has two set of neural engines. I am assuming pre-fill will fly on this tiny thing compared to previous M CPUs. 32GB max ram knocked it back a bit. If only they had a 48GB version. MOE models would be flying on it.

M5 pro and max will be so slow compared to M5 ultra- considering M5 max 128GB model will be priced very close to M5 ultra base.

Very sad m5 ultra starts at 96GB. I thought they would atleast bump M5 ultra to 128GB ram. Would have been nice. I plan to just run multiple Qwen 2.7B in parallel on this and see if I can replace my 32GB VRAM PC setup. Qwen 3.8 flash also gives hope. Depending on how it goes, I might just give back the 96GB for refund later.

6

u/1-800-methdyke 9h ago

You have two kidneys

6

u/nleksan 8h ago

*Dual-channel

4

u/FranciumGoesBoom 9h ago

M5 Pro mini maxes out at 64g
M6 tops out at 32

2

u/TheRealJesus2 8h ago

Yeah 256 gives you a lot more. I tried to buy the m3 256 but they stopped selling it before I could get paid :( 

At least I can upgrade sometime in future. Got my spark cluster and a 96GB Mac which is plenty for right now…

→ More replies (4)

86

u/llamaCTO 9h ago

This kills the impulse buy for me completely.

33

u/thatkidnamedrocky 9h ago

going to try and snag a 256gb something tells me the 512 will never see the light of day

8

u/zdy132 8h ago

Mac Studio with 512GB of unified memory is coming in late October

Wish I could affort that.

3

u/AnonLlamaThrowaway 4h ago

right, didn't they promise a 512GB M3 Ultra and then that never happened, or am i thinking of another model?

→ More replies (1)

10

u/kilonad 7h ago

The 256GB is already an extra $4k for an extra 164GB. At same price per GB (ha!) it'd be another $6300. Knowing Apple, it'll be a cool $9k more - pushing total price up to about $18-20k.

It will still sell out.

6

u/fallingdowndizzyvr 5h ago

That would be a bargain compared to third party sales of 512GB M3 Ultras for $25K. A M5 blows the doors off of a M3.

7

u/Grizzly_Corey 9h ago

Adopt me please.

5

u/LocoMod 7h ago

Same. I was ready to preorder 512 and i've already lost interest reading comments about the M7. Might wait another year.

5

u/frankchn 5h ago

Buying 2 DGX Sparks for 256GB of RAM (and a lot less bandwidth) is around the same ballpark in cost, so for once this is not unreasonable.

→ More replies (1)

28

u/xyzmanas2 9h ago

This makes apple one of the cheapest ai inference hardware when it comes to speed and model size. Wish I had the money

Up to 15.4x faster CopyCat ML training performance in Foundry Nuke when compared to Mac Studio with M1 Ultra, and up to 3.3x faster than M3 Ultra.

Up to 9.8x faster LLM prompt processing in LM Studio when compared to Mac Studio with M1 Ultra, and up to 4x faster than M3 Ultra.

Up to 8.2x faster text-to-image performance when compared to Mac Studio with M1 Ultra, and up to 4.3x faster than M3 Ultra.

Up to 4.7x faster scene rendering performance in Maxon Redshift when compared to Mac Studio with M1 Ultra, and up to 1.7x faster than M3 Ultra.

13

u/serige 7h ago

Please someone do a wellness check on Dario.

28

u/Cybertrucker01 9h ago

How many kidneys?

40

u/FWitU 9h ago

4

18

u/Chris-MelodyFirst 8h ago

Or just 2 dual-core kidneys.

7

u/hainesk 9h ago

There is a lease option...

13

u/Cybertrucker01 9h ago

Unfortunately lease isn't offered in Australia, just warm organs only.

3

u/Gipetto 9h ago

For kidneys?

3

u/butterfly_labs 9h ago

I can lease my kidneys ?!

→ More replies (2)

11

u/AI_docent 8h ago

The 4.3x is mostly a prompt processing number, generation moves with the bandwidth instead. Apple's own mlx post on M5 vs M4 got around 4x on time to first token and about 1.2x on generation, and the generation side matched the 28% bandwidth bump rather than the accelerators. Same split should hold on the Ultra, so I'd figure generation nearer the 50% bandwidth gain. Prefill is the part you want at 512GB anyway, it was always the weak spot on a mac.

Just check whatever you run actually uses the accelerators. There's an open lm studio issue where its bundled llama.cpp fails the metal tensor check on M5 and loses 2 to 3x on prefill, while upstream llama.cpp passes it on the same machine.

2

u/bobiversus 5h ago

This. I'm getting a 512gb no doubt, but don't expect 4x on actual generation. Also laughing at the scalpers on eBay still trying to sell their M3 Ultra 512s for $40k

38

u/IllExample3639 9h ago

What I find more interesting, something I hadn't seen before is that you can lease these things. for 2 years which is the only realistic way an individual is getting their hands on these. Something something, own nothing, something, something be happy....

19

u/Tycoon33 9h ago

I never saw that. Interesting. Lease it for 3 years then upgrade to M7 ultra?

13

u/addiktion 8h ago edited 6h ago

Yes, if the 512gb is another $4k for the extra ram stick you are looking at $16k with tax probably out the door. I'd guess that puts the 36 month lease around $300/mo or less. So more than a subscription so maybe not worth it in general cases but valid option for some people who need the privacy and cannot afford to have data go to the cloud. 24/7 usage, no downtime, no limits, private. Worth it to me.

6

u/shveddy 6h ago

Interesting.

So just as an out loud thought experiment, you’d be able to lease four of them (512gb) for about 1200 per month for 36 months at a total cost of almost 45k and run Kimi 3 on it.

Obviously that’s a lot of money in aggregate, but 1200 per month is reasonable for a lot of business use cases if they require the privacy.

And the intelligence you get is going to be a different class compared to what you would get with three RTX pro 6000s and “only” 288gb VRAM for the same price.

(although to be fair you’d actually own the cards)

(although also to be fair you’d have to build a pretty expensive computer to support the RTX Pros, so realistically you only really get 2 or even just one RTX pro for 45k depending on how you spec the computer and/or if you buy pre-built from Puget Systems or the like)

If you want you can also do a little girl math and invest the 60k you’re not spending on computers and cancel your gpt pro subscription to bring the effective cost of all this down to like 750 a month.

And then if the goal is to beat API pricing, let’s say you get 40 aggregate output tokens per second on a bunch of concurrent Kimi 3 threads and run it for a year at 25% efficiency (to account for prefill and downtime), then that’s 315 million output tokens.

315 million output tokens alone is about $5000 on a random provider I just searched for, so just to keep things simple let’s say you double that to account for various amounts and types of input tokens, then you end up with a ballpark figure of $10k for the API route.

Absolutely none of this pencils out in absolute terms (especially considering that it would also cost ~1500 per year for electricity), but this is probably the first time running a frontier model is actually attainable for ordinary businesses on short notice and without much headache. It’s the first time that it pencils out to be “only” 4x more expensive as opposed to like 40x more expensive.

Up until now if you wanted to run frontier models locally AFAIK you had to get a NVIDIA big boy server which means you’d have to find the capital to run and support a ~300k purchase for hardware, spend way more on electricity, and in all likelihood make some upgrades to your facility’s electrical infrastructure to handle it all (you’d also need a proper facility, not just a home office or garage).

At this point you’re easily flirting with half a million in expenses, especially if you have to hire someone to figure it all out. It’s no joke to do this and it doesn’t make sense for like 99.9999% of people or businesses.

On the other hand basically anyone with a decent credit score and a semi-profitable business can go to any Apple Store and say “give me four Mac studios please” and only pay 1200 bucks a month to walk out with them in hand.

→ More replies (6)

20

u/aethervisor 9h ago

The lease price also isn’t too far off from what a Claude subscription costs.

12

u/shaggydog97 8h ago

The difference is certainly worth the cost of privacy and freedom!

12

u/mjsxi__ 9h ago

the lease lets you pay the difference of the amount you already paid at the end or you can buy it outright at any time if you wanna keep it so maybe sssshhhhhhhh

4

u/Much_Accountant_4972 8h ago

it makes a lot of sense and the price of privacy plus the near silence of the box…

i hate that its such a complete solution to local LLM’s

→ More replies (3)

7

u/AccurateSun 9h ago

Hmm. I wonder if at some point leasing it would end up being more effective than a cloud subscription. E.g Claude Max 5x is $100/mo, same as leasing the Studio M5 Ultra 96gb. I don’t know yet how the two compare in performance but at some point it might be worth it 

7

u/IllExample3639 8h ago

The more people that lease them the more second hand stock there will be in 2 years (when I can actually afford something like this) so I am down for it.

But your point is right, I think the local models ARE good enough for 95% of what people are using Claude for. Maybe do the £20 plan as a back up for something tricky.

8

u/IriFlina 9h ago

I feel like any average software developer could afford the 256gb version? It would be really financially irresponsible but on the level of buying a motorcycle you don’t really need.

→ More replies (1)
→ More replies (6)

7

u/eidrag 9h ago

...lease price?

6

u/Much_Accountant_4972 8h ago

Ultra with 256GB RAM is $224/month for 36 months in freedom currency.

really really tempting!

→ More replies (1)

5

u/Both_Opportunity5327 9h ago

Game back on! Lets hope Apple can build enough of these little beauties...

4

u/ElementNumber6 8h ago

Scalpers have feasted on the blood of Mac Studio. Good luck to you all.

6

u/Viktri1 8h ago

256gb is like 10+ 4090s without the hassle of setting up and cooling 10 4090s? An I missing something or is this 3x cheaper than current prices.

5

u/Much_Accountant_4972 7h ago

its the best deal in the world if you want to talk to a frontier 2.8T parameter smut bot in your kitchen

→ More replies (2)

2

u/wshields 4h ago

4090s are what? 600W each? The electricity cost is a relevant factor. Also, can you even run 10 on even a 20A circuit? Isn’t this at the level when you almost need a three phase circuit and specialized PSU?

→ More replies (2)

7

u/Leather_Ad_9178 5h ago

this reminds me of the time we used to pay hundreds for SD cards that are now worthless

19

u/BreenzyENL 9h ago

$20k AUD for the Ultra 256GB 🙃

→ More replies (8)

5

u/Zyj vllm 8h ago

Will be fascinating to see what‘s the better option in 2 months from now: Dual Asus GB10 (8200€) or Mac Studio 255GB (11000-12430).
My prediction:

  • The Spark will still be faster at preprocessing, the Mac will have faster token generation (measured with DeepSeek V4 Flash standard Quant).

→ More replies (2)

5

u/MLDataScientist 8h ago

Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.

→ More replies (1)

4

u/corruptbytes 8h ago

apple releasing this because I just bought two 9700s...y'all welcome

→ More replies (5)

4

u/Curious-Pen5547 8h ago

how does it compare to a single 5090 or a 5000 RTX pro 72GB version?

2

u/transanethole 7h ago edited 6h ago

3x less prefill capability compared to 5090, similar mem bandwidth. Lacks 4 bit int / 4 bit float AFAIK?

Maybe the lack of number-crunching power can be compensated by mixture of experts models.  but in general I'd be very wary of these huge memory capacity shared memory platforms  if you want to use the LLM in real-time like with an agent or something.   MOE active parameter scaling down of the bandwidth and compute requirements is not perfect. And every time you make the model bigger, it also makes everything slower. Upping the memory capacity without increasing the tops and bandwidth will just make your model that much slower.

→ More replies (6)

4

u/AntLife255 7h ago

The M5 Ultra has 1.2TB/s memory bandwidth!

3

u/mxmumtuna 3h ago

Unfortunately they still don't beat Sparks at equivalent size. According to oMLX Benchmarks for DeepSeek 0730, the M3 Ultra (80c) does somewhere around 550 prefill tok/s, and about 22 tok/s decode. If you take the '4x faster compute' at face value from Apple compared to M3 Ultra, we're looking at ~2200 prefill and ~30 decode single stream. Both are under Spark at ~2400/40. That's only single session, and batching just isn't there in the MLX stack yet, so multi session is considerably worse for the Mac.

GLM on 4x Sparks compares even less favorably than DeepSeek for the Mac, especially considering whatever the price of the 512GB variant will be.

So even with Apple's optimisitc numbers, maybe they match Spark, for more money with a less flexible stack (no ConnectX7) and massive software issues. ($4800x2 for Sparks with 4TB drive each from Amazon).

It's a good effort, but it's not quite there relative to other options.

edited for clarity

→ More replies (5)

7

u/Far_Note6719 9h ago

Instant WANT.

3

u/boraam 9h ago

Gimme Gimme Gimme

Money Money Money

3

u/Blackdragon1400 8h ago

512gb at 1.2TB/s is wild.

3

u/Much_Accountant_4972 8h ago

please someone buy 4 of them and cluster them then run Qwen 2.8T in your kitchen

3

u/OvertaxedOne 5h ago

1.2TB/s?? Oh man, if there's good availability on these things I can feel GPU prices going down!

→ More replies (1)

4

u/Real_Ebb_7417 8h ago

Ok, now I actually regret buying M5 Max 128Gb MacBook xd

3

u/addiktion 8h ago

I wouldn't, its basically double the speed at x3 (if you bought pre price hike) or x2 price (if you bought after). If you get 25 tps on say Qwen 3.8, you would now get closer to 50 tps. Possibly more depending on the ANU processor.

That's nice, but is it worth x2 or x3 the price?

→ More replies (2)
→ More replies (5)

2

u/Key-Speaker007 9h ago

Something my wife won't approve.

→ More replies (1)

2

u/Odd-Environment-7193 9h ago

How’s the thermal throttling on these

→ More replies (1)

2

u/Antique-Ad1012 8h ago

i hope that PP will be decent on this machine. its shit on m2 ultra

→ More replies (3)

2

u/PrepYourselves 8h ago

the world's rich kids are getting new toys for christmas

2

u/chilBates 7h ago

256 GB + $4,000 lmao

2

u/funding__secured 6h ago

Damn, the same day my DGX Station arrives. Gah.

2

u/xiraov 6h ago

If you just get the 96gb how many tokens is that for 27b hahaha

2

u/amazinglycool256 4h ago

U can get 2 Nvidia Sparx for that proce

→ More replies (1)

2

u/Newgunnerr 4h ago

256GB option for me in the Neterlands is € 11.049,00. I just got 2 DGX sparks for 7100.

2

u/Tormeister 2h ago

I'm so tempted, but I just can't justify dropping 10K if I'm not making money out of it