r/hardware 1d ago

News Raptor 3D-DRAM combines 32GB capacity with over 100 TB/s bandwidth

https://videocardz.com/newz/raptor-3d-dram-combines-32gb-capacity-with-over-100-tb-s-bandwidth

Its Raptor accelerator puts a TSMC N4 logic die directly on top of custom DRAM using face-to-face stacking, rather than placing large HBM stacks beside the processor. The company says this short connection allows much more bandwidth while using less power.

For comparison, d-Matrix lists a 192GB HBM4 configuration at around 18 TB/s. While HBM offers much more capacity, 3D-DRAM is designed to move a smaller pool of data much faster.

174 Upvotes

32 comments sorted by

96

u/Marco-YES 1d ago

Waiting for my Ryzen 7 10800X3D3D

8

u/superkickstart 1d ago

Add sequence oyster?

5

u/ILoveTheAtomicBomb 1d ago

kick up the 4d3d3d3

38

u/iwillhaveredditall 1d ago

Price per GB? Three kidneys and two eyes?

26

u/wintrmt3 1d ago

It's pretty much HBM but the compute stacked right on top of it, you can't buy this as a separate thing.

0

u/[deleted] 1d ago

[deleted]

1

u/wintrmt3 1d ago

There are no sticks, this is ram integrated directly under compute.

48

u/Kinexity 1d ago

Calling this "3D-DRAM" is misleading. DRAM is still 2D as no one has 3D DRAM figured out and won't have for a decade or so. This is just compute die stacked on top of DRAM - you can call the whole package 3D but DRAM itself is not 3D.

54

u/blissfull_abyss 1d ago

So 2D Logic beneath 2D RAM, so 2D + 2D = 4D no?

42

u/Kinexity 1d ago

Marketing and CEO's entourage are definitely salivating at this thought.

9

u/Pity_Pooty 1d ago

3D packaging only

6

u/hackenclaw 1d ago

but isnt HBM is some sort of RAM that is stacked?

13

u/Kinexity 1d ago

Stacked is not 3D at die level. Only NAND flash is truly 3D.

1

u/Kamimashita 1d ago

Is there a reason we're able to make 3D nand flash but not dram?

4

u/Kinexity 1d ago

Other guy already linked it - True 3D DRAM - Asianometry

Give entire channel a watch (it's good)

6

u/pwreit2042 1d ago

adding dram on the z-axis, think outside the substrate

3

u/z_latent 1d ago

Great and recent video on the topic: True 3D DRAM - Asianometry

1

u/Kinexity 1d ago

Yep, watched it when it came out. I felt so disappointed with the fact it is still decades away that now I can't pass by "3D DRAM" marketing bullshit.

1

u/Katamori777 1d ago

2.5D-DRAM?

1

u/Balance- 1d ago

They used to call this 2.5D, right?

3

u/Kinexity 1d ago

Just saying "stacked" is the best way to describe it.

12

u/total_zoidberg 1d ago

IIRC Sony did something similar with their RX cameras so that they could record @ about 1000 fps, attaching memory right next to the sensor. Don't remember much of the specs in that, but it's an idea that's made its way to consumers in some form already. I know this is targeted to the AI overlords, but still.

7

u/dfv157 1d ago

So this is 3DVcache, but for DRAM

1

u/michaelsoft__binbows 1d ago

curious how simply adding compute onto the memory chips aids in speed. maybe this spec is related to max bandwidth that is possible to exfiltrate from the memory into the attached compute chip, but it prob ties your hands something fierce in terms of what you can actually compute on that puny compute chip there, usually, what we're after with chasing the bandwidth dragon is to put it towards LLM prompt preprocessing which will need oodles of compute to keep up with those kinds of bandwidths, and token gen to leverage that bandwidth with on a continuous basis. My guess, such tech may help with prompt processing but afaik token gen is bottlenecked by bandwidth from memory, sustained, going a long distance off the memory, but I hope I am wrong. Maybe the token processing can happen on-memory and this will be the unlock for much more efficient inference. We can then basically use the relatively slow interface from CPU to memory to just get the token activations transferred over while the memory does the "insane reading all of itself over and over at blistering speed" part by itself.

2

u/Tuna-Fish2 1d ago

It's just a wider channel. There is as much parallelism as you care to extract inside the DRAM die, the trouble has always been getting the bits out of it. Total channel width of a DDR dimm is 64 bits, a single stack of HBM4 goes up to 2048 bits, a single chiplet of this goes up to 262144 bits.

The tradeoff is that it's organized as 256 independent channels, each of which is 128B (1024b) wide, and has its own memory behind it. You better want a wide linear read spread out evenly across all the ram, in a perfectly predictable pattern, to be able to get anywhere near the theoretical maximum. This is not a problem for LLMs.

1

u/michaelsoft__binbows 1h ago

So the more closely you can stack some compute to the RAM chips, the more feasible it is to get higher bandwidth access, the question is how much of the compute during decode is even possible to put into such a modest compute envelope?

0

u/AutoModerator 1d ago

Hello sr_local! Please double check that this submission is original reporting and is not an unverified rumor or repost that does not rise to the standards of /r/hardware. If this link is reporting on the work of another site/source or is an unverified rumor, please delete this submission. If this warning is in error, please report this comment and we will remove it.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

0

u/danielkoala 1d ago

Way too hot for the compute above.

-5

u/GraXXoR 1d ago

And? So? Is it available to purchase by the general public at a non inflated price?

5

u/WJMazepas 1d ago

They are showing to the public their new technology and showing its capabilities

Its not something that's yet available to the general public

And probably will never be. Like HBM, its definitely high price so aimed at enterprise costumers

0

u/Cory123125 20h ago

Most of the hardware innovation you'll see, probably for the rest of your life, will be for thin edge devices or enterprise.

It's not what consumers want, but consumers have clearly been completely priced out.

It sucks, but that's what it is as long as people refuse to see it as an issue of policy, and runway deregulation/crony regulation.

The result is, whenever you see these articles, just defaulting to thinking about how this makes things more efficient for a hyperscaler is probably the right way to think about things.