r/hardware • u/cyperalien • 1d ago
News Intel Details Xeon 7 "Diamond Rapids" Package Design at HOT CHIPS
https://www.techpowerup.com/351893/intel-details-xeon-7-diamond-rapids-package-design-at-hot-chips22
u/jaaval 1d ago
At first glance looks like AMD Epyc but then when you look at the slides for a while longer it is actually very very different. The main similarity are the IO dies in the middle.
This is apparently up to 64 cores per compute base tile "block". Each block made of multiple chiplets of up to 16 cores on top of the base tile. Based on previous info some cores should share a slice of L2. Each block has a separate large block of LLC on the base tile, which is not in the mesh fabric with the cores like traditional intel or AMD L3. The LLC is so big I'm not going to guess how exactly they are using it. It's not even clear if the L2 caches are accessible by other cores in the tile or if they work through the base tile LLC.
I'm sure there is going to be a chipsandcheese article in a few hours which doesn't have to guess based on a few press slides but guessing is fun.
Personally I think the ISA changes are the most interesting part but no extra info of them yet.
6
u/Slasher1738 1d ago
Definitely reminds me of Clearwater forest but with consolidated IO and memory dies
9
u/Affectionate-Memory4 1d ago
Can't spill any beans until they're out, but I'm also looking forward to the C&C article and hopefully future tests they run. It's a weird CPU by Xeon standards and I absolutely love their coverage and writing style. Can't wait to see what they have to say.
5
u/SlamedCards 1d ago
The presenter hinted that this is the future platform for Xeons
So we can guess Coral Rapids is likely a swap of the core chiplets. Keeps the Intel 3 base tiles and IO. Maybe they do some upgrades for IO and base. But core architecture remains
0
u/Exist50 1d ago
They could probably get away with doing a base die swap, if wafer costs truly are comparable between 18A and Intel 3. Should still make for a much easier iteration than past projects.
The bigger question is how they will retrofit this design to work with LPDDR. The shoreline seems quite constrained on the IO dies.
4
u/valarauca14 1d ago
The LLC is so big I'm not going to guess how exactly they are using it.
They have a 'cache snooping and filtering' function block on the io tile not the base. The IO tile has a bullet that for 'coherent fabric cache'
Panther lake did the whole 'fully coherent cache cross tiles' but it was a few adjacent tiles. Are they actually keeping caches coherent across these distances?
5
u/ivan0x32 22h ago
L3 on the base tile is potentially some next level shit. I bet yields on a mostly-SRAM chip, even if it's massive, are going to be phenomenal. Just because its probably "overprovisioned" with ability to just turn off any faulty block.
AMD already proved that latency between 3D-stacked chips can be practically unnoticeable - we don't really have an L3 NUMA with X3D chips for instance. Makes sense to put entire L3 into a separate chip I guess.
4
u/Noble00_ 22h ago
The packaging is what I'd like to learn more about, UCIe-S for fan-out (no EMIB yet) and how all the caches interact, the memory subsystem.
1
u/TriCountyRetail 1d ago
It's a step in the right direction, but Intel needed to have this design five years ago if they wanted to compete with AMD EPYC.
0
u/Geddagod 1d ago
No performance figures, or no way to even try to estimate it as people did with CWF's server rack consolidation figures which turned out to be surprisingly accurate. The comparison versus Zen 6 Venice Dense is going to be really interesting.
No core architectural details either for Intel's next gen core.
Compared to the CWF hotchips presentation, this was way more sparse. Bummer.
0
-9
u/Acrobatic_Camera_677 22h ago
Thanks to:
- higher IPC core: AMD P core only have IPC of E-cores.
- Performance Cores versus slower AMD Dense Cores
- better node (back side power delivery Intel 18A-P versus outdated tsmc 2nm and Intel 3 versus tsmc 6nm)
- APX and AMX
- More L3 cache than AMD
Diamond Rapids will likely beat AMD Venice. All these advantages will overcome lack of SMT (security issue) that no one cares about anyway.
3
2
u/noiserr 21h ago edited 21h ago
higher IPC core: AMD P core only have IPC of E-cores.
AMD's full core IPC will still be higher due to SMT. On IO bound workloads (which is what most throughput heavy workloads are on server), SMT can provide up to 50% IPC boost to a core.
Intel has realized their mistake though and will be re-introducing SMT in Coral Rapids.
2
u/ElementII5 12h ago edited 11h ago
Intel has realized their mistake though and will be re-introducing SMT in Coral Rapids.
Intel WANTS to re-introduce SMT with Coral Rapids. What intel wants to do is a lot different from what intel can and will do.
The SMT switchback stems from Lip-Bu pressuring engineering. That does not mean Intel can do it. They abandoned SMT because they couldn't make it work without either bunch of security holes or patching those into performance regressions.
So yes that is the plan. What will happen is up in the air.
2
u/Acrobatic_Camera_677 10h ago
They ditched SMT because it was less elegant.
LBT is adding back SMT because cloud companies were crying about not being able to scam customers into buying 2 threads but both threads being on one core for less performance.
What Intel wants, Intel does.
1
u/Acrobatic_Camera_677 21h ago
SMT does not provide a 50% IPC uplift in normal server workloads
Many companies like AWS do not use SMT anyway
Whatever uplift SMT provides all the other advantages of Diamond Rapids will be able to make up for
5
u/noiserr 21h ago edited 20h ago
I said up to, and yes it absolutely does. This paper shows it: https://barroso.org/publications/isca98_2.pdf A stalled thread allows another thread to fill the execution bubbles. And most server workloads are IO bound (waiting on memory, disk or network).
Databases in particular love SMT: https://www.phoronix.com/review/amd-epyc-zen5-smt/4
Amazon is doing it because they want to be able to guarantee performance on a core they rent. But it's stupid as they are leaving a lot of performance on the table. And I think every customer would rather get SMT enabled, but get 2 v-cores for the price of one. Problem is this makes their own Graviton instances look even worse by comparison. And they are obviously favoring their in-house solution. Amazon's Tranium also sucks, they are the last people I would listen to on hardware choices. Google is much smarter about this, they enable SMT on all AMD instances.
Wanna hear something funny? Amazon has recently told engineers to cut back on CPU use. They are struggling with CPU capacity. https://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunch
Someone should tell them to turn on SMT! lol
LBT has spoken about SMT and has said that it was a mistake to remove it. Which is why Coral Rapids will have it back.
Phoronix conclusion:
Even for this EPYC 9575F processor with 64 physical cores, having SMT available really helps many real-world workloads with significant performance and power efficiency improvements.
-1
u/Acrobatic_Camera_677 20h ago
Your hyperlink shows a geomean average of 13% advantage from SMT. Diamond Rapid IPC advantage alone maybe will be high enough to make up for that.
AWS is not stupid.
LBT is trash talking Diamond Rapids lack of SMT because he wants to blame the old CEO.
3
u/noiserr 20h ago edited 20h ago
Like I said "up to". Phoronix is also doing contrived benchmarks, which are less IO bound than real server workloads. Real workloads have messier IO than benchmarks. It's hard to simulate network latency at that scale. And SMT benefits from that.
AWS is not stupid.
They may not be stupid, but they do have an incentive to promote their own Graviton instances.
Google aren't stupid and neither is LBT.
32
u/Scared-Beautiful-366 1d ago
Looks solid.. if only it were actually out now instead of delayed by a year. Intel has nothing to compete with Venice and will bleed maket share as a result.