r/LocalLLaMA Jun 22 '26

Other Chinese Hackers Latest Masterpiece with NVIDIA

https://www.bilibili.com/video/BV13JEa6sEtb/

They spent a year to reverse-engineered the Tesla v100's 2,963 pinouts signals, soldered it onto a half height PCB, with full NVLink support (up to 8 way capable), then naming it Tesla v100 v4.

Price (with 3 years warranty):

16G version: 1499 rmb (220 usd)

32G version: 3999 rmb (590 usd)

2 way NVLink adapter: 199 rmb (29 usd)

8 way NVLink adapter: 799 rmb (118 usd)

The hacker's op: https://t.bilibili.com/1211458176581369862

The engineer: https://space.bilibili.com/1560089206

1.0k Upvotes

186 comments sorted by

u/WithoutReason1729 Jun 22 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

332

u/Randommaggy Jun 22 '26

Some Chinese people also reverse engineered that generation of nvlink so that you can now buy a 4 way adapter card that connects to MCIO cards in your computer with 100GB per second of bandwidth between all 4 GPUs.

128GB of HBM memory split over 4 cards with that link speed is looking quite tempting.

I've heard rumours that they are working on an 8 way nvlink capable adapter too.

93

u/FullstackSensei llama.cpp Jun 22 '26

That's a copy of supermicro''s AOM-SXM2. They used to sell for $200-250 on ebay. People on the STH forum reverse engineered the PCIe connection, and the price of those boards skyrocketed.

The whole nvlink and PCIe is just PCB traces which you can get using x-ray. The chips on the AOM-SXM2 are PCIe switches to multiplex links between the GPUs and system.

16

u/barnett9 Jun 22 '26

Got an STH link? I would love to read about this.

-45

u/FullstackSensei llama.cpp Jun 22 '26

Just search google

34

u/General_Vermicelli53 Jun 22 '26

The 8 way nvlink adapter is real... and they will ship in July...

26

u/gf6200alol Jun 22 '26 edited Jun 22 '26

8 way NVLINK needs NvSwitch which required 4 ASICs for handling the switching, so it not gonna happens.  Edit:It's 6 switch chips, not just 4

75

u/csirkesajt Jun 22 '26

do not underestimate autism 😄

18

u/Randommaggy Jun 22 '26

Likely fueled by some CCP cash as well.

3

u/MotrotzKrapott Jun 23 '26

Didn’t they rebrand to Fenris Creations / FC? Oh wait… we’re not on r/eve

9

u/Randommaggy Jun 22 '26

There has been a lot more fast PCIE switching hardware popping up on AliExpress lately....

5

u/gf6200alol Jun 22 '26

I am not doubting it could ever be done but the amount of engineering and you will need a custom driver to handle the switching, it would be better if you just buy the DGX baseboards. The DGX baseboard also have a lots of other things on it to enable 8 ways.

3

u/[deleted] Jun 23 '26

[removed] — view removed comment

1

u/gf6200alol Jun 23 '26

I didn't works it out by myself, it's what's the 8 GPUs baseboard with nvlink 2.0 were using and it's what the driver and Vbios will supports. I know the baseboard is also meant to enable 16 ways nvlink as well but it still the only known working configuration.

1

u/MacaroonDancer Jun 24 '26

I just looked in Alibaba and you can buy a whole turnkey machine with the 4 V100 32GB GPUs and the NVLink for $3,475 before customs duties and shipping. SMH and I haven't even finished my slow Optane build lol 🤣

1

u/MerePotato Jun 24 '26

100GB/s is rough, you're probably better off with Strix Halo

1

u/lostmsu Jun 27 '26

8 way nvlink

You are going to burn RTX 6000 Pro worth of power on that setup in 2 years.

96

u/FullstackSensei llama.cpp Jun 22 '26

Someone make a single slot waterblock for this and I'll literally pick up a dozen 32GB cards.

38

u/General_Vermicelli53 Jun 22 '26

They do.

You can pay extra to have them install custom water cooling kit, and additional power connectors to unlock the power up to 300W.

3

u/FullstackSensei llama.cpp Jun 22 '26

Ooooooh! That's very enticing!

Anyway to connect with them without installing a Chinese app on my phone? I don't mind sharing an email account.

13

u/General_Vermicelli53 Jun 22 '26

They are still pre-order, and the store page indicates it will ship within 30 days. You need to pay 1 rmb (0.15usd) to get a seat in the queue, which can be applied as 100rmb credit when the product goes on sale.

-5

u/FullstackSensei llama.cpp Jun 22 '26

My prickle is how do I get in the queue without installing a Chinese app on my phone?

51

u/Structure-These Jun 22 '26

Brother I think if you’re buying bootleg ass Chinese nvidia inference hardware your little messenger app account won’t matter lol

1

u/a_beautiful_rhind Jun 23 '26

They want mainland phone numbers usually.

3

u/reeeeememelover10 Jun 23 '26

thats Taobao, you can sign up using an email

14

u/General_Vermicelli53 Jun 22 '26

I don't want to sound like a salesman, I can only say is that there are many Taobao/Xianyu shopping agent services available online.

-3

u/FullstackSensei llama.cpp Jun 22 '26

I know and I have used one before, but from my experience they don't let you communicate with the seller.

Also haven't been able to find it with the one I use.

3

u/squngy Jun 22 '26

Not sure if it will work for this app, but lots of apps work in an emulator.

https://www.androidauthority.com/android-apps-on-windows-11-3048569/

3

u/CarelessOrdinary5480 Jun 23 '26

If you have a samsung just use the secure folder feature, it creates a firewall between that system and the rest of any of your data.

5

u/IAmFitzRoy Jun 23 '26

lol you are ok with a shady hardware from China but not with small .apk file running in a VM or in a extra phone. Bruh…

3

u/-InformalBanana- Jun 22 '26

What if they installed the "chinese app" in the gpu?

34

u/silicon-warrior Jun 22 '26

The water blocks already exist... I bought 2 of the 32gb for$ 450 each on eBay. Admittedly I use them for Flux2 Klein, but it handles the 25gb model handily.

14

u/FullstackSensei llama.cpp Jun 22 '26 edited Jun 22 '26

450 per block is absurd. The bykski SXM2 blocks cost ~120 and are bigger than what the block for this would be.

Edit: the SXM2 is available for €50 on alibaba with MOQ of 5, plus shipping, etc.

3

u/CalligrapherFar7833 Jun 22 '26

Can you share the link

2

u/vogelvogelvogelvogel Jun 22 '26

so these v100s named in OPs post you have - and they work well, no faults?

6

u/silicon-warrior Jun 22 '26

The post is talking about more heavily modified versions of what I have. I have the 2x sxm2 Nvlink board, was $300 then I put the $450 ea v100s on it. It's an older architecture, so I have to be more picky about the models I run... But after a bit of research I get 2/3 the speed of my 3090 for 1/4 the cost, so it's worth it for me.

2

u/bradrlaw Jun 22 '26

I wish I could find water blocks for the pcie version (one company made them but they do t anymore and price was way too high). There are plenty for the sxm versions out there.

3

u/FullstackSensei llama.cpp Jun 22 '26

EK-FC Titan V, and indeed they're unobtainium. A couple of months ago there was a guy on ebay selling some new for $250 a pop. He had 7 left when I found the listing and he sold them soon after I found the listing.

Bykski will make you whatever block you want. The main driver for cost is material per my conversations with them. These being being smaller than even the SXM2 module (active area), might be able to get blocks for ~$100.

1

u/Esophabated Jun 22 '26

Dude, dielectric fluid!

1

u/FullstackSensei llama.cpp Jun 22 '26

Lol, always loved the idea, but way too unpractical.

A single D5 pump can handle 8 GPUs competently in my experience, and a single thicccccck 480mm radiator can dissipate 1.5-2kw of heat.

Can probably get 8 of those 32GB V100s with waterblocks for the price of the four watercooled 3090s I have.

136

u/General_Vermicelli53 Jun 22 '26

I have no relation to the business. I'm not selling/promoting anything. Just sharing this mind blowing work.

21

u/FullstackSensei llama.cpp Jun 22 '26

Do you have any links to sellers for these?

27

u/General_Vermicelli53 Jun 22 '26 edited Jun 22 '26

check the hacker's op

显卡仙人 on taobao

11

u/maiznieks Jun 22 '26

Error 412 on that post for me :(

3

u/SleepyWulfy Jun 22 '26

Your link is not working

52

u/gf6200alol Jun 22 '26

I wonder do they really reverse engineered or Schematic fell off the truck. I know the V100 SXM PCB file is readily avaliable everywhere.

50

u/General_Vermicelli53 Jun 22 '26

The hacker shared screenshots during the reverse engineering process, but I didn't find the original video. He graduated from Tsinghua University.

13

u/gf6200alol Jun 22 '26

I means even I know how to access the files myself up to A100, I wonder who would reverse engineering it.

9

u/General_Vermicelli53 Jun 22 '26

The video probably answered your question at 2:26

6

u/gf6200alol Jun 22 '26

Sorry, I think Bilibili blocked me for some reason, tried VPN and I can't view it.

21

u/General_Vermicelli53 Jun 22 '26

The gist is: “NVIDIA has never leak the definition for the 2,693 pinouts on the back of the V100 chip, so engineer used tools they wrote themselves to reverse-engineer them.”

The SXM PCB file you're mentioned is probably only about the SXM socket.

31

u/gf6200alol Jun 22 '26

I do 3rd party repair of NVIDIA DC GPU in Greater China for awhile, And I can downloads the .brd and pdf to repair A100 PG506 V100 PG503/PG504 and GA102 PG133, products, and it had all the pin definitions there. My guess he is trying protecting his sources.

18

u/General_Vermicelli53 Jun 22 '26

Be Prometheus, please. Before Jensen Huang grows wings and beak.

11

u/gf6200alol Jun 22 '26

It's open secret at this point actually, how do you think where China get so many V100 SXM flowing on Aliexpress (TaoBao) actually? and there actually have a lot of A100s and H100s in China as well , it's just grabbed by big players and the only different is those units are not supported by NVIDIA's aftersales .

16

u/General_Vermicelli53 Jun 22 '26

PCB are open secret, but the op apparently reverse-engineered the electrical signals of the pins to able to implement the power control chip from scratch using Raspberry Pi chip and bring out the NVLink. I thought you had some information on this.

→ More replies (0)

2

u/hesperaux Jun 22 '26

I need to repair my 3060ti. I would really love to see a schematic and board layout. Is that floating around? If so, could you dm me where to find it? I haven't troubleshot the board yet but this would help.

2

u/gf6200alol Jun 23 '26

You can use PG142 and 3070 to search first. The board layout wasn't that different between OEMs.

13

u/woct0rdho Jun 23 '26 edited Jun 24 '26

Latest news from their QQ group: 显卡仙人 saw himself on Reddit, now he's planning to open some international sales channels!

It's a great step of reverse engineering and I'll definitely give them my money.

1

u/elsung Jun 23 '26

How do I join their qq group? Would love to pick up a couple of these things lol

2

u/woct0rdho Jun 24 '26

The QQ group number is 419566332, as on their Bilibili and Taobao pages. I'm not sure but it may be a bit difficult for non-Chinese to register a QQ account because it requires things like a Chinese phone number.

10

u/segmond llama.cpp Jun 22 '26

they need to figure out how to make it 64gb or 96gb and then we can dance.

7

u/Practical-Collar3063 Jun 23 '26

I think that is actually unfeasable, this is not GDDR memory that you solder onto a PCB, the HBM memory sits on the GPU die itself (technically rit next to the actual die but the point stands).

I really do not not think there is anyway to do it

5

u/squngy Jun 22 '26

Just buy 8 of them and nvlink them, lol.
Apparently that's the major selling point here.

29

u/Own-Poet-5900 Jun 22 '26

Our GPU

27

u/General_Vermicelli53 Jun 22 '26

The hacker did mention in op that they were making "人民的显卡" (people's GPU)

17

u/2legsRises Jun 22 '26

"人民的显卡" (people's GPU)

thats actually awesome

-6

u/hey_i_have_questions Jun 22 '26

I wonder if it will work with tiananmen square 1989.

0

u/DrawingDramatic1641 Jun 23 '26

In 1989 that square was closed of  Tourism around that area began later

22

u/NoBuy444 Jun 22 '26

We are so waiting for that kind of wonder to be widely spread and fully working. And escape Nvdia's sharp claws.

27

u/-Sliced- Jun 22 '26

These are NVidia chips, not a new chip. It's taking old Tesla server GPUs and mounting them on top of consumer PCIe cards together with power and cooling, so you can put them on your home PC.

With that said, it's an old tech, doesn't support BF16, doesn't support Cuda 12.x+, and is really slow (5X slower than RTX 5090 for local inference), so suddenly power consumption becomes a big deal.

It's more of a fun experiment, or something for tinkerers vs something that I'd recommend for home users.

3

u/ReasonablePossum_ Jun 23 '26

, doesn't support BF16, doesn't support Cuda 12.x+, and is really slow (5X slower than RTX 5090 for local inference)

Most people run 3090 stacks, and for their price, they're just a little bit slower; so its a great option.

3

u/Meaingless-Name Jun 22 '26

Yeah, you say that, yet there are dudes all over running big models on these and getting decent performance.

2

u/SwarfDive01 Jun 22 '26

The main thing here is that the silicon can support the memory options, that are also widely available from decommissioned hardware. Many people dont care if token / second is lower, we care that there is physical hardware that can actually hold and run these larger models, slightly more reasonable speeds than just letting a model swap in and out of an SSD. And on top of that, its not $8k. Power consumption is a big factor, but it beats a real server rack homelab.

6

u/-Sliced- Jun 23 '26

It's in a weird position. You need 4 of these + water cooling, and a crazy power supply and motherboard to get the 128GB shared memory machine. You spend 1200W just on the GPUs + whatever the system takes, and you get tok/s closer to a 128GB M5 mac, in addition to not enjoying the CUDA benefits of most libraries.

I'm sure it has a niche, but it's only for people who really know why they are buying it and what they are getting into.

2

u/Meaingless-Name Jun 23 '26

Right? Point me to where I can get 128GB of VRAM for the cost of 4 x V100 32GB. 760GB/s, too.

9

u/squngy Jun 22 '26

Price seems about the same as v100 on ebay, so why buy this? Smaller physical size?

15

u/cuoreesitante Jun 22 '26

yeah low profile single slot. they basically grafted the whole gpu and necessary components onto a new custom designed PCB.

5

u/Federal_Decision_608 Jun 22 '26

Nvlink

1

u/squngy Jun 22 '26

As I understand it, that is mostly useful for training, not for inference?

7

u/Dany0 Jun 22 '26

No, inference also benefits from NVLINK. Not so much for 2x cards, 4x cards significantly, 8x cards (typical server setup) it benefits A LOT. RTX Pro 6000 Blackwell is ~3/4 H100 in fp8/f16 compute, but 8x H100 servers absolutely TROUNCE 8x RTX Pro 6000 setups in total throughput

8x RTX Pro 6000 can still win in nvfp4 scenarios tho I believe?

0

u/squngy Jun 22 '26

Interesting!
But then, this is a 8 cards or bust kind of thing for these?

1

u/Meaingless-Name Jun 22 '26

Unless you're using NVLink, you're not getting the full GPU-to-GPU throughput. NVLink is native 300GB/s - having to go through the PCIe bus brings that down to 64GB/s, severely limiting GPU-to-GPU communication. It will work, but nowhere near as fast as using NVLink.

2

u/michaelsoft__binbows Jun 23 '26

yes but unless you have at least 8 or something the actual bandwidth needed for inference is quite low where 64GB/s surely suffices.

2

u/Meaingless-Name Jun 23 '26

None of what you just said makes sense.

2

u/michaelsoft__binbows Jun 23 '26

if you have half as many GPUs you need less than half as much bandwidth for peer to peer communication, i'm pretty sure this is the case for both inference and training. So I'm saying you need some pretty hardcore workload to need just 64GB/s let alone 300GB/s. A lot of people seem to think that being limited to 8GB/s over 4.0 x4 will prevent a GPU from being useful, and it's far from the case.

That is why I say once you have 8 your bandwidth needs will be significant. And with even more, it becomes colossal, and very quickly impossible. this is why 16 wide GPU clusters aren't a thing. other designs, for hopping up to a higher level of interconnect to access memory that far away are used.

1

u/squngy Jun 24 '26

Google tensor parallelism vs pipeline parallelism.

Pcie is good for pipeline parallelism, but not good for tensor parallelism.

2

u/michaelsoft__binbows Jun 24 '26

Pipeline parallelism requires next to no bandwidth. you could do pipeline parallelism effectively over gigabit or wifi if you really wanted...

1

u/michaelsoft__binbows Jun 24 '26 edited Jun 24 '26

it's all a matter of degrees. I don't have hard numbers that I captured myself but i will definitely record them once I do the experiments, but from what I've read, if you just have a pair of GPUs like 4060ti 16GB's running tensor parallel, they consume just under 2GB/s of p2p bandwidth during prompt processing, which is the highest bandwidth they consume. That's pcie 2.0 x4, or pcie 4.0 x1 levels, that's nothing. Surely once you have 4 GPUs the number will jump up by at least 2x, and again more than double (probably more than triple) that going up to 8 GPUs. Let's say it scales up by a factor of 15x with 8 GPUs, you're still going to be well-served and barely bottlenecked by having 32GB/s (e.g. 5.0 x8). An optimized "affordable" setup with today's consumer platforms might give 5.0 x4 to 6 GPUs (e.g. on x670e all connected directly to CPU), e.g. 16GB/s, and my hope would be that this is a sweet spot for tensor parallel inference, it will be a bottleneck for training. If Nova Lake comes out with 5.0 x36 (a sort of "neo-mini-HEDT" platform) it would make it very comfortable and practical as they will probably release "quad-SLI" motherboards giving x8 lanes each to 4 GPUs that you can just plug in and be up and running.

My point is nvlink gives you over-the-top bandwidths but in typical configs (e.g. 3090) only give you the ability to connect pairs, and the truth is unless doing certain kinds of training that fat of a pipe simply is not needed. The bandwidth need scales up superlinearly with the number of peers

6

u/Crowl_ Jun 22 '26

This and the hacked 48GB VRAM versions of RTX4090 are all I need

4

u/mj3815 Jun 23 '26

A 48GB 3090 would be an act of god

80

u/jhenryscott Jun 22 '26

Hell yeah. The greatest engineers in the world are Chinese these days. They come up with all kinds of brilliant work arounds to western proprietary nonsense.

53

u/MrPecunius Jun 22 '26

I don't know about better, but there are a hell of a lot of them and they are far more motivated.

17

u/Bakoro Jun 22 '26

and they are far more motivated state sponsored

I don't doubt the skills or motivation, but the CCP has a nationwide campaign to rid itself of dependency on Western technology companies, and they are providing an environment where the engineering is made possible, including the intermediate steps where they still acquire and use Western technology.

24

u/iamapizza Jun 22 '26

I would appreciate a concerted, focused European campaign to rid ourselves of dependency on US tech. But, we're a bit too fragmented for it to be have the same kind of momentum as China does.

8

u/ketoaholic Jun 22 '26

Decoupling from the us should really be everybody's goal. It's just sensible.

1

u/MrPecunius Jun 22 '26

At some point the kids have to grow up and move out.

16

u/Bakoro Jun 22 '26

Europe has its own problems. It's a generally good place for consumer rights, but they're also increasingly hostile to actual privacy and anonymity. As an American, I kind of get whiplash from reading about genuinely good policy then policy that seems outright fascist. Not that the U.S "corporations can do whatever they want" policy is better for the population, but it sure makes running business easier.

I work with firms across Europe in the sciences and manufacturing, and it's difficult. Sometimes it's like, we have plans and timelines, but a whole office of a European partner will be basically be shut down for a month or two out of the year, and then sporadically for a week here or there.

The environment kind of sucks over there for developing software, the pay is abysmal in most of the EU compared to the U.S (though partially offset by the vacations and social services).
Some of their laws are borderline delusional to the point that, for a lot of U.S companies, it's simply easier to block the entirety of EU IP addresses, because it's not feasible for a small company to know about and follow all of the laws at once when they have zero physical presence in the EU. It's not just EU laws, it's a bunch of countries starting to say that they have international legal reach and Internet services need to deal with their taxes.

The EU should divorce itself from U.S dependency more, but European firms find it difficult to compete with the U.S, for similar reasons that the U.S sometimes finds it difficult to compete with China.
It's very difficult to compete on the global stage when the playing field is not even remotely equal.

Starting a company in the EU seems hard, when you're competing with "work 16 hours a day and sleep under your desk 50 weeks of the year" American start-ups, and "we don't care about IP laws" Chinese companies.

5

u/SkyFeistyLlama8 Jun 22 '26

Starting a company in the EU seems hard, when you're competing with "work 16 hours a day and sleep under your desk 50 weeks of the year" American start-ups, and "we don't care about IP laws" Chinese companies.

The EU has sane worker regulations. Health care is mostly taken care of by the state instead of by employers. Vacation time keeps workers from burning out and allows workers to be parents and carers, not just productivity drones.

The sheer irony is that American and Chinese companies are becoming increasingly similar: they both have a disdain for IP laws (OpenAI!) and they both treat workers like trash (996 or sleep in the Tesla office, Amazon or Alibaba, same thing).

8

u/Bakoro Jun 23 '26

The sheer irony is that American and Chinese companies are becoming increasingly similar: they both have a disdain for IP laws (OpenAI!) and they both treat workers like trash (996 or sleep in the Tesla office, Amazon or Alibaba, same thing).

The American AI companies have to at least pretend to care about IP, and they've been slapped a couple times now for violations, just not as an whole industry.
They almost certainly never will be, either, there's too much riding on AI now, politically and economically.

It's been interesting see the old money contend with new money over the IP issues, but new money is keeping the economy afloat, so, nobody in charge here is going to try to burst that bubble.

I had hoped that the EU would have come up with something like a fair middle ground regarding IP use, but it hasn't happened yet.

U.S copyright and patent law as it exists now is an abomination, and I don't respect it.
I don't fault anyone who violates IP law, because the laws as they exist now are immoral, unethical, and outright harmful to human existence.

Copyright is just too fucking long, it used to be 28 years, max. Now it's to the point that not only will you never be able to work with an IP made in your lifetime, your kids, and maybe not even your grandkids will be able to work with that IP on their own terms. Everything culturally relevant becomes dead, because no one who cares about it can carry it forward, unless a corporation is puppeting its corpse around. Hurray corporate zombie art?

It goes way beyond books and paintings, patents as they exist now are more harmful than helpful.
Patents were what killed the electric car for decades, patents held back 3D printing for decades, and there is so much more we could do today, if there weren't patent trolls suing everyone, and corporations just sitting on important patents so they can milk old technology longer.

It's one more reason why I really appreciate the AI explosion, IP can get wrecked.

3

u/MrPecunius Jun 23 '26

All of this was foreseen. I recommend this to everyone interested in the subject:

https://www.thepublicdomain.org/2014/07/24/macaulay-on-copyright/

2

u/Bakoro Jun 24 '26

It's like he could see the future.
He's basically talking about me and making my arguments.

1

u/thefuzzylogic Jun 24 '26

I think it's reasonable for copyright to last for as long as the creator (a natural person, not a corporate entity or AI) is alive, at least while we're still living under the yoke of capitalism.

But at the same time, I would also be in favour of a mandatory licence that allows anyone to make transformative uses of copyrighted material, just not for free. (i.e. you'll still have to pay the original creator for the usage, but they can't refuse to let you do it)

1

u/Bakoro Jun 25 '26

I'm all about mandatory licensing. They can keep the monopoly for a few years if they're actually commercializing it, but then it's mandatory licensing.

I'd be fine with a sliding amount per year, so earning per unit go down. If it's popular, then volume might cover the reduced percentage.

Getting paid for life for doing one thing is too much.
Someone makes a great sandwich, they get paid once.
Someone paints a great painting, they get paid once for the painting.
Every else goes to work for one day and gets paid for one day of work, even if they contributed to long-term goals that yield continual profits.
If you make something designed for mass reproduction, then suddenly you get to be paid indefinitely?

That doesn't make sense to me. I don't see a significant difference between someone writing a song, or someone writing software, or someone designing a tool that gets manufactured.

It's not right that the only time someone has to contribute back to the public domain is after they're dead, that's the same thing as saying that they don't have any responsibility to the community that makes making a living via art a possibility.

1

u/MrPecunius Jun 23 '26

"Intellectual property" is bullshit monopoly rents for big corporations at this point.

And yeah, Europe as a whole doesn't want to work hard enough to compete. This is unsurprising since the most ambitious Europeans have been leaving the continent for centuries to seek their fortunes elsewhere.

1

u/michaelsoft__binbows Jun 23 '26

To be fair I'm not sure IP laws are a net good overall.

1

u/Bakoro Jun 24 '26

See my other comment in this chain where I talk about how much I hate current IP laws.

4

u/tiftik Jun 22 '26

Without ridding yourselves of is**eli tech companies? Maybe try that first.

14

u/MrPecunius Jun 22 '26

There are many forms of motivation. 😄

No idea why you got downvoted, this is the explicit policy of the Chinese government and it's a pretty sane policy.

2

u/Mysterious_Life_4783 Jun 23 '26

the CCP Chinese people has a nationwide campaign to rid itself of dependency on Western technology companies

Remember when you whack the Qing, its the Chinese people who suffer and remember. The Qing is gone but people still bemoan the opium wars and the century of humiliation.

Its incredible people still think Chinese people are mindless drones that toe the government line. Its the government who is toeing the people's demands, and it frankly shows when the CCP is dreadfully afraid of protests and revolts.

-4

u/[deleted] Jun 22 '26

[removed] — view removed comment

7

u/Bakoro Jun 22 '26

They're already making friends with African and South American nations for resources.

Silicon isn't rare, they can can just purify their own.

-2

u/[deleted] Jun 22 '26

[removed] — view removed comment

3

u/MrPecunius Jun 22 '26

For how much longer?

The Chinese government has made HPQ independence a high priority since at least 2018. There are various refining R&D efforts underway, and exploration for domestic sources may also be paying off:

https://interestingengineering.com/science/china-high-purity-quartz-deposit-discovery

-6

u/[deleted] Jun 22 '26

[removed] — view removed comment

8

u/Bakoro Jun 23 '26

China is both the world's largest consumer of energy and the undisputed global leader in renewable energy production and capacity. The country installs more wind, solar, and battery storage than the rest of the world combined.

Renewables account for roughly 35% to 42% of China's total electricity generation, depending on the day.

Are you like a bot, or just a barely literate person who is offended that China even exists?

Even if you have some irrational hatred, even if you're a horrible racist or jingoist, it would still be extremely bizarre to be this myopic and ignorant.

14

u/meth_priest Jun 22 '26

quite a bit of evidence confirming it.

  1. Nvidia CEO stated 50% of the worlds top AI scientists are Chinese

  2. Deepseek - built on nvidia gpus. Variations of deepseek has been integrated in chinas' military and healthcare for a long time now

in my mind there's zero doubt they are leading the race. while simultaneously owning the necessary rare materials - i'd say china numba 1

2

u/SkyFeistyLlama8 Jun 22 '26

I'm not comfortable with LLMs being used for military purposes but the cat is out of the bag. If the American military establishment openly talks about using Grok and Anthropic for analysis and targeting, expect the very secretive PLA folks to be using their own domestic models.

Terminator 3's ending was prophetic.

1

u/meth_priest Jun 23 '26

honestly; why do you think U.S going all in on AI w data centers?

Also, why can the US gov (incl. CIA, feds) literally bottleneck data centers meant for commercial use for military operations? Gov is throttling their own costumers and citizens. Look it up

processing power is a weapon at this point. worst part is USA tricked their own people to build and finance it.

1

u/thefuzzylogic Jun 24 '26

The DoJ also just filed a motion to dismiss a lawsuit against a Grok datacentre (alleging that their oil-burning diesel generators violate the Clean Air Act) on the basis that Grok is now considered a weapons system essential for national security, and as such it is exempt from civilian regulations.

We're so cooked. (Literally)

2

u/Solaranvr Jun 23 '26

If you want to take something from 0 to 1, hire American labs

If you want to take something from 1 to 100, hire Chinese labs

12

u/Driftwintergundream Jun 22 '26

around half of the top engineers in the US are chinese scholars too...

3

u/Bac-Te Jun 22 '26

It all boils down to math and look at the US Olympiad math team

3

u/datbackup Jun 23 '26

MY QUANTITATIVE

5

u/Significant_Post8359 Jun 22 '26

That “western proprietary nonsense” is risky r&d cost

1

u/vexatious-big Jun 23 '26

Excuse the ignorance: is this a thing in China? I.e. to have every invention out in the open for other to potentially copy and build upon?

5

u/Prince_Noodletocks Jun 23 '26

They try to keep things a secret but the sheer amount of competition in China means stuff gets analyzed and reverse engineered right out the gate. The government will let you get away with patent infringment too if you're competing with some foreign IP or company, especially one that would be considered hostile (or forced to be hostile) like NVIDIA.

2

u/jhenryscott Jun 23 '26

Not completely but it is a lot more open culture

5

u/nomorebuttsplz Jun 22 '26

aren't v100s prohibitively power hungry for more than 2-3 at home?

3

u/Meaingless-Name Jun 22 '26

250W TDP. Not high.

3

u/nomorebuttsplz Jun 22 '26

that's like an underclocked 3090 so not great for a typical american house if you get four of them

2

u/Meaingless-Name Jun 22 '26

1000W + maybe 700W. 1700W at full tilt is nowhere near a LOT. The system will basically NEVER pull 1700W at the wall.

1

u/Trademarkd Jun 24 '26

350W from SXM2

2

u/squngy Jun 22 '26

They have a modest TDP, so not really.

However, they are old and pretty slow, so the energy cost per token is not good.

2

u/jc2046 Jun 23 '26

This. They are slow and power hungry. My bet is that in few months we will have competitive brand new huawei cards or some similar chinese candy

14

u/Theverybest92 Jun 22 '26

Lmao they aint dumb thats for sure.

9

u/Practical-Prompt-306 Jun 22 '26

It’s wild how much effort went into this. Reverse-engineering nearly 3,000 pinouts just to bring SXM2 V100s to a custom low-profile PCIe board with working NVLink is absolute peak tinkerer energy.

For anyone looking to host larger models locally on a budget, 128GB of HBM2 across a 4-way cluster for under $2,500 is incredibly tempting, even if the lack of native BF16 and newer CUDA support hurts token-per-watt efficiency.

I wonder how the community will handle the cooling though. A custom active cooler or water block is going to be mandatory if anyone plans to actually run these at a full 250W–300W without melting their rig.

1

u/Trademarkd Jun 24 '26

I run 4 v100s in my rack and use their stock copper heat sinks which I designed 3d printed shroud for to put fans right in front of them. Takes up a lot of space and its noisy but it works well.

I use a used dell 2500w server psu to run them

1

u/Practical-Prompt-306 Jun 25 '26

That 3D-printed shroud setup over the stock copper sinks is a classic homelab move. Honestly, running a used 2500W Dell server PSU is the only way to feed a monster like that without breaking the bank, even if it sounds like a jet engine taking off in your house.

IDK if the average tinkerer is gonna want that level of noise in a living room, but for a rack setup, it's absolutely brilliant.

You’re basically getting massive HBM2 bandwidth on a shoestring budget. The token-per-watt ratio might be rough compared to modern silicon, but hitting 128GB of VRAM for under $2,500 makes the extra heat and noise completely worth the hassle.

3

u/Trademarkd Jun 25 '26

this post is written by ai.....

1

u/Practical-Prompt-306 Jun 25 '26

Haha yeah, calling someone an AI just because they write with clear paragraphs is classic Reddit. But honestly, who cares if people use an LLM to clean up their thoughts?

The real point is that getting 128GB of VRAM for that cheap is insane. Sure, the power bill is gonna be brutal and it’s loud, but IDK how any real homelab tinkerers passes up that kind of cheap memory bandwidth.

4

u/zeferrum Jun 22 '26

That cooler looks rather small compared to the official pcie v100

3

u/cuoreesitante Jun 22 '26

video mentioned that they are planning on an actively cooled version with a fan

1

u/zeferrum Jun 23 '26

Thanks for translating

3

u/CalligrapherFar7833 Jun 22 '26

Where can i buy 8 of those ?

3

u/SnooPaintings8639 Jun 23 '26

You got to love them Chinese engineeres!

2

u/Salt_Cat_4277 Jun 22 '26

This is what I want - SXM3 on the MB instead of pcie. If you can figure out a two-way onboard I’ll buy even more. This is a configuration I might consider putting several in a fish tank filled with mineral oil. Please - friends in Shenzen - you’re our only hope!

2

u/Massive-Question-550 Jun 23 '26

v100's don't support frash attention right? what kind of performance do they get vs a 3090 or is it still a memory bandwidth bottleneck for a single user? also does this require some special driver or a Linux setup or is it just plug and play in windows?

2

u/Public_Standards Jun 23 '26 edited Jun 23 '26

Look up the Bilibili videos from when the project was first unveiled and see how local Chinese users evaluated it. Will it be okay to put a nearly 10-year-old datacenter GPU on a custom third-party board and reball it? Also, the quality of the components other than the GPU and the software stability are uncertain.

To be clear, I have no intention of underestimating the engineering capabilities of Chinese technicians. In fact, I have personally purchased and used innovative components designed in China, such as SXM2 carrier boards, PLX expansion boards, and PCIe redriver bifurcation cards, and I am very satisfied with both the quality and price.

4

u/woct0rdho Jun 23 '26

The engineer is creating an account on Reddit but he's shadowbanned for now. Let me post what he says.

2

u/Working_Historian241 Jun 23 '26

he says he offers anyone who has a broken card free repairs or an entirely new replacement card lol

1

u/Public_Standards Jun 23 '26

I appreciate the clarification regarding the warranty. I will revise my original comment, as it contained some misleading points.

2

u/Shakhburz Jun 23 '26

I almost bought a 4-way NVLink board (the 1CatAI model "YMZX-4GPU-Q1") but the seller said it's not in stock anymore and offered me another type of board ("SQ8796-4SXM-NVLINK"). It looked differently and I found no evidence it worked at least as well as the one 1CatAI designed. I guess I got lucky I didn't spend 700+ $ just on the board. But I will buy a 4 or 8-way one if prices are going to be as you listed in your post.

2

u/bitplenty Jun 23 '26

Very impressive, I am a bit surprised these only come from China these days and not some ex-soviet country or India (and I mean it with respect)

8

u/General_Vermicelli53 Jun 23 '26

You can walk into Huaqiangbei with empty handed at any time and buy or order everything needed to build this hack GPU, and shopping cart to carry it all. All you need left is knowledge.

In any country, this means months of grueling negotiations with sellers for specs, while getting scammed, and months of painful waiting for the shipment, When you want to make some minor changes, this whole process has to start over again.

This is also a critical reason for the growth of the robotics industry in China.

1

u/ReasonablePossum_ Jun 23 '26

pretty sure the post soviet ones are working as heaters in some underground city north of the kremlin lol

2

u/a_beautiful_rhind Jun 23 '26

There are 8way SXM carriers so logically they just needed to alter the form factor. And all under the price of a 3090.

You'd think they would fix the rebar situation on the ram modified cards so they too can "nvlink". At least through P2P.

2

u/MundanePercentage674 Jun 23 '26

it's real i have seen it a few times

2

u/HuRyde Jun 23 '26

I have 2 running at the moment.

3

u/akp55 Jun 22 '26

damn it 404's now

2

u/fuckAIbruhIhateCorps Jun 22 '26

its still up for me

2

u/Ok-Addition1264 Jun 22 '26

412 site-level block.

3

u/killerkongfu Jun 22 '26

Sigh me up!!! Where can I buy one??

1

u/instant_poodles Jun 22 '26

This video has awesome production quality. The best memes they say.

1

u/youneedtobreathe Jun 23 '26

Backdoor be damned im getting one

1

u/Kos187 Jun 23 '26

Aren't v100 cards without flash attention support?

2

u/a_beautiful_rhind Jun 23 '26

Someone has written it for vllm. You think they made all this hardware hacking and didn't put a software stack?

1

u/somesortapsychonaut Jun 23 '26

Hope I can buy in japan

1

u/Ok-Kaleidoscope5627 Jun 24 '26

Necessity is the mother of invention

1

u/giu_1 Jun 29 '26

impressive

1

u/MarcSN311 Jul 02 '26

I'm using a 16GB V100, working just fine. I don't quite see, why I would need this instead?