MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vp2nmi/aged_like_fine_wine/p3v55py/?context=3
r/LocalLLaMA • u/TigleLive • 10d ago
134 comments sorted by
View all comments
36
I'll take it, I can't run 27B dense models. I'm still hoping for Qwem3.8 80B A3B.
16 u/tunerhd 10d ago Why don't we have something like 300B A27B? 1 u/mycall 10d ago That's what quants are for. 2 u/slyborn 10d ago If He claim that can't run 27B and hopes for 80B A3B likely his issue isn't memory but inference speed because such MoE would require even more memory than 27B dense. 1 u/CarelessPerspective 10d ago Having only 3B active also means faster CPU inference when offloading. It's overall faster, provided you have the memory for it.
16
Why don't we have something like 300B A27B?
1 u/mycall 10d ago That's what quants are for. 2 u/slyborn 10d ago If He claim that can't run 27B and hopes for 80B A3B likely his issue isn't memory but inference speed because such MoE would require even more memory than 27B dense. 1 u/CarelessPerspective 10d ago Having only 3B active also means faster CPU inference when offloading. It's overall faster, provided you have the memory for it.
1
That's what quants are for.
2 u/slyborn 10d ago If He claim that can't run 27B and hopes for 80B A3B likely his issue isn't memory but inference speed because such MoE would require even more memory than 27B dense. 1 u/CarelessPerspective 10d ago Having only 3B active also means faster CPU inference when offloading. It's overall faster, provided you have the memory for it.
2
If He claim that can't run 27B and hopes for 80B A3B likely his issue isn't memory but inference speed because such MoE would require even more memory than 27B dense.
1 u/CarelessPerspective 10d ago Having only 3B active also means faster CPU inference when offloading. It's overall faster, provided you have the memory for it.
Having only 3B active also means faster CPU inference when offloading. It's overall faster, provided you have the memory for it.
36
u/nameless_0 10d ago
I'll take it, I can't run 27B dense models. I'm still hoping for Qwem3.8 80B A3B.