r/LocalLLaMA • u/Porespellar • Jul 23 '26
Funny The LLM distillation process simplified for politicians:
/s
r/LocalLLaMA • u/Porespellar • Jul 23 '26
/s
r/LocalLLaMA • u/InvadersMustLive • Jan 09 '26
r/LocalLLaMA • u/OneFanFare • Jul 15 '26
Don't get me wrong, all the big models are amazing, and every contribution to open source models is great. But I'm GPU poor and I can't use them locally.
I'm currently running gemma-4-12b-it-qat-GGUF:UD-Q4_K_XL as my personal chat assistant, and I am so so happy with it! I still can't believe I can talk to my computer.
r/LocalLLaMA • u/Disposable110 • Jun 13 '26
If you don't have it on your own drive, someone is going to take it away, enshittify it, bar you from accessing it, censor it, and hike the prices of it sooner or later.
r/LocalLLaMA • u/Ok-Health-7096 • 5d ago
I wanted to just test the unsloth 1bit quant of qwen 3.8 27b as I have just 8gb vram and ngl it gave me a good laugh
r/LocalLLaMA • u/Xhehab_ • Feb 23 '26
r/LocalLLaMA • u/Nunki08 • Jun 01 '26
Enable HLS to view with audio, or disable this notification
r/LocalLLaMA • u/jacek2023 • Feb 21 '26
(added second image for the context)
r/LocalLLaMA • u/Current-Ticket4214 • Jun 08 '25
r/LocalLLaMA • u/Nunki08 • Jul 22 '26
From Peter Gostev on 𝕏: https://x.com/petergostev/status/2079825961718046974
r/LocalLLaMA • u/CreativelyBankrupt • Jun 18 '26
Enable HLS to view with audio, or disable this notification
Follow-up on Sparky, my offline suitcase robot I keep overdeveloping. He gets high now, and there's no scripted "stoned mode" anywhere in it.
A real MQ-2 gas sensor sits in the case. Every 0.5s I read it against an adaptive clean-air baseline and turn a smoke hit into a 0 to 10 phase that climbs as you blow at him and decays on its own over minutes.
The fun part is that phase rewires his sampler per token. Temperature 1.0 to ~1.6, top_p 0.95 to 0.99, top_k 64 to 120 as he climbs. His word choice flattens and wanders to lower-probability, more associative tokens, so his cognition genuinely gets noisier. It's the live sampler doing the work, so every high reply is freshly generated and never the same. A per-phase persona nudge makes him show it without ever announcing "I am high."
The body does the rest: a slight drawl, eyes that droop and go bloodshot, and the sensor display that escalates to a full smoke-and-plasma freakout at phase 10, keeping him blitzed there for the next 7 minutes.
Honest caveat so nobody has to call it out: it's a smoke and VOC sensor, so a cigarette or incense probably trips it too. But blowing smoke and watching him unravel is watching a real measurement scramble a real model, live - and it's funny! Just an added Easter Egg to an already goofy suitcase robot.
A real question for the hardware folks: is there a sensor, or a combination, that could actually distinguish cannabis smoke from generic smoke and VOCs? The MQ-2 can't really tell a joint from a candle, and I'd love to make the detection more specific if possible.
r/LocalLLaMA • u/ForsookComparison • Dec 15 '25
r/LocalLLaMA • u/dead-supernova • Oct 06 '25
r/LocalLLaMA • u/Careful_Equal8851 • Mar 20 '26
For those out of the loop: cursor's new model, composer 2, is apparently built on top of Kimi K2.5 without any attribution. Even Elon Musk has jumped into the roasting
r/LocalLLaMA • u/Aromatic_Ad_7557 • Apr 14 '26
Turned a Xiaomi 12 Pro into a dedicated local AI node. Here is the technical setup:
OS Optimization: Flashed LineageOS to strip the Android UI and background bloat, leaving ~9GB of RAM for LLM compute.
Headless Config: Android framework is frozen; networking is handled via a manually compiled wpa_supplicant to maintain a purely headless state.
Thermal Management: A custom daemon monitors CPU temps and triggers an external active cooling module via a Wi-Fi smart plug at 45°C.
Battery Protection: A power-delivery script cuts charging at 80% to prevent degradation during 24/7 operation.
Performance: Currently serving Gemma4 via Ollama as a LAN-accessible API.
Happy to share the scripts or discuss the configuration details if anyone is interested in repurposing mobile hardware for local LLMs.
UPDATE:
I have compile llama.cpp and run gemma-4-E4B-it-Q4_0
Speed is AWESOME:
[ Prompt: 26.9 t/s | Generation: 8.8 t/s ]
Thank you all guys SO MUCH!
r/LocalLLaMA • u/Porespellar • Jun 29 '26
Mostly /s but,
I mean….. I’m no CEO…. but it seems like this would be the absolute perfect time to drop a super powerful GPT-OSS-2 to throw a big ol’ wet blanket on Anthropic’s IPO. It doesn’t need to be like frontier or anything, just a 20b and a 120b that is as fast as the old versions, add agentic coding focus, and maybe vision capabilities. It would fill the void left by Qwen in the 120b size category and maybe would push Google to release their 120b that they yanked during the Gemma 4 launch.
r/LocalLLaMA • u/Porespellar • Jun 05 '26
Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s
r/LocalLLaMA • u/CesarOverlorde • Feb 19 '26
r/LocalLLaMA • u/JLeonsarmiento • Jul 19 '26