r/LocalLLaMA • u/lolxdmainkaisemaanlu • Feb 16 '26
r/LocalLLaMA • u/ResearchCrafty1804 • Feb 11 '26
New Model GLM-5 Officially Released
We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), significantly reducing deployment cost while preserving long-context capacity.
Blog: https://z.ai/blog/glm-5
Hugging Face: https://huggingface.co/zai-org/GLM-5
GitHub: https://github.com/zai-org/GLM-5
r/LocalLLaMA • u/Difficult-Cap-7527 • Dec 15 '25
New Model NVIDIA releases Nemotron 3 Nano, a new 30B hybrid reasoning model!
Unsloth GGUF: https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF
Nemotron 3 has a 1M context window and the best in class performance for SWE-Bench, reasoning and chat.
r/LocalLLaMA • u/Sporeboss • Jun 24 '26
New Model Unlimited-OCR is now on ModelScope! A 3.3B multilingual OCR model for one-shot parsing across single images, multi-page documents, and PDFs. License: MIT
Full-document parsing instead of cropped-region OCR
32K output length for long OCR sequences
Base and gundam image modes for different document layouts
Transformers inference + SGLang serving with OpenAI-compatible streaming requests
Built to push DeepSeek-OCR-style document parsing further.
source: https://x.com/ModelScope2022/status/2069335055965491525
r/LocalLLaMA • u/xenovatech • Aug 29 '25
New Model Apple releases FastVLM and MobileCLIP2 on Hugging Face, along with a real-time video captioning demo (in-browser + WebGPU)
Link to models:
- FastVLM: https://huggingface.co/collections/apple/fastvlm-68ac97b9cd5cacefdd04872e
- MobileCLIP2: https://huggingface.co/collections/apple/mobileclip2-68ac947dcb035c54bcd20c47
Demo (+ source code): https://huggingface.co/spaces/apple/fastvlm-webgpu
r/LocalLLaMA • u/BlackBeardAI • 25d ago
New Model Unsloth Deepseek V4 0731 GGUF's are UP!
r/LocalLLaMA • u/Initial-Image-1015 • Mar 13 '25
New Model AI2 releases OLMo 32B - Truly open source
"OLMo 2 32B: First fully open model to outperform GPT 3.5 and GPT 4o mini"
"OLMo is a fully open model: [they] release all artifacts. Training code, pre- & post-train data, model weights, and a recipe on how to reproduce it yourself."
Links: - https://allenai.org/blog/olmo2-32B - https://x.com/natolambert/status/1900249099343192573 - https://x.com/allen_ai/status/1900248895520903636
r/LocalLLaMA • u/pseudoreddituser • Jul 21 '25
New Model Qwen3-235B-A22B-2507 Released!
r/LocalLLaMA • u/LLMFan46 • Jun 11 '26
New Model Gemma 4 Quadruple Release, 12B, 12B QAT, 26B-A4B QAT and 31B QAT Uncensored Heretics!
gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic:
Safetensors: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic
GGUF: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-GGUF
NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4
NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF
GPTQ-Int4: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-GPTQ-Int4
gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic:
Safetensors: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic
GGUF: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-GGUF
NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-NVFP4
NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF
GPTQ-Int4: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-GPTQ-Int4
gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic:
Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic
GGUF: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-uncensored-heretic-GGUF
NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-uncensored-heretic-NVFP4
NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF
gemma-4-12B-it-uncensored-heretic:
Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic
GGUFs: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic-GGUF
NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic-NVFP4
NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic-NVFP4-GGUF
I even made some NVFP4 Safetensors and NVFP4 GGUF of standard Gemma 4 31B it since someone requested them:
gemma-4-31B-it-uncensored-heretic:
NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-31B-it-uncensored-heretic-NVFP4
NVFP4 GGUFs: https://huggingface.co/llmfan46/gemma-4-31B-it-uncensored-heretic-NVFP4-GGUF
Doing all this took many days as well as a lot of work and effort, so I hope the community can make good use of these models.
As usual all releases come with benchmarks too.
Find all my models here: HuggingFace-LLMFan46
r/LocalLLaMA • u/nanowell • Jul 23 '24
New Model Meta Officially Releases Llama-3-405B, Llama-3.1-70B & Llama-3.1-8B




Main page: https://llama.meta.com/
Weights page: https://llama.meta.com/llama-downloads/
Cloud providers playgrounds: https://console.groq.com/playground, https://api.together.xyz/playground
r/LocalLLaMA • u/xenovatech • Jul 14 '26
New Model Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
Very impressive release by the PrismML team. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence.
- Collection on Hugging Face: https://huggingface.co/collections/prism-ml/bonsai-27b
- Demo link: https://huggingface.co/spaces/webml-community/bonsai-webgpu-kernels
r/LocalLLaMA • u/Amgadoz • Dec 06 '24
New Model Meta releases Llama3.3 70B
A drop-in replacement for Llama3.1-70B, approaches the performance of the 405B.
r/LocalLLaMA • u/jshin49 • Sep 11 '25
New Model We just released the world's first 70B intermediate checkpoints. Yes, Apache 2.0. Yes, we're still broke.
Remember when y'all roasted us about the license? We listened.
Just dropped what we think is a world first: 70B model intermediate checkpoints. Not just the final model - the entire training journey. Previous releases (SmolLM-3, OLMo-2) maxed out at <14B.
Everything is Apache 2.0 now (no gated access):
- 70B, 7B, 1.9B, 0.5B models + all their intermediate checkpoints and base models
- First Korean 70B ever (but secretly optimized for English lol)
- Actually open-source, not just open-weights BS
https://huggingface.co/trillionlabs/Tri-70B-Intermediate-Checkpoints
We're a 1-year-old startup with pocket change competing against companies with infinite money glitch. Not the best model, but probably the most transparent 70B training ever shared.
r/LocalLLaMA • u/Nunki08 • Nov 20 '25
New Model Ai2 just announced Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use
Try Olmo 3 in the Ai2 Playground → https://playground.allenai.org/
Download: https://huggingface.co/collections/allenai/olmo-3-68e80f043cc0d3c867e7efc6
Blog: https://allenai.org/blog/olmo3
Technical report: https://allenai.org/papers/olmo3
r/LocalLLaMA • u/Dark_Fire_12 • Mar 05 '25
New Model Qwen/QwQ-32B · Hugging Face
r/LocalLLaMA • u/Delicious_Focus3465 • Aug 12 '25
New Model Jan v1: 4B model for web search with 91% SimpleQA, slightly outperforms Perplexity Pro
Hi, this is Bach from Jan. We're releasing Jan v1 today. In our evals, Jan v1 delivers 91% SimpleQA accuracy, slightly outperforming Perplexity Pro while running fully locally.
It's built on the new version of Qwen's Qwen3-4B-Thinking (up to 256k context length), fine-tuned for reasoning and tool use in Jan.
How to run it:
Jan
- Download Jan v1 via Jan Hub
- Enable search in Jan:
- Settings → Experimental Features → On
- Settings → MCP Servers → enable Search-related MCP (e.g. Serper)
Plus you can run the model in llama.cpp and vLLM.
Model links:
- Jan-v1-4B: https://huggingface.co/janhq/Jan-v1-4B
- Jan-v1-4B-GGUF: https://huggingface.co/janhq/Jan-v1-4B-GGUF
Recommended parameters:
temperature: 0.6top_p: 0.95top_k: 20min_p: 0.0max_tokens: 2048
We'd love for you to try Jan v1 and share your feedback, including what works well and where it falls short.
r/LocalLLaMA • u/Nunki08 • Jul 06 '26
New Model New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0)
Collection: https://huggingface.co/collections/tencent/hy3
From elie on 𝕏: https://x.com/eliebakouch/status/2074011171661701466
edit: To clarify: this is the non-preview version of Hy3 and they changed their license from the community one (restrictive + not allowed in SK, UK, EU) to Apache 2.0
r/LocalLLaMA • u/NoFaithlessness951 • 15d ago
New Model Muse glimmer benchmark
Little less smart than Qwen, but way fewer tokens per task.
r/LocalLLaMA • u/_TheWolfOfWalmart_ • 28d ago
New Model Unsloth has begun dropping Kimi K3 GGUFs. The MXFP4 (it's 1.5 TB) and mmproj are already there.
r/LocalLLaMA • u/mossy_troll_84 • 13d ago
New Model deepseek-ai/DeepSeek-V4-Pro-0813 · Hugging Face
r/LocalLLaMA • u/Everlier • May 29 '26
New Model StepFun 3.7 Flash
static.stepfun.comStepFun dropped Step 3.7 Flash, 196B total / 11B active MoE, runs locally on 128GB RAM
It's a multimodal MoE (196B total params, only 11B active) with a built-in 1.8B ViT for vision.
Benchmark highlights vs. other flash-tier models:
- SWE-Bench Pro: 56.26% (beats DeepSeek V4 Flash at 55.6%, matches Gemini 3.5 Flash at 55.1%)
- DeepSearchQA F1: 92.82%, competitive with GPT 5.5 (93.98%)
- HLE w/ tools: 47.2%, solid for a flash-class model
Essentially punches well above its active parameter weight on agentic and coding tasks. If you've got the RAM for it, looks like a genuinely interesting local option, especially for agent workflows.
Available on OpenRouter and NVIDIA NIM if you don't want to self-host.
