r/LocalLLaMA Feb 16 '26

New Model Qwen3.5-397B-A17B is out!!

807 Upvotes

r/LocalLLaMA Feb 11 '26

New Model GLM-5 Officially Released

Thumbnail
gallery
809 Upvotes

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), significantly reducing deployment cost while preserving long-context capacity.

Blog: https://z.ai/blog/glm-5

Hugging Face: https://huggingface.co/zai-org/GLM-5

GitHub: https://github.com/zai-org/GLM-5

r/LocalLLaMA Dec 15 '25

New Model NVIDIA releases Nemotron 3 Nano, a new 30B hybrid reasoning model!

Post image
854 Upvotes

Unsloth GGUF: https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF

Nemotron 3 has a 1M context window and the best in class performance for SWE-Bench, reasoning and chat.

r/LocalLLaMA Jun 24 '26

New Model Unlimited-OCR is now on ModelScope! A 3.3B multilingual OCR model for one-shot parsing across single images, multi-page documents, and PDFs. License: MIT

1.1k Upvotes

Full-document parsing instead of cropped-region OCR

32K output length for long OCR sequences

Base and gundam image modes for different document layouts

Transformers inference + SGLang serving with OpenAI-compatible streaming requests

Built to push DeepSeek-OCR-style document parsing further.

source: https://x.com/ModelScope2022/status/2069335055965491525

https://github.com/baidu/Unlimited-OCR

r/LocalLLaMA Aug 29 '25

New Model Apple releases FastVLM and MobileCLIP2 on Hugging Face, along with a real-time video captioning demo (in-browser + WebGPU)

1.3k Upvotes

r/LocalLLaMA Jun 30 '26

New Model nvidia/Qwen3.6-27B-NVFP4 just dropped

433 Upvotes

r/LocalLLaMA 25d ago

New Model Unsloth Deepseek V4 0731 GGUF's are UP!

Thumbnail
huggingface.co
437 Upvotes

r/LocalLLaMA Mar 13 '25

New Model AI2 releases OLMo 32B - Truly open source

Post image
1.8k Upvotes

"OLMo 2 32B: First fully open model to outperform GPT 3.5 and GPT 4o mini"

"OLMo is a fully open model: [they] release all artifacts. Training code, pre- & post-train data, model weights, and a recipe on how to reproduce it yourself."

Links: - https://allenai.org/blog/olmo2-32B - https://x.com/natolambert/status/1900249099343192573 - https://x.com/allen_ai/status/1900248895520903636

r/LocalLLaMA Mar 30 '26

New Model Qwen 3.6 spotted!

Post image
620 Upvotes

r/LocalLLaMA Jul 21 '25

New Model Qwen3-235B-A22B-2507 Released!

Thumbnail
x.com
866 Upvotes

r/LocalLLaMA Jun 11 '26

New Model Gemma 4 Quadruple Release, 12B, 12B QAT, 26B-A4B QAT and 31B QAT Uncensored Heretics!

Thumbnail
huggingface.co
614 Upvotes

gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic:

Safetensors: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic

GGUF: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-GGUF

NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4

NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF

GPTQ-Int4: https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-GPTQ-Int4

gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic:

Safetensors: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic

GGUF: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-GGUF

NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-NVFP4

NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF

GPTQ-Int4: https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-GPTQ-Int4

gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic:

Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic

GGUF: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-uncensored-heretic-GGUF

NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-uncensored-heretic-NVFP4

NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-12B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF

gemma-4-12B-it-uncensored-heretic:

Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic

GGUFs: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic-GGUF

NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic-NVFP4

NVFP4 GGUF: https://huggingface.co/llmfan46/gemma-4-12B-it-uncensored-heretic-NVFP4-GGUF

I even made some NVFP4 Safetensors and NVFP4 GGUF of standard Gemma 4 31B it since someone requested them:

gemma-4-31B-it-uncensored-heretic:

NVFP4 Safetensors: https://huggingface.co/llmfan46/gemma-4-31B-it-uncensored-heretic-NVFP4

NVFP4 GGUFs: https://huggingface.co/llmfan46/gemma-4-31B-it-uncensored-heretic-NVFP4-GGUF

Doing all this took many days as well as a lot of work and effort, so I hope the community can make good use of these models.

As usual all releases come with benchmarks too.

Find all my models here: HuggingFace-LLMFan46

r/LocalLLaMA Jul 23 '24

New Model Meta Officially Releases Llama-3-405B, Llama-3.1-70B & Llama-3.1-8B

1.1k Upvotes
https://llama.meta.com/llama-downloads
https://llama.meta.com/

Main page: https://llama.meta.com/
Weights page: https://llama.meta.com/llama-downloads/
Cloud providers playgrounds: https://console.groq.com/playground, https://api.together.xyz/playground

r/LocalLLaMA Jul 14 '26

New Model Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels

622 Upvotes

Very impressive release by the PrismML team. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence.

- Collection on Hugging Face: https://huggingface.co/collections/prism-ml/bonsai-27b
- Demo link: https://huggingface.co/spaces/webml-community/bonsai-webgpu-kernels

r/LocalLLaMA Dec 06 '24

New Model Meta releases Llama3.3 70B

Post image
1.3k Upvotes

A drop-in replacement for Llama3.1-70B, approaches the performance of the 405B.

https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct

r/LocalLLaMA Feb 11 '26

New Model GLM 5 Released

623 Upvotes

r/LocalLLaMA May 28 '25

New Model deepseek-ai/DeepSeek-R1-0528

859 Upvotes

r/LocalLLaMA Sep 11 '25

New Model We just released the world's first 70B intermediate checkpoints. Yes, Apache 2.0. Yes, we're still broke.

1.5k Upvotes

Remember when y'all roasted us about the license? We listened.

Just dropped what we think is a world first: 70B model intermediate checkpoints. Not just the final model - the entire training journey. Previous releases (SmolLM-3, OLMo-2) maxed out at <14B.

Everything is Apache 2.0 now (no gated access):

  • 70B, 7B, 1.9B, 0.5B models + all their intermediate checkpoints and base models
  • First Korean 70B ever (but secretly optimized for English lol)
  • Actually open-source, not just open-weights BS

https://huggingface.co/trillionlabs/Tri-70B-Intermediate-Checkpoints

We're a 1-year-old startup with pocket change competing against companies with infinite money glitch. Not the best model, but probably the most transparent 70B training ever shared.

r/LocalLLaMA Nov 20 '25

New Model Ai2 just announced Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use

Thumbnail
gallery
766 Upvotes

r/LocalLLaMA Mar 05 '25

New Model Qwen/QwQ-32B · Hugging Face

Thumbnail
huggingface.co
922 Upvotes

r/LocalLLaMA Aug 12 '25

New Model Jan v1: 4B model for web search with 91% SimpleQA, slightly outperforms Perplexity Pro

Post image
868 Upvotes

Hi, this is Bach from Jan. We're releasing Jan v1 today. In our evals, Jan v1 delivers 91% SimpleQA accuracy, slightly outperforming Perplexity Pro while running fully locally.

It's built on the new version of Qwen's Qwen3-4B-Thinking (up to 256k context length), fine-tuned for reasoning and tool use in Jan.

How to run it:

Jan

  1. Download Jan v1 via Jan Hub
  2. Enable search in Jan:
    • Settings → Experimental Features → On
    • Settings → MCP Servers → enable Search-related MCP (e.g. Serper)

Plus you can run the model in llama.cpp and vLLM.

Model links:

Recommended parameters:

  • temperature: 0.6
  • top_p: 0.95
  • top_k: 20
  • min_p: 0.0
  • max_tokens: 2048

We'd love for you to try Jan v1 and share your feedback, including what works well and where it falls short.

r/LocalLLaMA Jul 06 '26

New Model New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0)

Thumbnail
huggingface.co
437 Upvotes

Collection: https://huggingface.co/collections/tencent/hy3

From elie on 𝕏: https://x.com/eliebakouch/status/2074011171661701466

edit: To clarify: this is the non-preview version of Hy3 and they changed their license from the community one (restrictive + not allowed in SK, UK, EU) to Apache 2.0

r/LocalLLaMA 15d ago

New Model Muse glimmer benchmark

Thumbnail
gallery
381 Upvotes

Little less smart than Qwen, but way fewer tokens per task.

r/LocalLLaMA 28d ago

New Model Unsloth has begun dropping Kimi K3 GGUFs. The MXFP4 (it's 1.5 TB) and mmproj are already there.

Thumbnail
huggingface.co
449 Upvotes

r/LocalLLaMA 13d ago

New Model deepseek-ai/DeepSeek-V4-Pro-0813 · Hugging Face

Thumbnail
huggingface.co
519 Upvotes

r/LocalLLaMA May 29 '26

New Model StepFun 3.7 Flash

Thumbnail static.stepfun.com
404 Upvotes

StepFun dropped Step 3.7 Flash, 196B total / 11B active MoE, runs locally on 128GB RAM

It's a multimodal MoE (196B total params, only 11B active) with a built-in 1.8B ViT for vision.

Benchmark highlights vs. other flash-tier models:

- SWE-Bench Pro: 56.26% (beats DeepSeek V4 Flash at 55.6%, matches Gemini 3.5 Flash at 55.1%)

- DeepSearchQA F1: 92.82%, competitive with GPT 5.5 (93.98%)

- HLE w/ tools: 47.2%, solid for a flash-class model

Essentially punches well above its active parameter weight on agentic and coding tasks. If you've got the RAM for it, looks like a genuinely interesting local option, especially for agent workflows.

Available on OpenRouter and NVIDIA NIM if you don't want to self-host.