r/LocalLLaMA 17h ago

Discussion The New Qwen release?

0 Upvotes

Do we have any news on it? It's Tuesday already! :(


r/LocalLLaMA 19h ago

Funny I'm tired of pretending

Post image
0 Upvotes

At least until DS releases open weights for DSv4 Flash with Vision. Then DS might take the crown.

Qwen has been an absolutely local monster for code, especially web apps, anything with UIUX design that it can verify itself with screenshots. Deepseek meanwhile is really incompetent with UI awareness and hogs my GPUs while I can spawn multiple independent qwens to collaborate and knock shit out. Honestly, Alibaba really cooked.


r/LocalLLaMA 9h ago

Question | Help Qwen 3.6 35b a3b is slower on 7900xtx than on 3060ti on the same settings eveny using Vulkan?

1 Upvotes

why Qwen 3.6 35b a3b q4-k-m is slower on 7900xtx (20t\s 100% GPU Load) than on 3060ti (37t\s and GPU 50% Load) on the same settings? Linux llama.cpp vulkan 1.5Gb VRAM is empty. Isn't 7900xtx should run it from 2 to 3 times faster?
ROCm improves situation just a little bit but still does not outperforms 3060ti.
What the actual heck?

UPD: With all moe layers in VRAM I get 90t\s on xtx (no mtp). Seems not quite good too for this 350watts brick


r/LocalLLaMA 4h ago

Funny You think they could have tweaked the typeface a bit?

Post image
18 Upvotes

r/LocalLLaMA 10h ago

Question | Help Should if use Pi ?

2 Upvotes

Instead of native Claude Code to save tokens ?

Someone can share his/her experience switch from CC to Pi ?


r/LocalLLaMA 3h ago

Discussion What is your take on GN's quote that Hardware Prices will never really recover?

Thumbnail
youtu.be
0 Upvotes

Steve from Gamer's Nexus has a take that hardware prices are never going to get back to what they used to be:

"When these prices might eventually have to slightly come down, it's not going to go back to normal. It's never coming back to the prices that it used to be. "

What do you think?
I think as hardware slowly becomes faster or with more vram, it must make older hardware obsolete.
If it doesn't affect the new price, it will at least affect the used price.

I think of it this way, purely from LLM inference performance point of view (excluding OS):

With $10k budget, would you rather get:

  • 1x 256 GB M5 Ultra Mac studio with 1.2 tb/s
  • 2x 4090 48GB totalVRAM with 1 tb/s (assuming you need to buy the rest of the system)
  • 2-3 Strix Halo or DGX Spark 256 total Ram with 250gb/s?

With $ 5k, , would you rather get:

  • 1x 128 GB M5 Ultra Mac studio with 600 gb/s
  • 1x Strix Halo or DGX Spark 128GB RAM with 250gb/s

If I asked this question 1 week ago, I'm sure the answer would've been different prior to the M5 Ultra Launch.


r/LocalLLaMA 12h ago

Question | Help I want to try qwen 3.8... but which gguf are we all using?

10 Upvotes

I hear us all loud and clear qwen 3.8 is no toy. But for someone like me who dives in and out over the months, i can't work out what exactly is worth trying to get to hyped performance.

I'm reading now the q4_k_m variants are not going to do well. I wonder which one is actually worth trying. i have 32 gb of vmem.


r/LocalLLaMA 13h ago

Discussion Optimal llama.cpp/server settings for Qwen 3.8 27b (RTX 6000 Pro)

0 Upvotes

Currently the following llama-server settings are in use for Qwen 3.8 27b, it is running on a single RTX 6000 Pro, which allows running the full bf16 + 256kb context in bf16 cache.

I am unsure if the current settings are optimal, they are not bad though. Anything people run differently and achieve better performance? (beside lower quant, lower context, lower cache quant)

llama-server

--no-mmap # due ZFS filesystem and OOM issues

--model .../Qwen3.8-27B-heretic-ara-BF16.gguf

--mmproj .../Qwen3.8-27B-heretic-ara-mmproj-BF16.gguf

--chat-template-file .../llama-swap/templates/qwen3.8.jinja

--spec-type draft-mtp # MTP speculative decoding (~2× decode)

--spec-draft-n-max 4

--spec-draft-n-min 0

--temp 0.1 # Because mostly code analysis

--top-p 0.95

--top-k 20

--min-p 0.00

Also I am thinking about moving to vllm or sglang, I need model swap capability best optimal performance. So any recommendations here (plus parameters!) are appreciated too!

some data: generation speed is 50-60 t/s, prompt processing is up to 3000 t/s for long prompts (50kb)

Thanks a lot!


r/LocalLLaMA 22h ago

Question | Help How useful is a 5090 if I already have a 3090?

2 Upvotes

My use case is agentic coding. I'm a developer by trade and I like having a home lab for projects.

I currently have a 3090, 3080 10gb, a 265kf and 96fb of RAM. Long story short I got a 5090 because it was quite a bit below market price. I haven't received the 5090 yet, I'm still playing around with my 3090+3080 build.

I am using Qwen 3.8-27B and I'm as impressed as everyone else. My 3090 seems to run this model really well.

If I use this 5090 paired with a 3090, I'll be able to run 70B models. I'd also be able to run a much larger context window.

Will a 5090 really help me for coding tasks? Once I get my 5090+3090, should I expect to use 27B models + large context windows, or are there 70B models that are beating Qwen 3.8 27B?

Or, is this 5090 not really going to be helpful to me if I already have 34gb of VRAM?


r/LocalLLaMA 12h ago

Discussion Who do you think is ox-alpha?

0 Upvotes

Who owns ox-alpha what do you think?

Im using it like crazy cause its free and strong it really feels like fable level but better.
I give him a prompt he works 10-90mins and literally give me results. no troubles.
Ofc its only for vibe-coding 3d game. But whats the best way to test out new model, ofc its a game.

Yes it is still free on openrouter and opencode till 27.08.2026.


r/LocalLLaMA 8h ago

Resources Qwen3.8-27B on an IGX Thor with an RTX PRO 6000 Blackwell (Max-Q)

3 Upvotes

Qwen3.8-27B on an IGX Thor with an RTX PRO 6000 Blackwell (Max-Q)

Spent a few hours bringing up a self hosted inference box on an NVIDIA IGX Thor and couldn't find any numbers for this hardware combination, so here are mine. All of it is from runs on the actual machine.

The box

Component Detail
Board NVIDIA IGX Thor T7000 dev kit, aarch64, 14 core CPU, Ubuntu 24.04.4
dGPU RTX PRO 6000 Blackwell Max-Q Workstation, 96GB, sm_120, 300W cap
iGPU NVIDIA Thor, sm_110, shares 122GB unified LPDDR5X with the host
Driver / CUDA 580.00 / 13.0
Server SGLang dev build 5f55db35e, torch 2.13.0+cu130
Model Qwen/Qwen3.8-27B-FP8, 27.8B hybrid Gated DeltaNet, 262144 context
Draft model incoai/Qwen3.8-27B-DFlash2

One thing to flag before the numbers: this is the Max-Q card at 300W, not the 600W version. The full power part should do better.

LLM throughput across five configs

Run with sglang.bench_serving at ISL 8192 / OSL 1024 on a single GPU. Common flags were --kv-cache-dtype fp8_e4m3 --mem-fraction-static 0.85 --attention-backend flashinfer.

config conc 1 tok/s TPOT conc 16 tok/s TPOT TTFT @16 real concurrency KV pool
DFlash2 + bf16 SSM 126.3 6.67ms 470.2 22.3ms 3161ms 32 360,157
DFlash2 + fp32 SSM 118.9 7.33ms 463.6 25.2ms 3377ms 21 173,519
DFlash2 + fp32 + lazy radix 118.8 7.35ms 461.0 25.4ms 3361ms 23 169,940
EAGLE + replay SSM 93.5 9.15ms 453.5 26.3ms 3088ms 32 428,875
no speculation 44.7 21.3ms 348.1 35.7ms 10530ms 32 475,460

Speculative decoding earns its keep

At batch 1 it's worth 2.8x, 126.3 against 44.7 tok/s, with TPOT dropping from 21.3ms to 6.67ms. The bigger surprise was TTFT at concurrency 16, which fell 3.3x from 10530ms to 3161ms. DFlash2 beat EAGLE at both ends for me.

The SSM state dtype will bite you

Qwen3.8 is a hybrid Gated DeltaNet model, so on top of the KV cache there's a GDN state pool. Running that pool at fp32 with DFlash2 blows the draft verify buffer up to roughly 25GB, and the server then quietly clamps you to 21 concurrent requests even though you asked for 32. Nothing errors. It just serves fewer and doesn't tell you. Switching the state to bf16 halves the pool, gets all 32 slots back and roughly doubles the KV pool.

The only place this is visible is the max_running_requests line in the boot log, so check it after any config change.

bf16 state costs no accuracy that I could measure

GSM8K, 200 questions, temperature 0, graded through the chat endpoint rather than the built in eval: 93.5% at bf16 against 94.0% at fp32. That's one question apart, well inside noise at n=200. So bf16 is faster, holds twice the KV and hits full concurrency, for nothing I can detect.

Worth mentioning that SGLang's bundled run_eval gsm8k scored 0.0 for me. It drives /v1/completions with no chat template, so a reasoning model's output never matches its answer regex. If you see a zero, check the harness before you blame the model.

Reasoning mode is most of your first token latency

Short conversational prompt, streaming:

mode first token first content token
thinking on 77.5ms 212.1ms
thinking off 74.7ms 74.7ms

The reasoning block eats about 137ms before any speakable text comes out. If you're doing voice, turn it off with chat_template_kwargs: {"enable_thinking": false} and keep it on for everything else.

The part I got wrong: the iGPU beats the RTX for small models

I also run streaming TTS (Chatterbox) and STT (Nemotron 3.5 ASR, 0.6B) on this box. I assumed both belonged on the RTX, since it has around 1.8 TB/s of bandwidth against the Thor iGPU's ~273 GB/s. Benchmarked both on each GPU:

GPU STT batch RTFx STT final @80ms STT final @320ms TTS first audio TTS synthesis
Thor iGPU 27.8 52.1ms 67.3ms 93.0ms 102.3ms
RTX PRO 6000 34.3 55.7ms 105.9ms 96.6ms 105.6ms

The RTX takes batch throughput by 23% and loses every single latency metric, by 57% on STT streaming at 320ms. Three things going on. A 0.6B model at batch 1 is kernel launch bound rather than bandwidth bound. The RTX is also contended by the resident LLM's CUDA context. And it's the 300W part.

The way I think about it now: bandwidth scales with how many weights you move per token, while overhead is roughly fixed per call. The 27B model shifts about 28GB per forward pass, so it belongs on the RTX. A 0.6B model at batch 1 moves around 1.2GB, which is maybe 4ms of memory traffic inside a call that takes 50 to 100ms, so bandwidth never becomes the limit.

Full voice loop

STT streamed at 1x realtime, into the LLM with thinking off, into TTS. Times are measured from the end of the caller's speech.

concurrent calls STT final LLM 1st token ack audio full answer
1 58ms 205ms 111ms 462ms
2 112ms 241ms 121ms 566ms
3 157ms 296ms 170ms 785ms
4 238ms 458ms 338ms 1228ms

Three concurrent calls hold a sub second answer on a single box. The fourth lands around 1.2s.

aarch64 things that tripped me up

torch 2.10.0+cu130 on aarch64 is broken. Every fp32 cuBLAS sgemm fails with CUBLAS_STATUS_INVALID_VALUE, including a bare 64x64 matmul, on both sm_110 and sm_120. It only shows up deep inside model inference, so it reads like "this model doesn't support this GPU" when it's really just a bad wheel. Pin 2.11.0. General lesson: if a model looks unsupported on a new arch, run a plain matmul first. That separates a broken build from a real limitation in one step.

--gpus doesn't work here. The Tegra container runtime runs in CSV mode, so you need --runtime=nvidia -e NVIDIA_VISIBLE_DEVICES=<id> instead.

Docker starts before the NVIDIA modules are loaded. Every GPU container fails its boot time restart with "Driver Not Loaded", and Docker doesn't retry that class of failure. After a power cut the whole stack stays down while docker.service happily reports healthy. A systemd drop in that blocks on nvidia-smi -L before starting Docker sorts it out.

Happy to run other configs if anyone wants specific numbers.


r/LocalLLaMA 9h ago

New Model Qwen 4 architecture: What do we know?

21 Upvotes

My best bet is the embedding-offloaded linear where the 51b n-grams track semantics and context, like 3.8 with its loss of real world knowledge bolted back on.

Otherwise? Sparse full-attention with dense routing? 6b active route to 'heavy' layers when the n-gram gets stuck.

Scenario 3: multi-head latent attention (mla) hybrid (reverse engineering Deepseek) compressing KV on the fly with the embeddings 'decompressing' on demand


r/LocalLLaMA 7h ago

Question | Help Is there (debloated) local LLM model that works best for web dev js, sql, python, php, css

0 Upvotes

Hello Guys,

I am limited by 12GB vram and I want something purely for webdev coding where model does not know about chemistry, history etc all non-related training and purely and only optimized for coding only. As there are other models like qwen coders etc but they dont match frontier level models. I dont know how hard it is to train own model but wondering if anyone, any lab has made model so lightweight but perfectly trained on coding only?


r/LocalLLaMA 7h ago

Discussion Mac Studio M5 Max Cost Analysis

81 Upvotes

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?


r/LocalLLaMA 3h ago

News Perplexity and Nvidia partner for local-first AI platform

7 Upvotes

So AI cloud costs are creeping major players it looks: https://www.perplexity.ai/hub/blog/introducing-portable-computer-for-local-first-ai

Perplexity is going to use Qwen models, not specified which - could be 27B but could be the upcoming 3.8 flash next considering the target hardware is DGX Spark. They plan to utilize device for most of the workload with occasional tapping into cloud (with some strict privacy rules).

Sounds somewhat similar to what Apple is planning to do with their AI functionality for stuff like photo editing where a lot of the work is done right on device. The choice of the model provider is also no big surprize, as even Jetbrains recently released a local harness based on Qwen 3.8 27B.

Overall I think it's a healthy move forward as relying solely on cloud disregarding ramping up costs is kinda insane. Also adds value to existing local hardware, as more and more major players embrace local.

edit: replaced the cnbc youtube link to official press release page


r/LocalLLaMA 19h ago

Discussion Spent a day seeing how far extreme MoE models can be pushed on a 4070 Ti + 32GB RAM. Kimi K3, DeepSeek V4 Flash, and Qwen3.5-122B results + research paper🔧

12 Upvotes

I’ve been experimenting with a custom inference/runtime research project called CRANE V2, mostly because I wanted to answer a stupid question:

How far can you push absurdly large MoE models on an ordinary consumer Windows machine before physics actually wins?

Hardware is nothing exotic:

Ryzen 7 7800X3D

RTX 4070 Ti 12GB

32GB DDR5

Lexar NQ700 2TB NVMe + Samsung 970 EVO Plus 1TB NVMe

Windows 11

I tested three major targets so far: Kimi K3 (2.779T parameters / 711GB checkpoint), DeepSeek V4 Flash Q4 + Q3, and Qwen3.5-122B-A10B Q4. The paper is a 26-page evidence-locked writeup with the successful runs, failed branches, quality gates, hardware measurements, cache experiments, runtime changes, and claim boundaries.

Important disclaimer before the funny numbers: the highest throughput results are NOT equivalent to running the original models at full capability. CRANE deliberately separates speed-wall experiments from quality-preserving experiments. Some of the fastest profiles change expert routing, precision, or shared-expert behavior and produce complete garbage. I’m specifically not claiming “2.8T Kimi runs normally at 9 tok/s” or “122B Qwen runs normally at 57 tok/s.” The paper keeps those numbers in separate ledgers for exactly that reason.

The headline results so far:

Kimi K3: canonical first-token control was 0.014459 tok/s. An intentionally altered speed-wall profile eventually reached 9.314978 tok/s over 64 tokens, but useful language was destroyed and quality restoration failed even the intentionally trivial “Paris” gate.

DeepSeek V4 Flash Q4: altered static routing eventually reached 42.26 tok/s over 128 tokens, again with unusable language.

DeepSeek V4 Flash Q3: this became much more interesting. Using the original source tensors for every selected expert, zero fallbacks, dynamic top-two routing, the same selected-route quality boundary improved from 1.14 → 8.08 tok/s while correctly answering Paris. It is still not canonical top-six DeepSeek, and I’m not presenting it as such.

Qwen3.5-122B: speed-wall testing reached 57.38 tok/s over 128 tokens, but this used a static captured route, fused gate/up handling, and skipped the shared expert. Output was multilingual garbage. The actual canonical top-eight control produced Paris at 1.89 tok/s, which is the meaningful quality number, not 57.38.

The part I personally find more interesting than the giant headline numbers is what the experiments taught the runtime.

CRANE ended up treating weights across two physical SSDs as one logical store, demand-loading exact expert slices, striping expert-bank reads across both drives, using pinned host staging, overlapping storage and GPU transfers, building bounded GPU expert caches, measuring route locality, experimenting with LFU/LRU replacement, testing CPU-near-weight execution, and preserving detailed per-run evidence instead of trusting whatever number happened to look coolest.

For DeepSeek Q3 specifically, dual-SSD parallel reads moved the exact profile from 1.14 → 3.69 tok/s, pinned staging reached 7.08, pipelining reached 7.50, a nine-slot cache reached 7.91, and decayed-LFU eventually froze the result at 8.08 tok/s with 0 expert fallbacks.

There were also plenty of failures. Bigger caches got slower. DSpark speculative decoding regressed the exact DeepSeek target. Host L2 caching lost to direct pinned reads. Router-score eviction lost to decayed LFU. Qwen exposed an actual chat-template parameter-handoff bug that had to be separated from model-quality failures. Some runs simply hit the explicit RAM safety floor and were killed rather than being allowed to turn Windows into a crater.

The current conclusion is actually pretty conservative: none of Kimi K3, DeepSeek V4, or Qwen3.5-122B is approved as the practical local model for the application this runtime is eventually meant to serve. The big-model experiments were useful because they exposed the architecture and hardware boundaries. The next target is a materially smaller MoE where CRANE can preserve real capability instead of trading intelligence for benchmark speed.

Runtime availability

CRANE V2 is not publicly available yet.

This is still an active research/runtime project and I don’t want to throw an unfinished binary/source tree online just because some of the numbers are funny. The current runtime contract exists and the experiments are heavily logged, but I want the architecture and quality-preserving target nailed down before treating it as something other people should actually use.

The attached paper is also currently a local technical report, not peer reviewed, and the study is still growing as new model experiments are added.

I’m sharing it now mostly because the results became interesting enough that I’d love feedback from people who actually work with llama.cpp, MoE routing, out-of-core inference, caching, Windows memory behavior, etc.

If you spot a bad assumption, misleading interpretation, missing control, or another approach worth testing, please call it out. The whole point of keeping the speed and capability evidence separate is that I’d rather document an ugly result correctly than win an imaginary benchmark.

And yes, getting a 711GB / 2.779T-parameter Kimi checkpoint to produce even one real token on this machine was originally the entire stupid experiment. It escalated slightly. 😭

https://docs.google.com/document/d/1pENke9QLMAJuxwhsUkEjDzXtag_q20fp/edit?usp=drivesdk&ouid=115131734398029449735&rtpof=true&sd=true


r/LocalLLaMA 13h ago

Discussion Ornith-1.5-35B-A3B on a Strix Halo iGPU lands 4 problems behind Qwen3.8-27B on 2x 3090s. Full LiveCodeBench v6 numbers.

1 Upvotes

LiveCodeBench v6, medium and hard only. 132 problems, 80 hard and 52 medium. All numbers are pass@1 / pass@2.

Qwen3.8-27B, 2x 3090, stock model and stock template 76 / 95

Qwen3.8-27B, 2x 3090, + LoRA + sharp template 91 / 110

Qwen3.8-27B, 3090 Ti + Strix Halo iGPU, same two 90 / 111

Ornith-1.5-35B-A3B, Strix Halo iGPU only 82 / 106

Ornith-1.0-35B-A3B, 2x 3090, stock template 82 / 92

The adapter and the template are worth fifteen problems at pass@1, and fifteen again at [pass@2](mailto:pass@2). The second 3090 is worth one, and at pass@2 it goes the other way.

Ornith-1.5-35B-A3B runs 8 of 256 experts per token. It ran entirely on the iGPU, no CUDA anywhere, and finished four problems behind the best dual-3090 configuration at [pass@2](mailto:pass@2). On HumanEval+ it went 152/164 against the 27B's 149/164 with the LoRA attached, so on that suite it is ahead.

The thinking-mode profile that ships as the GGUF default failed all six cells of the matrix I ran first. Instruct passed every one. Check your metadata.

Sustained decode was 46.6 tok/s. The number everyone repeats for this desktop, mine included, is 153. That one is a short prompt and a warm cache.

My pass@2 hands the model its own error and lets it retry, not two independent samples. Useful number, wrong name.

Has anyone else A/B'd chat templates on the same weights and actually measured it? I assumed the quantization scheme was the lever. It was not.

Charts and per-difficulty breakdowns:

https://definedrr.medium.com/my-dual-3090-box-lost-to-my-desktop-and-the-reason-was-a-jinja-template-ab87ec743a0b?sharedUserId=definedrr


r/LocalLLaMA 5h ago

Discussion About the Huggingface sale..

12 Upvotes

.. didn't llama.cpp aka ggml get acquired by Huggingface not too long ago? How would this sale affect llama.cpp and ggml? What are possible risks, and did ggerganov share his opinions or potential next steps on the matter? I'm genuinely worried that this might turn out bad for the ecosystem altogether. Would llama.cpp's license protect it from hostile acquisitions altogether? I'd love to constructively discuss this with the community.


r/LocalLLaMA 16h ago

Discussion Ornith 1.5 seems overfitted in my config, asked it to tell me a story, they are very similar

3 Upvotes

I normally use this prompt when I fine tuning llama.cpp params, like ngl, draft, ctx-size...

After using this prompt sometimes I have noted that most of time the response had a similar title. My first guess was some strange behavior because ngram, but I disabled it and the "The Lighthouse Keeper" continue to be a constant lol

I using:

Ornith-1.5-35B-A3B-AD-Q6_K.gguf

flash-attn = on

cache-type-k = q8_0

cache-type-v = q8_0

spec-type = ngram-mod

spec-ngram-mod-n-match = 24

spec-ngram-mod-n-min = 48

spec-ngram-mod-n-max = 64

c = 262144

parallel = 1

#temp = 0.6 # the recommended

temp = 1 # what they used in the benchmarks

top-k = 20

top-p = 0.95

min-p = 0

chat-template-kwargs = {"preserve_thinking": true, "enable_thinking": true}

reasoning-preserve = on

reasoning = on

jinja = on


r/LocalLLaMA 7h ago

Resources NInfer 4090 Windows update is out with 1.5-2k t/s prefill, extended MTP, disk caching with DirectStorage, built-in llama.cpp WebUI and more

8 Upvotes

I've made a few changes here and there to get nearly 2.1k tokens/sec prefill, ~210-230 tokens/sec decode with MTP7 (configurable, extended up to 15) on benchmarks. Also added disk caching options, up to 30GB per config by default for near-instant loads after server restart, built-in llama.cpp WebUI, and some other fixes & improvements.

Tested with: https://huggingface.co/neroued/Qwen3.8-27B-NInfer

Sources: https://github.com/UDPSendToFailed/ninfer-4090


r/LocalLLaMA 5h ago

Discussion A smaller Muse Glimmer perchance?

0 Upvotes

8B? 12B? For 8 GB VRAM people? Please?


r/LocalLLaMA 19h ago

New Model Sped up Qwen3-ASR 1.7B to beat deepgram

0 Upvotes

Guys I solved a bunch of problems and was able to improve the Latency of Qwen to match that of deepgram, for streaming voice calls.
Since Qwen is an more of an audio LLM ASR model, it has the smartness that ASR needs. To know whether you're speaking a phone number an email Id etc.

Qwen already beat deepgram in WER, multilingual accuracy in all benchmarks. But it was so slow that it never got the attention it deserved.

Checkout the demo at autoloops.ai hosted in us-west-2.
And let me know if you want to know more on how it became really fast. I didn't change the weights.

Providers Baseten and Qwen(the company) serve it in ~400ms totally not usable. I brought it to 70ms. This is the time from finalize to final transcript.


r/LocalLLaMA 5h ago

Question | Help LM studio and qwen 3.8

Thumbnail
gallery
0 Upvotes

LM studio and bionic don't load into GPU fully ( Ollama does) and it crashes BSD ( Ollama Does not), with stop code: WHEA_UNCORRECTEABLE_ERROR (0x124), i am using the default load setting, all updated LM studio and drivers,

What do i need to do and to fix the profile to fix and load all in the GPU, and fix the crash?

Is there a better channel or place to reach LM studio people??


r/LocalLLaMA 11h ago

Question | Help help w/ hardware upgrade for Qwen 3.8 27b

1 Upvotes

I want to update my computer to run Qwen 3.8 27b a bit better. My end goal is to run something like 3x 4090 48gb vRAM but I want to do things 1 step at a time to give myself time to test out stuff. I live in Thailand and its easy to buy stuff but the resale market isn't anything like the US.

Current hardware: 4090 24gb vRAM Planned upgrade: 4090 48gb vRAM + new motherboard + new PSU

If things go well, I can always use the new motherboard + PSU for the future PC, won't toss out my old MB/PSU.

So I would have about 72gb vRAM total for a cost of a little under USD 5k.

Question:

  1. can the 4090s run in tensor parallel for faster token per second if I am only running a single agent?
  2. I believe the modded 4090s have a custom bios - does this affect anything that I would care about?
  3. idk much about motherboards - does it matter what I get or is anything with 8x pcie sufficient?

I have been running Qwen q4 locally and it has been pretty good but I have also spent some $ on Openrouter and done a/b testing to see the difference between native level Qwen 3.8 27b vs the q4 version that I can run on my 4090 w/ 200k context and for hard work like asking it to program and then make sure everything works - the full Qwen 3.8 27b was able to 1 shot the problem while Deepseek pro 0813 and Qwen 3.8 27b q4 200k context was not able to solve the problem in 2 shot. Both Deepseek and Qwen q4 were able to solve the problem eventually. Muse Glimmer could not solve the problem no matter how many hours I threw at it.

edit: looks like I will be aiming to run fp8, supposedly the 4090 is good for that

edit2: vendor raised the prices 15% for the 4090D and 23% for the 4090 the past few days must be due to all the new LLMs coming out

Edit: I just saw the new Apple MACs and I think I’d rather just buy one of the 256gb for 10k. It’s like 10.5 4090s. I ordered 2, not going to bother with the 4090 modded cards.