Qwen 3.8 27B
Everyone argues Q4 against Q5. The thinking budget moves quality seven times more than the quant does. Same weights. Same tasks. Effort off: 61.3%. Effort max: 80.3%. Twenty points. No quant in the same study moved it more than three.
Just dropped MTP for Qwen 3.7 Flash Next (125B A6B)!
25 tokens/sec on a single RTX 4090! I took the 125B (6B active) setup from below, plugged in the new shared-Q8_0 MTP drafter, and pushed decode throughput to 25.35 tokens/sec at an 80k context window on a single
So i've been banging my head against this all day
Qwen 3.8 Flash Next Q8 GGUF active inference GPUs won't break 200W utilization looks lazy as hell had Fable 5.1 run autoresearch for 8+ hours built a whole custom recipe from scratch and i'm STILL not maxing these cards is
Qwen 3.8 27B is pretty awesome
So I figured I’d share my thoughts on local models and the power you unlock by pairing one with a harness in your own environment.
You have reached the end of the archive
All of qwen38