Just dropped MTP for Qwen 3.7 Flash Next (125B A6B)!
25 tokens/sec on a single RTX 4090! I took the 125B (6B active) setup from below, plugged in the new shared-Q8_0 MTP drafter, and pushed decode throughput to 25.35 tokens/sec at an 80k context window on a single
So i've been banging my head against this all day
Qwen 3.8 Flash Next Q8 GGUF active inference GPUs won't break 200W utilization looks lazy as hell had Fable 5.1 run autoresearch for 8+ hours built a whole custom recipe from scratch and i'm STILL not maxing these cards is
Qwen 3.8 27B is pretty awesome
So I figured I’d share my thoughts on local models and the power you unlock by pairing one with a harness in your own environment.
Qwen 3.8 Max Full COURSE 1 HOUR (Build & Automate Anything)
You have reached the end of the archive
All of qwen38