On my Strix Halo PC, Qwen 3.8 Flash-Next on Halogen passed 66/72 coding jobs and generated at 38.0 tok/s versus 13.9 in my earlier setup.
Different quant, backend and drafting make this a configuration comparison, not an engine-only test. Runs:
Qwen 3.8 flash next built this 3d octopus invaders game in about 2 hours on my 2x dgx sparks and i didn't even try hard.
The model is soo tasteful you can feel it in responses. here are the data and specs: qwen 3.8 flash next, 180b moe, official fp8 45 tok/s fresh, 35 tok/s
Qwen 3.8 Flash-Next resolved 17/18 coding jobs on my Strix Halo PC.
But as context grew, generation fell from about 15 to 8 tok/s across three harnesses. One run per task, so I would check context before blaming the agent. The 18-run data:
New Mac Studio day one: Qwen 3.8 27B, fully local, running on Bionic.
The speed is absolutely CRAZY 🤯 Local AI is no longer a compromise.
You have reached the end of the archive
All of qwen38