Qwen 3.8 flash next just made 125B parameters look cheap.
The wild part isn’t its size. It’s how little of that size it needs every time it thinks. The Architecture: → 125B main parameters → Only 6B activated per token → Another 51B parameters in its Engram embedding
I tested GLM 5.3 Flash and Qwen 3.8 Flash on two DGX Sparks.
One model crawled. The other built a playable game. Here’s the honest answer on whether local AI coding is worth the hardware.
The local Qwen 3.8-Flash-Next comparison.
I had the local AI (Qwen 3.8 27B) create a ``5-minute caravan-style shooting game where the rank…
I had the local AI (Qwen 3.8 27B) create a ``5-minute caravan-style shooting game where the rank increases as time passes and is destroyed, and decreases when you get hit.'' To be honest, there are only problems with the gameplay since the AI is directly taken out, but I feel like an interesting game could be made based on this.
You have reached the end of the archive
All of qwen38