Qwen 3.8 Flash Next BF16 on 4x Nvidia DGX Sparks is also great at making dioramas
A 27 billion parameter AI model running on a $300 GPU with just 8GB of memory.
That's what Unsloth's new quantization has achieved with Qwen 3.8 27B — straight to PC, no cloud, no subscription.
Qwen 3.8 flash next just made 125B parameters look cheap.
The wild part isn’t its size. It’s how little of that size it needs every time it thinks. The Architecture: → 125B main parameters → Only 6B activated per token → Another 51B parameters in its Engram embedding
I tested GLM 5.3 Flash and Qwen 3.8 Flash on two DGX Sparks.
One model crawled. The other built a playable game. Here’s the honest answer on whether local AI coding is worth the hardware.
You have reached the end of the archive
All of qwen38