Qwen 3.8 Flash Next 180B (80gb) on android mobile phone 12gb RAM (400$ phone)...
Qwen 3.8 Flash Next BF16 on 4x Nvidia DGX Sparks is also great at making dioramas
A 27 billion parameter AI model running on a $300 GPU with just 8GB of memory.
That's what Unsloth's new quantization has achieved with Qwen 3.8 27B — straight to PC, no cloud, no subscription.
Qwen 3.8 flash next just made 125B parameters look cheap.
The wild part isn’t its size. It’s how little of that size it needs every time it thinks. The Architecture: → 125B main parameters → Only 6B activated per token → Another 51B parameters in its Engram embedding
You have reached the end of the archive
All of qwen38