INSANE. Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM.
64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk. Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP. Just 25 GPU layers offloaded to stay inside
16G Beggars Edition M1 deployment Qwen 3.8-27B actual test report
No surprise, it’s like a typewriter, but my IQ is okay😄 📊 Measured configuration and data Device: MacBook Pro (M1/16GB) Model: Qwen3.8-27B (UD-Q2_K_XL, 9.15GB) GPU: Full layer Metal offloading Memory: Stable around 11 GB First Token delay: about 850ms Pre-filling speed: 10~12.5 tok/s…
You have reached the end of the archive
All of qwen38