FRAMEWIREIndonesiaUpdated Aug 20Live wire
0:00 / 0:00

INSANE. Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM.

64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk. Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP. Just 25 GPU layers offloaded to stay inside

CyrilXBTAug 2088
0:00 / 0:00

16G Beggars Edition M1 deployment Qwen 3.8-27B actual test report

No surprise, it’s like a typewriter, but my IQ is okay😄 📊 Measured configuration and data Device: MacBook Pro (M1/16GB) Model: Qwen3.8-27B (UD-Q2_K_XL, 9.15GB) GPU: Full layer Metal offloading Memory: Stable around 11 GB First Token delay: about 850ms Pre-filling speed: 10~12.5 tok/s…

LonelyAug 20207