FRAMEWIREIndonesiaUpdated Aug 20Live wire
0:00 / 0:00
0:00 / 0:00

INSANE. Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM.

64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk. Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP. Just 25 GPU layers offloaded to stay inside

CyrilXBTAug 2087
0:00 / 0:00

And now I also have a swarm running Qwen 3.8 27b in my nvidia RTX 5090.

I like where this is going so far. Pretty locked in. gm

NooNe0xAug 205
0:00 / 0:00

16G Beggars Edition M1 deployment Qwen 3.8-27B actual test report

No surprise, it’s like a typewriter, but my IQ is okay😄 📊 Measured configuration and data Device: MacBook Pro (M1/16GB) Model: Qwen3.8-27B (UD-Q2_K_XL, 9.15GB) GPU: Full layer Metal offloading Memory: Stable around 11 GB First Token delay: about 850ms Pre-filling speed: 10~12.5 tok/s…

LonelyAug 20199