MLX guys, do we have some good quants for Qwen 3.8 27B to run on MBP Pro MAX4?
Its time to replace my Gemma in Hermés. Would love to have local power of almost Opus 4.6 while preserving the speed. Any tips?
INSANE. Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM.
64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk. Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP. Just 25 GPU layers offloaded to stay inside
16G Beggars Edition M1 deployment Qwen 3.8-27B actual test report
No surprise, it’s like a typewriter, but my IQ is okay😄 📊 Measured configuration and data Device: MacBook Pro (M1/16GB) Model: Qwen3.8-27B (UD-Q2_K_XL, 9.15GB) GPU: Full layer Metal offloading Memory: Stable around 11 GB First Token delay: about 850ms Pre-filling speed: 10~12.5 tok/s…
You have reached the end of the archive
All of qwen38