Qwen-3.8-27B is used in Claudecode of M2Max. Since the memory bandwidth is only 400GB/s
Qwen-3.8-27B is used in Claudecode of M2Max. Since the memory bandwidth is only 400GB/s, after unremitting efforts, it can only reach the speed of 16token/s for the time being. With cache, the local machine is still too slow. It takes half a day to process 20,000 tokens at a time.
Qwen 3.8-27B just matched a model that was the best in the world 6 months ago.
And it runs on your laptop. 27 billion parameters. Tiny by AI standards. It keeps pace with Claude Opus 4.6 on coding, agent work, and reading images. Here's the crazy part: → Apache 2.0 license.
GLM 5.3 vs Qwen 3.8 Max vs Gemini 3.7
I gave all three models the exact same one-line prompt. Results • GLM 5.3 → ~10 min • 735 lines of code • Qwen 3.8 Max → ~8 min • 625 lines of code • Gemini 3.7 → ~2 min 15 sec • 874 lines of code My take: Qwen 3.8 Max surprised
If you're curious what ~48 tok/s looks like.
Running MTPLX Qwen 3.8 4 bit quant -- fully locally on 128gb MBP M5 from
You have reached the end of the archive
All of qwen38