Damn it, old Mac is like fire and ice.
Yesterday, my M1 max studio 32G was deployed with Ling-3.0-tiny INT4, 64 token/s, it was smooth and stress-free, and the Mac was still as cold as ever. Today I tried a Qwen 3.8-27B. The Unsloth desktop client UnslothAI is used, Deployment of test models is very convenient. The model chosen is unsloth/Qwen3.8-27B-GGUF UD-IQ2_S…
I've been trying to find a good solution local AI.
I tried using Qwen 3.8 27B on my M5 Max w/ 128GB and it has not been a good experience. I have a desktop computer running Ubuntu with a 5090 and quite honestly that experience has been much better, but it is a remote connection
Wanted To See Max Speeds of 6 Instances of Qwen 3.8 27B w/ 2 x 3090 NVLINK and DFLASH 2
Qwen3.8-27B • 6 parallel coding agents • Total: 8,352 tokens in 10s • Aggregate: 330.6 tok/s peak · 812.0 tok/s sustained • Per-stream: 147.8 high / 145.7 low / 147.1 avg tok/s • TTFT
Qwen 3.8 saves you $20/month on Claude.
All it takes is a $5,000 Mac Studio/DGX Spark upgrade. Now, how long until we break even?
You have reached the end of the archive
All of qwen38