I've been trying to find a good solution local AI.
I tried using Qwen 3.8 27B on my M5 Max w/ 128GB and it has not been a good experience. I have a desktop computer running Ubuntu with a 5090 and quite honestly that experience has been much better, but it is a remote connection
Wanted To See Max Speeds of 6 Instances of Qwen 3.8 27B w/ 2 x 3090 NVLINK and DFLASH 2
Qwen3.8-27B • 6 parallel coding agents • Total: 8,352 tokens in 10s • Aggregate: 330.6 tok/s peak · 812.0 tok/s sustained • Per-stream: 147.8 high / 145.7 low / 147.1 avg tok/s • TTFT
Qwen 3.8 saves you $20/month on Claude.
All it takes is a $5,000 Mac Studio/DGX Spark upgrade. Now, how long until we break even?
Most people are comparing Gemini 3.7 Flash vs Qwen 3.8.
They're asking the wrong question. The real question: How do you combine them into a business system? → Gemini 3.7 Flash = strategy + planning + orchestration. → Qwen 3.8 27B = production + deployment + repeatable
You have reached the end of the archive
All of qwen38