I ran Qwen 3.8 27B, Nemotron 3.5 Lightning and Muse Glimmer on 16 hard problems.
One RTX 5090, Ollama, 64K context, Q4_K_M, and every model loaded alone on the GPU. Full write up with all the charts, the method, and the raw results:
Qwen 3.8 Max Full COURSE 1 HOUR (Build & Automate Anything)
Same prompt with Qwen 3.8 27b.
Took 5x longer but its running on my macbook m5 max :)
GLM 5.3 had no business looking this good.
OmedTheVibeCoder tested GLM 5.3 Max against Fable 5, Qwen 3.8 Max, and Grok 4.6 on the same task. Fable 5 still takes it for me. But GLM 5.3 is suddenly much closer — and against Grok 4.6, I’d pick GLM pretty comfortably here. That’s
You have reached the end of the archive
All of qwen38