Meta just released Muse Spark 1.2 + Muse Code.
So I tested it against GPT-5.6 Sol, Qwen 3.8 and Fable 5. The benchmark scores aren’t the interesting part. The real-world builds told a very different story.
DeepSeek V4 Pro (0813) vs Qwen 3.8 Max
GLM 5.3 Max vs Fable 5 vs Qwen 3.8 Max vs Grok 4.6
GLM 5.3 is a massive jump over 5.2. Still not beating Fable 5 for me, but it’s way closer now. And vs Grok 4.6? I’d take GLM 5.3 pretty easily on this test.
Advanced Frontier LLM Coding Benchmark 2 14.08.2026
Modeller: - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - GLM 5.3 Max - Zcode (App) - GPT 5.6 Sol Ultra - Codex (App) Görev: Ferrofluid: Metaball yüzeyi + mıknatıs imlecine doğru yükselen
You have reached the end of the archive
All of qwen38