I ran Qwen 3.8, Fable 5, GPT-5.6 and Kimi K3 through 23 identical real-world builds.
Games. Landing pages. Physics simulations. Web experiences. The winner changed constantly. And that's the useful part.
DeepSeek V4 Pro (0813) vs Qwen 3.8 Max
GLM 5.3 Max vs Fable 5 vs Qwen 3.8 Max vs Grok 4.6
GLM 5.3 is a massive jump over 5.2. Still not beating Fable 5 for me, but it’s way closer now. And vs Grok 4.6? I’d take GLM 5.3 pretty easily on this test.
Advanced Frontier LLM Coding Benchmark 2 14.08.2026
Models: - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - GLM 5.3 Max - Zcode (App) - GPT 5.6 Sol Ultra - Codex (App) Task: Ferrofluid: Rising towards metaball surface + magnet cursor
You have reached the end of the archive
All of qwen38