Claude built the only game that made me care about the cats.
Same prompt. One shot. No fixes. Claude Opus 5 xhigh: ~28m · ~170K tokens · ~$4.60 Qwen 3.8 Max: ~35m · ~50K · ~$0.60 GPT-5.6 Sol: ~19m · ~22.5K · ~$0.68 Grok 4.6 xhigh: 23:48 · 113K · $0.68 All four built playable
This is the comparison that everyone is waiting for, four strong models facing each other 👀
1. The Qwen 3.8 is coming in strong 2. The GLM 5.3 is coming to take the position 3. Grok 4.6 has its own people who follow it 4. Gemini 3.7 It is not surprising that it is banned What are your expectations, who will emerge the strongest?
Ox Alpha is definitely NOT GLM 5.5, it’s honestly way worse than GLM 5.3
Just ran the benchmarks: Qwen 3.8 27B choked. Opus 5 still claps everything easily
Same prompt, four models.
Two are open-weight. Can you tell which? Gemini 3.7 Flash, Qwen 3.8, Grok 4.6, GLM 5.3. One test isn't a benchmark, but it shows why open models belong in the mix. Leaderboards get you a shortlist. Your own prompts tell you which one actually fits.
You have reached the end of the archive
All of qwen38