Claude Opus 5 spent 24 hours building a complete 3D roguelite.
Just 3 prompts, one every ~8 hours, with no intervention during each session. By the end it had written 60K lines of TypeScript, generated its own design docs, built the game, and even created a cinematic trailer by
Grok 4.6 just locked in #1 on CursorBench for real-world coding.
Not only is it sitting at the top of the performance chart, it’s also on the efficiency frontier. Frontier-level coding results at a cost most models can’t touch. Claude Fable 5, Opus 5, GPT-5.6 Sol… all behind.
GLM 5.3 Max vs Qwen 3.8 Max vs Grok 4.6 vs Opus 5
GLM 5.3 Max is actually holding up really well here. My ranking for this run: Opus 5 > Qwen 3.8 Max > GLM 5.3 Max > Grok 4.6 Definitely competitive now. GLM is getting scary close.
I programmed a new single prompt game with Claude, but this time I made it much more challenging
I gave him a photo of my living room and asked him to create a 3D game in which the protagonist is a baby and his goal is to throw and eat as many things as possible before I catch him.
You have reached the end of the archive
All of Claude Opus 5