Grok 4.6 just locked in #1 on CursorBench for real-world coding.
Not only is it sitting at the top of the performance chart, it’s also on the efficiency frontier. Frontier-level coding results at a cost most models can’t touch. Claude Fable 5, Opus 5, GPT-5.6 Sol… all behind.
GLM 5.3 Max vs Qwen 3.8 Max vs Grok 4.6 vs Opus 5
GLM 5.3 Max is actually holding up really well here. My ranking for this run: Opus 5 > Qwen 3.8 Max > GLM 5.3 Max > Grok 4.6 Definitely competitive now. GLM is getting scary close.
I programmed a new single prompt game with Claude, but this time I made it much more challenging
I gave him a photo of my living room and asked him to create a 3D game in which the protagonist is a baby and his goal is to throw and eat as many things as possible before I catch him.
All models are so creative now
I gave Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, and Claude Opus 5 an ace of clubs card. And told them to be creative and draw inside the card, using the card as inspo and here are their results
You have reached the end of the archive
All of Claude Opus 5