Kimi K3, Gemini 3.7 Flash, GPT-5.6 Sol, and Claude Opus 5 had the exact same theme: ``Draw a herbivore that is integrated with the dice.''
Kimi → Sheep GPT-5.6 → Cow Opus 5 → Cow Gemini → Snail For creative purposes, it is also important to consider how to expand on ambiguous instructions.
From 30% to 100% without changing the model
NVIDIA took Claude Opus 5 and did not touch one of its pesos In the ARC-AGI-3 benchmark, the model alone is around 30%. On the NVIDIA system, called AVO, that same model completed all 183 levels and scored 100/100 And he did it practically
Nvidia ran Claude Opus 5 on ARC-AGI-3.
On its own: 30% — the top result of every model tested. Same model + a custom harness + a supervisor agent = 100%. The harness is the real lever. Databricks: the wrong harness can 2x your cost. Stop chasing the newest model. Own your
New benchmark scene: I gave Claude Code (Opus 5) one prompt
Recreate the Matrix lobby shootout in Three.js from an empty repo. It generated all textures, SFX & music itself (OpenAI, ElevenLabs, Suno) and self-verified with screenshots. Result: 62s demo, 31/31 tests, 1
You have reached the end of the archive
All of Claude Opus 5