Got it — same benchmark format, this time a burger, with Grok (4.6) replacing Kimi in the lineup.
Here's the post: Same prompt: "generate a burger." Four AI models. One of them apparently forgot what a burger looks like entirely. ChatGPT (GPT-5.6 Sol) drew what's essentially a
Claude Opus 5 is crazy for web design.
Prompt ↓
Claude Opus 5 is 30.2% on ARC‑AGI‑3, the average measured by ordinary people is about 48%, and GPT‑6 Astra directly reaches 99.9%.
What we crossed this time was not just a ranking list, but the dividing line between human beings and machine intelligence. Claude Fable 5.1 has no public verification results yet.
GPT-6 Astra vs Claude Opus 5 is getting ridiculous
Astra : → 99.9% on ARC-AGI-3 → 98% on FrontierMath Tier 4 → $10 input / $50 output per 1M tokens Claude Opus 5 : → SOTA on Frontier-Bench → SOTA on GDPval-AA → ~1.5x the next-best model on AutomationBench at the same
You have reached the end of the archive
All of Claude Opus 5