FRAMEWIREIndonesiaUpdated Sep 6Live wire
0:00 / 0:00

Got it — same benchmark format, this time a burger, with Grok (4.6) replacing Kimi in the lineup.

Here's the post: Same prompt: "generate a burger." Four AI models. One of them apparently forgot what a burger looks like entirely. ChatGPT (GPT-5.6 Sol) drew what's essentially a

MaciavellySep 6
0:00 / 0:00

Claude Opus 5 is crazy for web design.

Prompt ↓

Faisal AhmedSep 67
0:00 / 0:00

Claude Opus 5 is 30.2% on ARC‑AGI‑3, the average measured by ordinary people is about 48%, and GPT‑6 Astra directly reaches 99.9%.

What we crossed this time was not just a ranking list, but the dividing line between human beings and machine intelligence. Claude Fable 5.1 has no public verification results yet.

岩创AISep 6
0:00 / 0:00

GPT-6 Astra vs Claude Opus 5 is getting ridiculous

Astra : → 99.9% on ARC-AGI-3 → 98% on FrontierMath Tier 4 → $10 input / $50 output per 1M tokens Claude Opus 5 : → SOTA on Frontier-Bench → SOTA on GDPval-AA → ~1.5x the next-best model on AutomationBench at the same

LummoxSep 619