Retiring the snowboarder test: Fable 5.1 and GPT-6 Astra both one-shot it, no notes.
March 2026: the first reaction was "this is fake." September 2026: nobody blinks. Both the models and the expectations moved that fast.
I'll be working 8-5 starting tomorrow, I'm starting my job, don't go crazy without me, calm down
Benchmarks don't show a model's true capabilities, but specific tasks do.
The same task for GPT-6 Astra, Fable 5.1, Grok 4.6, and Opus 5 A tank game with voxel graphics, for 2 players via WebSocket and three.js
GPT-6 Astra is truly terrifying 😨
As a designer, I was confident that artificial intelligence would not outperform me in design... Until Astra came and turned everything around 🔄 He actually uses Figma, builds entire pages, from BI to UI with very high quality. Fable 5, Opus 5, even GPT-5.6 haven't reached this level. The question that concerns me…
You have reached the end of the archive
All of Claude Opus 5