Grok 4.6 crushed Opus 5 in this BridgeBench comparison.
One benchmark is a starting signal. The useful test is how each model performs on the workflow you actually need.
I switched my main AI coding model from Opus 5 to Fable yesterday.
It's much smoother. I didn't use it much before due to cost and speed concerns. Even though the token price doubled, you can solve an issue or PR with fewer tokens. Turns out it's even?
Gave my 11 year old son a claudeai account, and he came back with this gold.
He would like you to play it: He used Opus 5, ElevenLabs and threejs.
GPT-5.6 Sol does seem to have gotten better
In the video, you’ll see GPT-5.6 Sol’s result from when it originally launched, followed by the result from the exact same prompt today. Same test, same prompt, same harness. The difference is notable. I know I’ve been posting a lot
You have reached the end of the archive
All of Claude Opus 5