Anthropic built the hand-off into the product.
Flip one API flag and anything Fable 5.1 refuses goes to a cheaper Claude, with a credit so you don't pay twice. Score those hand-offs as failures and Terminal-Bench drops from about 85 to 79, under Opus 5 at half the price.
Benchmarks are starting to matter less.
GPT-6 Astra is hitting 99.9% on ARC-AGI-3 and 98% on FrontierMath Tier 4. Claude Opus 5 is SOTA on Frontier-Bench and GDPval-AA, while dominating AutomationBench. But Astra costs 2× more per output token: $50/M output vs $25/M for Opus.
The dispute is open: Opus 5 vs GPT-5.6 Sol.
🔥 They both received the same challenge. But which one delivered the best result? At KAIROGEN you have access to 30+ AI models.
Claude Opus 5 just keeps getting better
I had an idea in my head. Turned it into a prompt. And somehow… this came out.
You have reached the end of the archive
All of Claude Opus 5