In the Opus 5.5 thinking gears
In the Opus 5.5 thinking gears, I ran five gears from the lowest to the highest, and ran the same question three times in each gear: a performance summary with four mistakes. Fifteen times, all four errors were caught. The lowest gear took about 19 seconds, and the highest gear took 102 seconds.
RTX 5070 Ti(16GB).
We had local models make music 🎧 ◼︎ Flow ・Song 1: It's a bit strange. Looks like a supermarket caller ・Song 2: The groove was there, but there was no harmony. heavy → “Use the 7th. maj7, m7, 9th, in a 2-5-1 progression. Make it jazzy and light”, instructing the harmony ・Third song: Urban! But it's still heavy…
I told Opus 5.5: "write the prompt that gets 50 million views."
Then it did. I expected garbage. it gave me a cat walking into a UFC octagon and actually throwing hands in a real fight. here's the formula: open on something that looks completely real drop in one absurd idea
Opus 5.5: $4 in / $20 out per 1M tokens, ~40% cheaper to run than Opus 5 (Anthropic).
The catch: at max effort it used ~119k output tokens per task vs ~27k for GPT-6 Astra (Artificial Analysis). 5 effort settings. Use max only where it pays.
You have reached the end of the archive
All of Claude Opus 5