5/ Next dataset: what one dollar buys.
Four models, one identical task, measured by bridgebench. - GPT-6 Luna: over 100 runs. Sol: 12. Astra: 1.6. Opus 5.5: 0.8 - Time per run: Luna 43 seconds, Opus 5.5 nine and a half minutes - 175 seconds, 43 cents, one turn, no tool calls
2/ The writing test. Everyone says the slop is gone, so I checked.
Two briefs, a release email and a launch post, to Opus 5, Opus 5.5 and GPT-6 Sol. Three drafts each, 18 pieces. Counted em dashes and 37 stock phrases. - Em dashes: zero in 18. Stock phrases: one "excited to",
1/ The scoreboard. Artificial Analysis rolls ten evals into one number, and Opus 5.5 leads it.
Opus 5.5 at 57.6, Fable 5.1 at 53.4, Opus 5 at 50.8, GPT-6 Sol at 47.5 - Sol scores level with GPT-5.6 on this index, at half the list price of Opus 5.5 - Running the whole index
Opus 5.5 came out Tuesday and took the week.
What it means if you pay for Claude: - Your plan goes further. People run it all day and use around 10% of their weekly limit - 40% cheaper than Opus 5 and 30% faster. Claude Code made it the default on Pro and Team - Cursor, Devin
You have reached the end of the archive
All of Claude Opus 5