FRAMEWIREIndonesiaUpdated Sep 28Live wire
0:00 / 0:00

1/ The scores, from Anthropic's own page.

Sonnet 5.5 against Opus 5.5. - Terminal coding: 70.6% to 66.4% - CursorBench: 55.5% to 57.8% - Computer use: 80.1% to 81.8% - Knowledge work, GDPval-AA: 1844 to 1846. GPT-6 Sol: 1487 - Reading charts: 61.6% to 64.4%. Sonnet 5 scored

Parker RexSep 283
0:00 / 0:00

Claude Sonnet 5.5 is out.

It costs the same as GPT-6 Sol and half of Opus 5.5. On Anthropic's terminal coding test it scores 70.6%. Opus 5.5 scores 66.4%. Sonnet 5 scored 10.3%. Thursday I called Sol the better deal. It has a challenger at the same price.

Parker RexSep 2811
0:00 / 0:00

Sonnet 5.5 just dropped.

💀 30%+ faster, up to 30% lower cost per task, and the benchmark jump from Sonnet 5 is looking serious. The wild part? Sonnet 5.5 is getting surprisingly close to Opus 5.5 on some workloads. Anthropic just made the middle tier a lot more dangerous.

EdgexSep 282
0:00 / 0:00

Xiaomi's open weight model matches Claude Opus and GPT 5 across most benchmarks and you can download it and run it on your own server

"Xiaomi put out Mimo On September twenty second Two days ago. Mimo Pro performs on par with Claude's Opus V, which is like this insane model that

Worth WatchingSep 28