Kode Claude × Opus 5.5
Buat video pengenalan untuk layanan AI fiktif “setsuna” dalam 30 detik Temanya adalah "Saatnya menunggu respon AI". AI diatur untuk hanya berpikir sekitar 6% dari 2,4 detik yang diperlukan untuk merespons. ・Buat grafik batang rincian waktu tunggu dan potong satu per satu.
5/ Next dataset: what one dollar buys.
Four models, one identical task, measured by bridgebench. - GPT-6 Luna: over 100 runs. Sol: 12. Astra: 1.6. Opus 5.5: 0.8 - Time per run: Luna 43 seconds, Opus 5.5 nine and a half minutes - 175 seconds, 43 cents, one turn, no tool calls
2/ The writing test. Everyone says the slop is gone, so I checked.
Two briefs, a release email and a launch post, to Opus 5, Opus 5.5 and GPT-6 Sol. Three drafts each, 18 pieces. Counted em dashes and 37 stock phrases. - Em dashes: zero in 18. Stock phrases: one "excited to",
1/ The scoreboard. Artificial Analysis rolls ten evals into one number, and Opus 5.5 leads it.
Opus 5.5 at 57.6, Fable 5.1 at 53.4, Opus 5 at 50.8, GPT-6 Sol at 47.5 - Sol scores level with GPT-5.6 on this index, at half the list price of Opus 5.5 - Running the whole index
Sudah sampai ujung arsip
Semua Claude Opus 5