Benchmaxxing exists because nobody outside a lab can tell a good benchmark from a popular one, so popularity wins.
Nickheiner, VP of RL Environments at Surge AI, takes that apart in "When Will The Benchmaxxing Plague End?", on aiDotEngineer's YouTube. It's a working tour of
Qwen3.8 Max vs Opus 5 (Max)
Claude Code × Opus 5's Excel creation ability was too high...
I was able to automatically generate a 3-year PL with 19 tabs and 36 months. If you were to do this work before generation AI, it would probably take a week. The functions are perfect, and the database and cell references are very beautiful. Opus from Fable 5
Claude Opus 5, GPT-5.6 Sol and Kimi K3 - all in one place, same prompt, real answers.
Arena lets you throw the exact same prompt at top models and compare the answers side by side. I tried: “Give me 3 reasons why AI stocks could crash.” • same question. • different models.
You have reached the end of the archive
All of Claude Opus 5