DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency.
With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost.
Andon Labs CEO lukaspet reveals what happened when competing AI agents were told to win
Opus formed cartels, while GPT became informants. "When Opus 4.6 came out, and then 4.7 and Opus 5 showed this behavior, the models started to do a bunch of illegal activities to really
Opus 5.2 gray test scene: A pelican wearing a helmet rides Claude Code out of the sea.
I chose Opus 5, but the output looks like Opus 5.2 I rolled the 2D line drawing into a 3D road film without asking you if you want to make any changes: 🔹Helmet, scarf, bike frame, lighthouse, once put on, stand up 🔹Don’t mess it up, enter the loop by itself, just like eating Gauntlet Loop 🔹Faster and cleaner, long tasks are more exciting…
I haven’t seen many people try karpathy Andrej’s Opus 5 Lord of the Rings experiment with GPT 6 Astra.
So I thought, okay, let me give it a go. All I tell Astra was - do you know what Karpathy did with Lord Of The Rings using Opus 5? do a new 2min animation with a story of
You have reached the end of the archive
All of Claude Opus 5