40 AI agents entered
Only 8 made it out I put 20 Grok 4.6 agents against 20 Claude Opus 5 agents inside a 60-second PvP arena. Both teams began in vertical ranks, then split into four linked five-agent squads coordinated through a shared command bus. Every fighter could call
Anthropic just launched Claude Fable 5.1 and Mythos 5.1.
On Anthropic’s evals, Fable 5.1 scores 55.8% on Terminal-Bench 4.0 vs 42.0% for Fable 5, 52.3% for Opus 5 and 37.3% for GPT-5.6 Sol. Mythos 5.1 reaches 60.9%. Fable is GA today.
Would be interesting to eval how opus 5 performs on tasks autonomously vs iterative tasks
GLM 5.3, GPT 5.6 Luna, Claude Opus 5, Grok 4.6, DeepSeek V4, Kimi K3, and a cloud browser for agents - all FREE.
DuckDuckGo: GPT 5.6 Luna and older models, no signup: LM Arena: free side-by-side, Opus 5, GPT-5.6 Sol, Grok 4.6, Qwen 3.8 Max:
You have reached the end of the archive
All of Claude Opus 5