Four models. Three tasks.
ZERO FIXES. everything looked the same at the start. but after one prompt, the difference became obvious. GPT-5.6, Claude Opus 5, Qwen 3.8 Max and Grok 4.6 had to build a premium product page, a playable browser game, and an animated rainy city scene
Opus 5 + higgsfield just killed the design team.
Saved the $35K invoice. One operator. A staffed web design shop doing $35k a month usually keeps around $10k after payroll. The same revenue run by one operator on a tight AI stack keeps almost all of it. Same clients. Same
Nvidia's coding agent solved the ARC-AGI-3-BENCHMARK complete
ARC-AGI-3 places an AI in an alien environment and doesn't explain anything to it. There are no rules to read and no goal written anywhere - just a grid, a few buttons and the task of figuring out what's here yourself
Anthropic just deleted more than 80% of the system prompt for Claude Code.
In this 2-minute breakdown, a DevOps engineer explains exactly why giving strict rules to newer models like Claude Opus 5 and Claude Fable 5 actually hurts their performance. they found that old
You have reached the end of the archive
All of Claude Opus 5