The gap between “here’s the design” and “here’s the working app” is getting very small.
Grok 4.6 took #1 on the VISTA leaderboard, beating Claude Fable 5, Opus 5, and GPT-5.6 Sol. It has to look at a real Figma design, figure out the structure, recreate the visuals, and build
Can't believe claude made this remotion with one prompt.
Opus 5 is definately the best for remotions beating fable 5. I used this remotion skill to get this output check it out:
Oh? ! Claude Design x Opus 5 Mont Saint Michel
Isn't it pretty good? ! Well, if you search for the real thing, that's it 😂
Harness-bench with e2b and herdrdev
Compare Claude Opus 5 vs Codex GPT 5.6 Sol performance when debugging an ambiguous parsing edge case. Give each its own E2B sandbox, so you can spin up as many harness evals as you want in parallel. Results we got from our run: - Opus used
You have reached the end of the archive
All of Claude Opus 5