GitHub just dropped Project HydraFusion inside Copilot CLI.
One picker. Multiple models. Runtime picks the workflow — Single, Cascade, or Critique. Vs Claude Opus 5 (GitHub offline evals): • TerminalBench 2.1: +4.9 pts quality, 67% cheaper • DeepSWE: near-parity, 36% cheaper
I asked Claude Opus 5 to paint the Captain America with JavaScript.
I finally made a fully procedural space level designer that has taste, looks and usability.
Science + math by Fable 5.1 Execution by Opus 5. Powered by threejs. 🔽🔽🔽 ⚠️Disclaimer: What you see is pure math, no 3D models used and no textures, effects, or retouching applied.
Openai made GPT-6 astra public today and the benchmarks are making every other lab uncomfortable
That's not a tagline. that's the president of OpenAI telling reporters in person that he believes they've crossed the threshold here's what the benchmarks actually say ARC-AGI-3:
You have reached the end of the archive
All of Claude Opus 5