So I was trying to write some benchmark tests for AI models.
Base line, golden line, full eval on multi run setup. Goal was to break Astra and Fable with pass rate of 20-40%. I've iterated with Astra and Fable to actually write evals as I designed them - it keeps breaking on
Thought I'd share my weekend Astra project
A "graphic novel harness" for exploring a sci-fi world and story. Links in the reply below — reply and I'll DM you the password. What it is: A generative graphic novel that extends as you read it, forward or inward. If you highlight a
An open-source model just landed 0.1 points behind claude opus 5 on workflow automation
Nex-N2.5 Max scored 50.2 on AutomationBench v1.0.6. Opus 5 scored 50.3. That gap is small. The model is not. The family starts with a 35B Mini, moves to a 397B multimodal Pro, and ends with
When opus 5 tells me the flux capacitor belongs inside the tailwind config as a optional arg inside the TNMNT method.
You have reached the end of the archive
All of Claude Opus 5