One benchmark. Nine sub-metrics.
Four frontier models. No winner. We scored GPT-5.6, Gemini 3.1 Pro, Claude Opus 5, and Muse Spark 1.2 on RxScribe-Bench, then broke every axis into its parts.
I like to use AI for visualization.
I’ve talked with Fable about creating DIY hardware. Together, we created a parts list and a plan for how to assemble it. Then I asked Fable to guide Opus 5 sub-agents in creating a visualization website.
Half of my timeline is angry at Claude Opus 5.
Too verbose, too eager to overengineer, walls of text. The fixes are sitting in Anthropic's own docs. Nobody reads docs, so I did it for you. Three things, all quoted from their prompting guide. 1. Verbosity is real, and the knob
This multiplayer FPS was built in under 2 hours
Opus 5 created a Minecraft-inspired shooter with: 5 maps 12 weapons multiplayer voice lines + SFX full gameplay systems ElevenLabs handled the audio We went from “AI can prototype games” to “AI can ship multiplayer
You have reached the end of the archive
All of Claude Opus 5