Grok 4.6 team cooked.
It came out of nowhere and overtook OpenAI Sol on our fully autonomous threat-hunting benchmark. 🎯 And this is not a benchmark built around hand-holding hints disguised as questions, fuzzy LLM-as-a-judge scoring, or unscalable emulations. We built it
This is what running a 24hr prompt to create an open-world version of GTA VI with Opus 5 looks like
Access prompt here: trying the prompt out myself (r!p to my usage tokens) sadly the developer hasn't made it playable yet hopefully i get the same
Grok 4.6 Kimi K3
Opus 5 It is true that Grok 4.6 is a good model considering speed, price, and intelligence. Not yet the best intelligence. Grok 4.7 seems to be something to look forward to.
Claude opus 5 is officially the dumbest genius in this entire experiment
Moon dev: "the smartest model on paper is dead last and bleeding faster than everyone else combined" "it just slammed the CLOSE button on BTC at the worst possible second and i physically winced"
You have reached the end of the archive
All of Claude Opus 5