Grok 4.6 just did something no other AI could.
It's the first model to beat Kimi K3 at real coding work. Claude Fable 5 tried. Failed. Claude Opus 5 tried. Failed. GPT Soul Max tried. Failed. Grok 4.6 walked past all of them. Here's what's inside: → It scored 61 on the
Grok 4.6 team cooked.
It came out of nowhere and overtook OpenAI Sol on our fully autonomous threat-hunting benchmark. 🎯 And this is not a benchmark built around hand-holding hints disguised as questions, fuzzy LLM-as-a-judge scoring, or unscalable emulations. We built it
This is what running a 24hr prompt to create an open-world version of GTA VI with Opus 5 looks like
Access prompt here: trying the prompt out myself (r!p to my usage tokens) sadly the developer hasn't made it playable yet hopefully i get the same
Grok 4.6 Kimi K3
Opus 5 It is true that Grok 4.6 is a good model considering speed, price, and intelligence. Not yet the best intelligence. Grok 4.7 seems to be something to look forward to.
You have reached the end of the archive
All of Claude Opus 5