Kimi k3 outperformed claude opus 5 by 57x
Claude Opus 5 vs Opus 4.8 vs Opus 4.7 vs Opus 4.6 - on italian architecture
Kimi 2.6's agent swarm quietly turned "hire a research team" into a single prompt...
300 sub-agents, one input, coordinated across thousands of steps, and the model plans the whole workflow itself - no roles to assign, nothing to configure here's what people are actually
Trying out some visualisations for autoresearch to understand how models approach this.
Inferring and plotting references to previous trials is already interesting: both bigger and higher-effort models seem to make more longer-range inferences. Claude Opus 5 at low vs xhigh:
You have reached the end of the archive
All of Claude Opus 5