FRAMEWIREIndonesiaUpdated Aug 19Live wire
0:00 / 0:00

Kimi k3 outperformed claude opus 5 by 57x

Fede(URU) 🇺🇾Aug 19
0:00 / 0:00

Claude Opus 5 vs Opus 4.8 vs Opus 4.7 vs Opus 4.6 - on italian architecture

Fede(URU) 🇺🇾Aug 19
0:00 / 0:00

Kimi 2.6's agent swarm quietly turned "hire a research team" into a single prompt...

300 sub-agents, one input, coordinated across thousands of steps, and the model plans the whole workflow itself - no roles to assign, nothing to configure here's what people are actually

rewindAug 1921
0:00 / 0:00

Trying out some visualisations for autoresearch to understand how models approach this.

Inferring and plotting references to previous trials is already interesting: both bigger and higher-effort models seem to make more longer-range inferences. Claude Opus 5 at low vs xhigh:

Sam GijsenAug 192