FRAMEWIREIndonesiaUpdated Sep 22Live wire
0:00 / 0:00

Have you ever played an AI benchmark?

Here, have fun: ARC AGI 3 is no longer a grid puzzle. François Chollet and ARC Prize have turned them into interactive mini worlds. The agent doesn't know the rules, has to act, read feedback, build a world model and

Dr.-Ing. Martin SchieleSep 22
0:00 / 0:00

I asked Claude Opus 5.5 to create a webapp that turns a picture into vectors

Then blows it apart into a 3D exploded view Absolutely amazing! Took a lot of tokens though, opus 5.5 doesn't like to give up or do the job half heartedly, it's a hardworker. You can explore here:

Pradeep KapoorSep 229
0:00 / 0:00

What model is the best at finding vulnerabilities inside agents?

👽 Enoki Labs's attacker runs on a harness plus a model. The harness is ours. We put 10 models through it, against two agents with 23 known vulnerabilities. One writes code, one moves money. Only the model

Enoki LabsSep 223
0:00 / 0:00

Grok 4.7 kept the price.

It did not keep the token count. Artificial Analysis scored Grok 4.7 at 46 on the Intelligence Index, +2 over 4.6. On the Coding Agent Index it gained +9 and now sits right behind Claude Opus 5. Then the token line. Output tokens per Intelligence Index

KravnSep 22