Our internal writing benchmark uses 3 AI judges: Claude, GPT and DeepSeek.
So do they favor their own family? The GPT judge scores OpenAI models 2.3 points higher (out of 100) than the other two judges do. The DeepSeek judge gives DeepSeek models +1.7. The Claude judge gives
I'm getting better and better results out of XiaomiMiMo MiMo-v2.6-flash
Tested with a bit difficult prompt, to create a Koi-Kingdom, difficult cause qwen3.8-flash and deepseek-v4.1-flash did struggle this one. But not MiMo. One-shot, three.js, no external assets, single html
MiMo-v2.6-flash is proving to be really good model overall.
It's performing on par with DeepSeek-v4.1-flash on most of my tests. "A Three.js scene of a mountain with a long stair of red torii gates climbing into mist, a shrine at the peak." This is
I made two AIs play Snake
Same question every move. Straight, left, or right? DeepSeek V4 Flash writes a paragraph, then picks. Jev just picks After 30 seconds: 122 moves vs 16 244 ms vs 1.71 s 13 apples vs 2 1,150 characters of slop vs 0 Jev is TypeSafe's decision model.
You have reached the end of the archive
All of deepseek