A model nobody was talking about just closed the gap on Claude.
It happened quietly on August 12th, and most people slept on it. 😳 DeepSeek V4 Pro vs Claude Fable 5: → Terminal Bench 2.1: 87.9 vs 88.0 (basically tied) → AutomationBench: 31.8 vs 29.1 (DeepSeek wins) →
Deepseek Harness is still fresh and new but I have high hopes!
Deepseek V4 flash just beat its own pro model on all 9 agent benchmarks.
And the biggest upgrade came without adding more parameters. What changed: → DeepSWE jumped from 7.3 to 54.4 after retraining → Terminal Bench 2.1: Flash scored 82.7 vs Pro at 72.1 → Same
Kimi K3 is getting ridiculous at 3D generation.
Compared to Gemini 3.7, Grok 4.6, DeepSeek V4, and GPT-5.6; Kimi K3 is cheaper. One prompt turned into a full military armory with detailed props, weapon racks, animated monitors, and dynamic lighting.
You have reached the end of the archive
All of deepseek