DeepSeek V4 Flash: the free model that beats its own Pro version.
Same architecture. Same parameter count. DeepSeek just retrained it. Now it wins on all 9 agent benchmarks against the bigger paid model. → DeepSWE: jumped from 7.3 to 54.4 on the SAME model → Terminal Bench:
Grok 4.6 ranks #1 on RuntimeWire’s Newsroom Reliability v0.2 benchmark with a score of 0.79
Beating GPT-5.6 Sol, Claude Opus 4.8, Gemini and DeepSeek.
DeepSeek V4 FULL 1 Hour 50 min Course
Kimi K3 Qwen 3.8 Max
DeepSeek V4 Pro 0831 I am surprised that the performance of Qwen 3.8 Max is better than I expected. But it is the most expensive and takes a long time. DS V4 Pro failed.
You have reached the end of the archive
All of deepseek