DeepSeek V4 Pro just got dangerously close to the expensive models.
The benchmark gap is now tiny on agent work. Here’s the headline: → Terminal Bench 2.1: 87.9. → Claude Fable 5: 88.0. → Earlier DeepSeek Pro: 72.1. → New DeepSeek Pro: 87.9. → Humanity’s Last Exam with
The three numbers from one model are not a deepseek story, they are every leaderboard you have ever read
87.9, 78.7, mid fifties. Same weights, same test. The harness moved and the score moved thirty points. Berkeley put a size on this in April. Researchers stress tested the
After Liang Wenfeng raised his price, he finally had a common language with Sam Altman. Sam: Welcome to the "reasonable profit" club!
Uncle Liang: The computing power of V4 Flash is really unbearable after it goes online. If it doesn’t increase, it will become a free server. Sam: I understand. From Liang Sheng to Liang Shu, I grew up very quickly.
Just ran my first multiplayer duel game with my son.
I am shocked 98% of features work like a charm. We even had a major fight for the central planets and everyone just lost their starter fleet 😭 All done by spec driven development by #claude and #deepseek. Now mostly claude
You have reached the end of the archive
All of deepseek