FRAMEWIREIndonesiaUpdated Aug 19Live wire
0:00 / 0:00

DeepSeek V4 Pro just got dangerously close to the expensive models.

The benchmark gap is now tiny on agent work. Here’s the headline: → Terminal Bench 2.1: 87.9. → Claude Fable 5: 88.0. → Earlier DeepSeek Pro: 72.1. → New DeepSeek Pro: 87.9. → Humanity’s Last Exam with

Julian Goldie SEOAug 191
0:00 / 0:00

The three numbers from one model are not a deepseek story, they are every leaderboard you have ever read

87.9, 78.7, mid fifties. Same weights, same test. The harness moved and the score moved thirty points. Berkeley put a size on this in April. Researchers stress tested the

StarHazeAug 19
0:00 / 0:00

After Liang Wenfeng raised his price, he finally had a common language with Sam Altman. Sam: Welcome to the "reasonable profit" club!

Uncle Liang: The computing power of V4 Flash is really unbearable after it goes online. If it doesn’t increase, it will become a free server. Sam: I understand. From Liang Sheng to Liang Shu, I grew up very quickly.

SuSu_酥酥👅Aug 199
0:00 / 0:00

Just ran my first multiplayer duel game with my son.

I am shocked 98% of features work like a charm. We even had a major fight for the central planets and everyone just lost their starter fleet 😭 All done by spec driven development by #claude and #deepseek. Now mostly claude

Dmytro GladkyiAug 19