Is model scaling the only source of agent improvement?
We henryqin1997 YaxinLu1997 VITAGroupUT VictorKaiWang1 are glad to share our work: We reach 95.3% Raw Accuracy, or a $15 Frontier Run (88.8% with DeepSeek-v4-Flash, matching GPT-5.6 Sol Max), on Terminal-Bench 2.1 via
DeepSeek-V4-Pro is here — and it's redefining cost-performance.
✅ 1.6T total parameters / 49B activated ✅ 1/10th the price of Claude Opus 5 ✅ FIM code completion mastery ✅ Full Agent toolchain support We tested it head-to-head against Claude Opus 5.
The Most Comprehensive DeepSeek Harness Plugin Store—Built Right Into DSH.
Deepseek_ai Finding and installing DSH plugins is still difficult. So our founder Lafe8088 built DSH Market, a plugin store you can install directly inside DeepSeek Harness. It indexes 7,210
Deepseek V4 pro just scored 87.9 on terminal bench
This AI model was built to actually finish tasks inside a computer terminal.
You have reached the end of the archive
All of deepseek