DeepSeek V4.1 Flash: open weights, MIT, 552B params activating 8B/16B per token.
KV cache 890 bytes/token, ~1/4 of the previous Flash - the weights did NOT shrink. On Terminal-Bench 2.1 at maximum effort it scores 90.6 vs Opus 5.0's 89.1, but 30.0 and 31.2 on 3.0 and 4.0.
4 DGX Sparks running DeepSeek V4.1 Flash
6 concurrent sessions, 170.73 tok/s coding throughput. 🔥 1,200 output tokens in 7.028 seconds. Warm single-stream: 💻 Code: 77.61 tok/s 🧮 Math: 74.08 🧠 Reasoning: 48.75 ✍️ Prose: 33.02 6-stream aggregate: 💻 Code: 170.73 tok/s 🔢
You have reached the end of the archive
All of deepseek