I’m ready for DeepSeek V4.1 Flash to fit on my DGX Spark…
But until it does, I’m enjoying good quality, affordable API access. Open source local AI is the future!
Is it possible? 😱 Deepseek v4.1 flash score equal to 98% of GPT got 6 points
Now a site comes, you register, it gives you 6 dollars and you chat right there, with this model, you can spend 5.6 million tokens per hour on this site. Step 1: Go to this site
DeepSeek V4.1 Flash: open weights, MIT, 552B params activating 8B/16B per token.
KV cache 890 bytes/token, ~1/4 of the previous Flash - the weights did NOT shrink. On Terminal-Bench 2.1 at maximum effort it scores 90.6 vs Opus 5.0's 89.1, but 30.0 and 31.2 on 3.0 and 4.0.
4 DGX Sparks running DeepSeek V4.1 Flash
6 concurrent sessions, 170.73 tok/s coding throughput. 🔥 1,200 output tokens in 7.028 seconds. Warm single-stream: 💻 Code: 77.61 tok/s 🧮 Math: 74.08 🧠 Reasoning: 48.75 ✍️ Prose: 33.02 6-stream aggregate: 💻 Code: 170.73 tok/s 🔢
You have reached the end of the archive
All of deepseek