DwarfStar running DeepSeek v4.1 Flash on a 128GB M5 Max.
I didn't expect with SSD streaming it could be so fast. Recent SSD streaming changes to retain the right experts surely helped, but also maybe DS4.1 uses the same experts more. Will push online when ready QA > ASAP.
DeepSeek 4.1 Flash vs. GPT-6 Astra + Higgsfield in a black hole simulation.
DeepSeek V4.1 Flash is FREE on Apinex 😳
And they have more free models too: - DeepSeek V4 Pro - DeepSeek V4 Flash - Gemini 3.8 Flash - GLM 5.3 Flash - Muse Spark 1.3 - GPT-5.6 Luna - Qwen 3.8 Max check them here: free access is live 👀
DeepSeek-V4.1-Flash on four DGX Sparks, single stream, live 🐋
The predecessor (V4 Flash + DSpark) peaked around 50 to 70 tok/s on code. V4.1 is 552B total, 16B active on the way out, 8B active on prefill, and about half the checkpoint is n-gram memory tables that we keep on
You have reached the end of the archive
All of deepseek