Deepseek v4.1 Flash on 24GB Macbook.
I think I reached my limit, 1.36 tok/s. It is far from being usable but the feeling of having a SOTA model on my consumer machine is awesome. I started from 0.74 tok/s and almost double it. It was a great journey.
I'm using Codex as an orchestrator on Herdr, assigning implementation to DS and Agy, and reviewing the responses without permission.
510 GB model. 128 GB machine.
256k context, thinking on. DeepSeek-V4.1-Flash on one DGX Spark, no smaller model, no cloud. The trick: the model only uses 6 of 384 experts per token. So I tell the box what it should be good at, and it keeps only the experts that job actually
Which one is better, DeepSeek V4.1 Flash or GLM 5.3 Flash?
More cost-effective? I did a test to help you make a better decision on which model to choose.
You have reached the end of the archive
All of deepseek