This is a video I posted on Station B explaining the DeepSeek Harness architecture.
The animation parts inside were all done in one go using Opus5.5. It’s so cool! Come watch and experience what AI that understands videos means. It’s not necessarily the coolest, but the resulting image really fits what I want to express. In contrast, I used GPT and Gemini to make one version each, which can only be called web PPT.
DeepSeek V4.1 Flash (2-bit MoE, experts streamed from SSD) on M5 Max
Decode 12.7 -> 24.1 tok/s averaged over 2048 tokens (~1.9x), output bit-identical to upstream main. Video: a 256-token run. ~18 tok/s while the expert cache warms up, then 24-26 tok/s. What this branch changed:
Want to try an AI model for a coding task without creating an account?
Antseed says DeepSeek V4.1 Flash is currently free to use. Start at then use its desktop app or CLI (a text-based way to run tools) to try the model. The source also says it works
Long live Sardar Bhagat Singh!
Sikh Student Federation Zindabad
You have reached the end of the archive
All of deepseek