You can use DeepSeek V4.1 Flash for FREE right now 😳
The model is showing up on xplabs: - deepseek-v4.1-flash how to try it: 1. go to 2. find deepseek-v4.1-flash 3. select it and start chatting no need to pay for the models you can try it right now
35B MoE. 8GB of VRAM.
39.3 tokens/s decode. That is Qwen3.6-35B-A3B NVFP4 on a laptop RTX 4060 — not an H100 rack. The hot experts sit in a GPU cache. The rest of the pool lives in system RAM. The paper’s coding-agent run even clears the 33 tok/s Codex median they cite. 🆕
Horse tinder one shot Deepseek V4.1 Flash.
I told it to make a video too and it did!
Every time a new model such as DeepSeek V4.1 Flash comes out, founder Liang Wenfeng seems to be forced to do this dance for the time being.
You have reached the end of the archive
All of deepseek