DeepSeek V4.1 Flash weights are out!
We are shipping day-0 inference and RL support in SGLang and Miles. V4.1 extends the V4 stack with compressed KV shared across layers, a two-stage sparse indexer, and a 196B Engram lookup memory. It is natively multimodal with 552B backbone
GPT6 astra reduces intelligence, my PRO20X, Pelican Cycling actually drew this for me?
The second one is the latest deepseek 4.1 flash. I hope OPENAI will fix the behavior of labeling normal users for no reason and demotivating them as soon as possible. Such rude actions will personally drive away your users! sama thsottiaux
GPT-6 Astra, DeepSeek-V4.1 Flash, Fabel 5.1...
As long as a model is released, everyone will test the Pelican bicycle, but why do they actively ignore Xiaomi Mimo? But Mimo seems to be the first one I’ve ever seen who stepped on the reverse wheel😅
When it comes to DeepSeek, everyone only remembers one number: DeepSeek-V3 training cost $5.576 million.
Now it has just raised approximately US$7.4 billion and is preparing for new financing and IPO. Many people’s first reaction is: What about the promised low cost? It's like seeing that a tank of fuel for a racing car is not expensive, and then thinking that running a team is not expensive either.
You have reached the end of the archive
All of deepseek