Dylan Patel reveals DeepSeek tunes its models so tightly to Nvidia GPUs that they run far worse on TPUs and Trainium
"A simple low hanging sort of optimization that we tend to do with our models is we make mixture of experts models. People do mixtures of experts, but they make
DeepSeek-V4.1-Flash has did one of the best cloth sim i have ever seen
It one-shot'd this cloth hanging sim, did the best texture work, good physics, pretty good result. Stay tuned, more video's coming soon!
DeepSeek V4.1 Flash is already hitting nearly 400 tokens/sec.
DeepSeek has started internal testing of V4.1 Flash, featuring a new architecture, built-in multimodality, higher performance, faster inference, and lower operating costs. API access: keep the same base_url and set
Here is a 15 second long example of local Qwen 3.8 Q3 XSS (massive aura loss from just typing that out) vs DeepSeek V4.1 Flash limited preview
Tfw you are a local tokencel
You have reached the end of the archive
All of deepseek