Ultimate compression KV Cache Compression
What has changed in DeepSeek-V4.1-Flash’s Causal Encoder–Decoder? In a regular Transformer, each layer generates its own K and V based on the hidden states entering this layer. DeepSeek-V4.1-Flash's Causal Encoder–Decoder adjusts this dependency: the decoder's global KV is directly determined by the causal encoder…
DeepSeek V4.1 Flash made DOOM in Atomic Chat 🐳
Deepseek_ai dropped a new architecture for the flash model so we tested it by making Doom in an html file and it built a real raycasting shooter with gun mechanics and enemy waves Run AI models locally ->
Three snowy builds. One model.
DeepSeek V4.1 Flash passed our initial market, cabin and Supply Run checks through OpenCode Go. No extra repair round. Three small case studies, with tokens, times and limits in the video. Thanks for watching, fam!
DeepSeek V4.1 Flash is live on RunBiOS — Day 0.
1M context · 384K output · vision · tool calling · prompt caching at $0.08/M Input $0.40/M · Output $1.60/M. Agentic and coding runs resend context every step, and cached input costs 80% less. Get 50% extra credit on your wallet
You have reached the end of the archive
All of deepseek