FRAMEWIREIndonesiaUpdated Sep 10Live wire
0:00 / 0:00

To understand attention, you can first distinguish the roles of K and V

K participates in calculating the matching score between Q and each entry. V is weighted and summed according to attention weights to form the output. The core attention of DeepSeek-V4.1-Flash uses the same storage vector and plays the role of K and V at the same time. Use Q first…

Dongxi 东锡 NLPSep 107
0:00 / 0:00

DeepSeek V4.1 Flash made DOOM in Atomic Chat 🐳

Deepseek_ai dropped a new architecture for the flash model so we tested it by making Doom in an html file and it built a real raycasting shooter with gun mechanics and enemy waves Run AI models locally ->

atomic.chatSep 103
0:00 / 0:00

Three snowy builds. One model.

DeepSeek V4.1 Flash passed our initial market, cabin and Supply Run checks through OpenCode Go. No extra repair round. Three small case studies, with tokens, times and limits in the video. Thanks for watching, fam!

TonkaToyXL | AI Test CabinSep 10
0:00 / 0:00

DeepSeek V4.1 Flash is live on RunBiOS — Day 0.

1M context · 384K output · vision · tool calling · prompt caching at $0.08/M Input $0.40/M · Output $1.60/M. Agentic and coding runs resend context every step, and cached input costs 80% less. Get 50% extra credit on your wallet

UltraSafe AISep 102