April: V4-Pro launches at $3.48 per million output
May: DeepSeek cuts it 75% and calls the discount permanent today: $3.96 at peak the price is now higher than it was before the discount 😭
Local showdown: Qwen vs Deepseek
Fireworks one-shot: "build a fireworks show over a city, single HTML file, no libraries." two locals on DGX Sparks, temp 0.6, one attempt, no edits. DeepSeek-V4-Flash (2x GB10): 1m20s, 5,214 tokens, Qwen3.8-27B NVFP4 (1x GB10, xhigh): thought
This is my assistant Szayelaporro.
What it can do for now: Roleplay chat as Szayelaporro. I can chat in Telegram and via a PC window implemented as an animated "Desktop Pet". The "Desktop Pet" simulates a separate Telegram DM. Additionally, I can communicate with him by voice
DGX Station: this is just the start of what it is possible to do (and why I claim DwarfStar could be *the* Station inference engine).
DeepSeek v4 PRO Q2 with routed experts split among VRAM / RAM with kernels optimized for this peculiar setup. 45 t/s but can go faster.
You have reached the end of the archive
All of deepseek