Exploring Microsoft Foundry (AIFoundryDevs) to run DeepSeek V4 Flash and GPT-Realtime-2.1 with zero data retention at scale…
Availability is just *abysmal*.
Based on using codex, I made a short video introducing the dsv4.1 flash architecture.
Ran FlappyBench on DeepSeek V4.1-Flash, GLM 5.3, and Kimi K3 with the same /design prompt.
🔹 DSV4.1-Flash: 9/10 · $0.0089 one-shot with the lowest cost 🔹 GLM-5.3: 9.5/10 · $0.0184 · nailed the UI but cost 2x more 🔹 Kimi K3: 8/10 · $0.0740 · hard-coded gameplay and 8x more
Got deepseek V4.1 flash running locally on a 16GB m1 mac mini
Original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)
You have reached the end of the archive
All of deepseek