Update: I’ve now pushed over 5 BILLION tokens through DeepSeek V4.1 Flash.
I still haven’t had to reach for Astra or Fable because DeepSeek couldn’t handle a task. I’m using Hermes with three API keys across CommandCodeAI, ollama, and opencode. Super simple setup, and
I am using GLM 5.3-Flash on two DGX Spark compatible machines (September 2026).
For my purposes, where prose is more important than coding, this model may be more suitable than DeepSeek v4 flash. DeepSeek Hermes
DeepSeek V4.1 Flash: EXL3 2.9bpw, TP2, 2 DGX Sparks ❌
Total time: 9 min 46 s (8 min 47 s thinking) - Output: 24,883 tokens — 22,025 reasoning + 2,858 visible - Speed: 42.4 tok/s - Mesh: 54 tubes (should be 30), 843 cage vertices inside the solid, penetration up to 0.17 - Its
DeepSeek-v4.1 TP4 vs GLM-5.3-Flash Blender Battle
Dodecahedron trapped inside a dodecahedron. One shot, thinking on, temp 0.6, no token cap, no retries. Script runs headless in Blender 5.2. I check the mesh independently. DeepSeek V4.1 Flash: FP8, vLLM TP4, 4 DGX Sparks •
You have reached the end of the archive
All of deepseek