And this is why we run Jinx and Luna's cognitive systems on Qwen 3.8 Flash Next on our 2x DGX Spark configuration.
Here's a staggered-prefill concurrency comparison between it, Deepseek v4 Flash Vision, and GLM 5.3 Flash.
Top-5 Best Value #LLM Models of the Day at UTC-04
| Model | #AAII | Price | | GLM 5.3 Flash | 42 | $0.09 | | GLM 5.3 | 45 | $1.23 | | GPT-5.6 Luna | 38 | $0.18 | | Gemini 3.8 Flash | 41 | $0.57 | | DeepSeek V4 Flash 0731 | 34 | $0.06 |
I tested DeepSeek V4 Flash with Vision (DS4FV) using the two DGX Sparks together with TP‑2, with four staggered requests.
It took 57.14 s to complete all four, compared to 36.35 s for Q38FN and 83.87 s for GLM53F. If you mean two independent instances of DeepSeek, one per
Yesterday, DeepSeek announced a temporary trial version of DeepSeek V4.1 Flash in the WeChat group.
Now we can add a model configuration to DeepSeek Harness and use it. I tried it myself and the speed is very fast, reaching 300-400 tps. In the past, when I was thinking about it in max mode, there was a lot of crackling output, but now even if I turn on max
You have reached the end of the archive
All of deepseek