Now runs with 4 threads in parallel. The prefill cache is also working perfectly.
satisfaction. Claude Code's restrictions have been reinstated, but it still works well enough to continue using it. DeepSeek v4.1 on A100x4
Mandarin Assistant - an inference router across China’s frontier AI models.
The problem is simple: there is no single best model anymore. DeepSeek, Kimi, Qwen, GLM and MiniMax all behave differently across coding, long-context, research and cost. Mandarin Assistant fixes it.
And ofc Astra absolutely mogs it but deepseek is like 1/10th the price on api
Wondered how deepseek v4.1 flash would do on this, man this model is really goooooood
You have reached the end of the archive
All of deepseek