Qwen 3.8 Flash and GLM 5.3 Flash just dropped
So TonyD2Wild breaks them down against the reigning DeepSeek V4 Flash on 2 DGX Sparks (GB10). 60 to 80 tok/s on two, up to 120 on four.
The good performance of the Hermes Agent AI model with DeepSeek V4 Flash 0731 on local hardware stands out.
Qwen 3.8 Flash Next and GLM 5.3 Flash will be tested soon.
Alibaba just turned AI from a chatbot into a worker.
And the most important part isn’t the 2.4 trillion parameters. It’s what Qwenwork can finish without you. What it actually does: → Researches the task, writes the document, creates images and builds the web page →
Your phone can't run the new Qwen 3.8 27B model.
Not alone, anyway. I built SwarmLLM. I taught the model to share. It splits a large language model across the devices around you and runs it in browser tabs. In this demo it's Qwen 3.8 27B on my MacBook and my iPhone together.
You have reached the end of the archive
All of qwen38