Calculate the recent popular models and hardware data: M5 Max Qwen 3.8 27B MLX 4bit (dense model)
Basic short contextPP about 900-925 tok/sTPS 32-33 tok/s ~32k TPS and 28-29 tok/s In most cases, there is a significant improvement after turning on MTP or DFlash. The actual measurement of MTPLX is mostly 55-65 tok/sDFlash2, some people have reached about 70 tok/s M5 Ultra Today…
Warcraft III menu round 2
Gemini 3.1 Pro DSV4 Pro GLM 5.3 Grok 4.6 Fable Sol Kimi K3 I have opinions on which did the best. I think it's worth doing a 4 square comparison with Qwen 3.8 as well.
416 tok/s on a 3090. That is what promises for Qwen 3.6 35B-A3B, and it is the one column in that tool you should ignore.
The tool went round yesterday and it deserves the attention. Free, no signup, reads your GPU straight from the browser, 284 devices
In today’s mainstream model, creativity is no longer a shortcoming. Gemini 3.7 Flash, Grok 4.6
In today’s mainstream model, creativity is no longer a shortcoming. Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, Claude Opus 5, the same Ace of Spades, required to be drawn into the card. This kind of same-question test can better see the differences in models than running scores. The question is not who draws better, but who thinks more like a human being. 🔥…
You have reached the end of the archive
All of qwen38