FRAMEWIREIndonesiaUpdated Aug 25Live wire
0:00 / 0:00

Tried running the Qwen 3.8 2-bit model locally on my Book4 Pro running Omarchy✌️

It works, but I got around 1.2 tkn/sec ... way too slow to actually use it (expected).. this is pretty much the limit of what I can run on my current hardware... might just wait for 35B MoE model

ManuAug 25
0:00 / 0:00

Calculate the recent popular models and hardware data: M5 Max Qwen 3.8 27B MLX 4bit (dense model)

Basic short contextPP about 900-925 tok/sTPS 32-33 tok/s ~32k TPS and 28-29 tok/s In most cases, there is a significant improvement after turning on MTP or DFlash. The actual measurement of MTPLX is mostly 55-65 tok/sDFlash2, some people have reached about 70 tok/s M5 Ultra Today…

花椰菜Aug 2510
0:00 / 0:00

Warcraft III menu round 2

Gemini 3.1 Pro DSV4 Pro GLM 5.3 Grok 4.6 Fable Sol Kimi K3 I have opinions on which did the best. I think it's worth doing a 4 square comparison with Qwen 3.8 as well.

Loktar 🇺🇸Aug 2522
0:00 / 0:00

416 tok/s on a 3090. That is what promises for Qwen 3.6 35B-A3B, and it is the one column in that tool you should ignore.

The tool went round yesterday and it deserves the attention. Free, no signup, reads your GPU straight from the browser, 284 devices

BountyAug 253