I'm running Qwen 3.8 Flash Next at 25-30 tokens/sec with MTP
But this speed is measured after a 20k-token context is loaded. For a single RTX 3090 and 64 GB DDR4, that's pretty remarkable. This isn't about the premature speed, which is usually just the first 2k tokens; those
Qodercli + /goal + computer use + qwen 3.8 max 0902
/goal use computer use to checkout weather app and recreate it. compare with the original and keep polish util they feel the same. including the visual and interacive.
The big showdown documented
All based on forensic analysis of all 6 runs - multiple hours of zcode session materials Milions of tokens. Qwen 3.8 27b, Qwen 3.6 35B A3B, Ornith 1.5 35B A3B and Gemma 4 26B A4B All analyzed and compared All 4 bit All on a M2 Max Macbook Pro
COMPETITION 1 RESULTS: — I would actually argue Muse Spark 1.3 could be tied with Qwen 3.8 Max given cost and speed…
My verdict is that GLM 5.3 is all around best when balancing performance on this visual task, cost, speed and token usage. Muse Spark 1.3 was insanely cheap
You have reached the end of the archive
All of qwen38