I got 20% improvement in token/second with Qwen 3.8 Flash Next.
The caviat it's my custom fork. I added mtp support but I'm trying to improve vanilla first.
Qwen 3.8 27B
Thinking vs Code Some say this model thinks a lot... 🤣
Qwen 3.8 27B running locally on an RTX 5090 beats Opus 4.8 in personal benchmarks at up to 200 tokens…
Qwen 3.8 27B running locally on an RTX 5090 beats Opus 4.8 in personal benchmarks at up to 200 tokens per second with no internet, subscription or API required and is now powering a Hermes agent working 24/7 for free.
Here is my first local Qwen 3.8 Flash Next result.
I am disappointed to learn that the current releases that can run on my Mac do not have vision so my images ended up getting tossed for SVG's which I almost never like. Using this quant: 25.3
You have reached the end of the archive
All of qwen38