A 125B MoE model Just hit 25 tokens/sec on a single RTX 4090 at home.
Qwen 3.8 Flash Next + MTP speculative decoding at - 80k context. - 25.35 t/s decode - 471 t/s prefill Running a 125B Mixture-of-Experts model on a single consumer GPU
Took 15 minutes to make this with Qwen 3.8 27B locally running on my RTX Laptop and serving to my Macbook .
Welcome to the Real World , MATRIX!!!
Gemini 3.8 Flash made a Minecraft clone.
It’s slop. I ran the same test I run on every model. One shot, one prompt: create a Minecraft clone in a single HTML file. It actually looks pretty good at first. The textures are solid and it added a bunch of stuff I didn’t ask for,
Gemini 3.7 Flash High vs Claude Opus 5 High vs Grok 4.6 High vs Qwen 3.8 27B
You have reached the end of the archive
All of qwen38