Bro… someone just built in one shot a Mario clone locally with Qwen 3.8 27b
Link:
Gemini 3.8 Flash High effort Vs Qwen 3.8 27B Max effort
A 125B MoE model Just hit 25 tokens/sec on a single RTX 4090 at home.
Qwen 3.8 Flash Next + MTP speculative decoding at - 80k context. - 25.35 t/s decode - 471 t/s prefill Running a 125B Mixture-of-Experts model on a single consumer GPU
Took 15 minutes to make this with Qwen 3.8 27B locally running on my RTX Laptop and serving to my Macbook .
Welcome to the Real World , MATRIX!!!
You have reached the end of the archive
All of qwen38