My company's harness for Qwen 3.8-27B, Muse Glimmer 30B and Gemma 4:31B.
Local Laptop LLM
Bro… someone just built in one shot a Mario clone locally with Qwen 3.8 27b
Link:
Gemini 3.8 Flash High effort Vs Qwen 3.8 27B Max effort
A 125B MoE model Just hit 25 tokens/sec on a single RTX 4090 at home.
Qwen 3.8 Flash Next + MTP speculative decoding at - 80k context. - 25.35 t/s decode - 471 t/s prefill Running a 125B Mixture-of-Experts model on a single consumer GPU
You have reached the end of the archive
All of qwen38