Qwen 3.8 27B Q4 is now running on an RTX 4060 with just 8GB VRAM and a 64,000 token context window using Unsloth's new IQ4_XS quant at 14.6GB on disk.
→ Prefill at 150 tokens per second, decode at 5 tokens per second via native MTP → Only 25 GPU layers offloaded to stay within
GLM 5.3 vs Gemini 3.7 Flash vs Qwen 3.8 vs Grok 4.6
Hermes Agent got Qwen 3.8 installed (with Grok)
Attached it to a Hermes Bot, and the first job I gave it was to brainstorm this animation with me for it's persona. I can't even believe this is running on my own computer. SO fast on my 5090. thanks NousResearch Teknium
Qwen 3.8 27B running on a 3090 at Q5 just beat gemini 3.7 flash on a water surface simulation.
Same prompt, one shot, not full precision. One of the best water simulations seen across every local model tested.
You have reached the end of the archive
All of qwen38