Doing some live testing not finished but current state .
Do you think its useable ? 2 b70s Qwen 3.8 27b FP16 Deepseek harness
Qwen 3.8 27B Q4 is now running on an RTX 4060 with just 8GB VRAM and a 64,000 token context window using Unsloth's new IQ4_XS quant at 14.6GB on disk.
→ Prefill at 150 tokens per second, decode at 5 tokens per second via native MTP → Only 25 GPU layers offloaded to stay within
GLM 5.3 vs Gemini 3.7 Flash vs Qwen 3.8 vs Grok 4.6
Hermes Agent got Qwen 3.8 installed (with Grok)
Attached it to a Hermes Bot, and the first job I gave it was to brainstorm this animation with me for it's persona. I can't even believe this is running on my own computer. SO fast on my 5090. thanks NousResearch Teknium
You have reached the end of the archive
All of qwen38