50% more context unlocked for Qwen 3.8 27b Q4_K_XL dflash 2 on a single RTX 4090 (24 GB VRAM)
I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B
Deployed qwen 3.8 27b with dflash2 on 2 rtx 5090s
Steady generation: ~220 tok/s Normal range: ~190–257 tok/s Best observed: ~343 tok/s sgl_project with dflash2 is blazing fast
Qwen 3.8 27B can now run locally on RTX 4060 8GB?!
This is what made me stop scrolling. The 27B model which usually sounds "too big for a small GPU" can actually be forced to run on an RTX 4060 with only 8GB VRAM using a new quantization from Unsloth: IQ4_XS.
Gemini 3.7 Flash vs Qwen 3.8 vs DeepSeek V4 Pro vs GLM 5.3
Same prompt: busy train interior to leaving the station Qwen 3.8 & GLM 5.3 look the best visually
You have reached the end of the archive
All of qwen38