Deployed qwen 3.8 27b with dflash2 on 2 rtx 5090s
Steady generation: ~220 tok/s Normal range: ~190–257 tok/s Best observed: ~343 tok/s sgl_project with dflash2 is blazing fast
Qwen 3.8 27B can now run locally on RTX 4060 8GB?!
This is what made me stop scrolling. The 27B model which usually sounds "too big for a small GPU" can actually be forced to run on an RTX 4060 with only 8GB VRAM using a new quantization from Unsloth: IQ4_XS.
Gemini 3.7 Flash vs Qwen 3.8 vs DeepSeek V4 Pro vs GLM 5.3
Same prompt: busy train interior to leaving the station Qwen 3.8 & GLM 5.3 look the best visually
Wait… you can run Qwen 3.8 27B locally with basically a single gaming GPU.
You need: • 24 GB VRAM • ~17 GB disk • RTX 3090 / 4090 / 5090 no cloud. no API bill. no server farm. your GPU is probably more capable than you think.
You have reached the end of the archive
All of qwen38