Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2!
(MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and
Open source Dsh plug-in, Liang Wenfeng vs. Liang Wengu.
DeepSeek V4 FULL 1 Hour 50 min Course
You have reached the end of the archive
All of deepseek