DFlash 2 just dropped, and local inference is now on another level.
According to AA, Qwen3.8 27B sits roughly between GLM-5.2 Max and DeepSeek Flash 0731 Max. Q4_K_M running on a single 32GB RTX 5090: 75 tok/s sustained, up to 200 tok/s peak, 131K context, one active session.
Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2!
(MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and
Open source Dsh plug-in, Liang Wenfeng vs. Liang Wengu.
You have reached the end of the archive
All of deepseek