I unlocked the full MTP potential in this new update of cafe-llama.cpp, it doubles decoding speed.
Qwen 3.8 Flash Next is over 50t/s, it was 25t/s Qwen 3.6 35B Q8 100t/s was 40t/s Qwen 3.8 27B 80t/s was 40t/s recommended --spec-draft-n-max 4
Local model test on my DGX Spark: Qwen 3.8 Flash vs. a V8 engine.
One HTML file. Code-only geometry. 8 pistons in firing order, exploded view, X-ray mode, touch orbit on iPhone. No cloud. Full prompt in the first comment.
Wtf Turns out Qwen 3.8 27B can make and edit videos like Opus 5.5
Literally, why did no one think to try this before?
Another GenAI SNES banger feat.
RealAstropulse! I used Claude Code to learn retro game dev. How Qwen 3.8 is taking over.
You have reached the end of the archive
All of qwen38