Try out cache-to-cache here
ONE DGX Spark Working with MiaAI_lab recipe of Qwen 3.8 flash next, but you can use any model Simple test of an essay about the research paper:
No single-stream headline today, so here's the part builders actually argue about
Two boxes running the same Gemma 4 12B IT QAT at 4-bit, and the spread between them is 3%. 29 tok/s on Strix Halo (Q4_K_XL). 28 tok/s on DGX Spark (Q4_K_M). Same model, different machines, and the
Gained a lot of respect for Zuck lately.
Muse and Qwen 3.8 27B are now daily drivers. Seriously impressed with their latest AI developments.
Qwen 3.8 27B doing an agentic task with subagent at ~120 tok/s on M5 Max MacBook Pro in lmstudio Bionic using inco_ai Splash engine
At launch a little more than a month ago the model was running around ~20tk/s It’s 6x faster now, local AI is progressing at lightning speed!
You have reached the end of the archive
All of qwen38