No single-stream headline today, so here's the part builders actually argue about
Two boxes running the same Gemma 4 12B IT QAT at 4-bit, and the spread between them is 3%. 29 tok/s on Strix Halo (Q4_K_XL). 28 tok/s on DGX Spark (Q4_K_M). Same model, different machines, and the
Gained a lot of respect for Zuck lately.
Muse and Qwen 3.8 27B are now daily drivers. Seriously impressed with their latest AI developments.
Qwen 3.8 27B doing an agentic task with subagent at ~120 tok/s on M5 Max MacBook Pro in lmstudio Bionic using inco_ai Splash engine
At launch a little more than a month ago the model was running around ~20tk/s It’s 6x faster now, local AI is progressing at lightning speed!
Like any distilled model, Ternary Bonsai 2 27B inherits refusals and censorship from its teacher model: Qwen 3.8 27B.
Fortunately, there’s a very simple and elegant way to reduce or perhaps even completely eliminate them. 🧵
You have reached the end of the archive
All of qwen38