FRAMEWIREIndonesiaUpdated Sep 22Live wire
0:00 / 0:00

No single-stream headline today, so here's the part builders actually argue about

Two boxes running the same Gemma 4 12B IT QAT at 4-bit, and the spread between them is 3%. 29 tok/s on Strix Halo (Q4_K_XL). 28 tok/s on DGX Spark (Q4_K_M). Same model, different machines, and the

Zach ASep 22
0:00 / 0:00

Gained a lot of respect for Zuck lately.

Muse and Qwen 3.8 27B are now daily drivers. Seriously impressed with their latest AI developments.

StartupHakkSep 21
0:00 / 0:00

Qwen 3.8 27B doing an agentic task with subagent at ~120 tok/s on M5 Max MacBook Pro in lmstudio Bionic using inco_ai Splash engine

At launch a little more than a month ago the model was running around ~20tk/s It’s 6x faster now, local AI is progressing at lightning speed!

Adrien GrondinSep 21181
0:00 / 0:00

Like any distilled model, Ternary Bonsai 2 27B inherits refusals and censorship from its teacher model: Qwen 3.8 27B.

Fortunately, there’s a very simple and elegant way to reduce or perhaps even completely eliminate them. 🧵

Private LLMSep 2127