I had some fun playing around with the Qwen 3.8 model yesterday.
What blows my mind is that it's nearly a frontier model yet it runs on my 5 year M1 Max laptop! Now it's not super fast, 16 tokens/s but still! Check out the little video I made about how to set it up:
I'm doing exactly this, my voice goes thru my home server running Qwen 3.8
So I’m working on making decoding and prefill as fast as possible on my veloGB10 engine for the 3.8 27b model.
My 4x DGX Spark networking is switched (Mikrotik CRS812 DDQ ), High quality QFSP56 cables with all paths verified to run at max speed. This is a pure Rust/CUDA (with
I'm building my son 'roblox' but entirely local AIs.
Doing Qwen 3.8 on a 5090 as its chat backbone.
You have reached the end of the archive
All of qwen38