If you're curious what ~48 tok/s looks like.
Running MTPLX Qwen 3.8 4 bit quant -- fully locally on 128gb MBP M5 from
Nice! According to actual testing, a large model with a scale of 100 billion was successfully deployed on DGX Spark!
! 100 billion per machine, soaring speed! (See screen recording👇🏻) However, SGLang (sgl_project) is said to be faster! I have to say that the domestic open source ecosystem is really getting stronger and stronger now! It is strongly recommended that everyone learns about local model deployment: install Ling-3.0, Qwen 3.8 27B, Minimax…
Make a nice gaussian splat 3d demo.
Use shaders. three.js, multiple files. qwen 3.8 27B BF16 on 4x 3090s one shot:
Beautiful visual of somebody running, qwen 3.8 27B locally on a RTX 5090 32 GB VRAM system with 115 tokens/sec
Note, Qwen3.8-27B's official BF16 checkpoint is 55.6 GB
You have reached the end of the archive
All of qwen38