Nice! According to actual testing, a large model with a scale of 100 billion was successfully deployed on DGX Spark!
! 100 billion per machine, soaring speed! (See screen recording👇🏻) However, SGLang (sgl_project) is said to be faster! I have to say that the domestic open source ecosystem is really getting stronger and stronger now! It is strongly recommended that everyone learns about local model deployment: install Ling-3.0, Qwen 3.8 27B, Minimax…
Make a nice gaussian splat 3d demo.
Use shaders. three.js, multiple files. qwen 3.8 27B BF16 on 4x 3090s one shot:
Beautiful visual of somebody running, qwen 3.8 27B locally on a RTX 5090 32 GB VRAM system with 115 tokens/sec
Note, Qwen3.8-27B's official BF16 checkpoint is 55.6 GB
Qwen 3.8 27b
DeepSeek V4 Flash 0731 GPT 5.6 Luna Most inexpensive models fail these tests.
You have reached the end of the archive
All of qwen38