Awesome! Qwen-3.8-27B NVFP4 quantization, enable MTP
The throughput speed on DGX Spark has reached 17tok/s, nearly doubled again If coupled with SGLang’s MTP, wouldn’t it be possible to take off✈️😄
You are paying Opus 4.8 prices for steps a 27B model finishes just as well.
NVIDIA shipped an open source router this week, NeMo SwitchYard. It reads the task and picks the model. Their own numbers: SwitchYard across Opus 4.8 plus smaller models completed more tasks than Opus
Qwen 3.8 27B Q5 XYZ, a paradise survival island, incredible graphics, everything generated by the model.
It’s impressive that it has only 27B parameters; I think this is the best graphical demonstration of this model I’ve shown so far.
Qwen 3.8 27B on a 3090. This is the stack while it answers.
You have reached the end of the archive
All of qwen38