Gemini 3.7 Flash vs Qwen 3.8 vs Grok 4.6 vs GLM 5.3
Qwen 3.8 27B vs. DeepSeek V4 Flash 0731
Even though DeepSeek is 10x the size, Qwen actually produced better results
Awesome! Qwen-3.8-27B NVFP4 quantization, enable MTP
The throughput speed on DGX Spark has reached 17tok/s, nearly doubled again If coupled with SGLang’s MTP, wouldn’t it be possible to take off✈️😄
You are paying Opus 4.8 prices for steps a 27B model finishes just as well.
NVIDIA shipped an open source router this week, NeMo SwitchYard. It reads the task and picks the model. Their own numbers: SwitchYard across Opus 4.8 plus smaller models completed more tasks than Opus
You have reached the end of the archive
All of qwen38