This was one shorted with qwen 3.8 max Alibaba_Qwen
Here's Qwen 3.8 27b in action using the DeepSeek harness.
This is self-hosted from our facility. 55-65 tokens per second. Text go woosh. This is FP16 precision 262k context vision enabled. We will be adding more open weight models soon, next up will be Qwen 3.8 Flash Next
I put MCDMA through its paces today and tested Qwen 3.8 Flash Next on my dual Spark / Mac Studio cluster.
My first test was disaggregated prefill across my two Sparks, then decode onto my Studio. PP 2,100 tok/s Decode 80 tok/s at 20k context, higher on short replies. That's
This is an actual operation of Qwen-3.8-Flash-Next running on DGX Spark with MTP support + llama.cpp on-direct configuration.
What I want you to see in the video is not the generation speed itself, but the speed from when a task is thrown until it starts moving. Once you send it, Reasoning goes into Reasoning without any feeling of waiting, and the text starts flowing straight away.
You have reached the end of the archive
All of qwen38