I put MCDMA through its paces today and tested Qwen 3.8 Flash Next on my dual Spark / Mac Studio cluster.
My first test was disaggregated prefill across my two Sparks, then decode onto my Studio. PP 2,100 tok/s Decode 80 tok/s at 20k context, higher on short replies. That's
This is an actual operation of Qwen-3.8-Flash-Next running on DGX Spark with MTP support + llama.cpp on-direct configuration.
What I want you to see in the video is not the generation speed itself, but the speed from when a task is thrown until it starts moving. Once you send it, Reasoning goes into Reasoning without any feeling of waiting, and the text starts flowing straight away.
Testing qwen 3.8 27B on cerebras doing a wiki speedrun w/ ego lite
High reasoning: 4 clicks, 20.9s no reasoning: 10 clicks, 15.8s + missed the win once an idiot in motion will go further than a genius at rest
MiniMax H3, DaVinci Resolve, and Qwen 3.8 27B for a stylized “Words” animation.
The interesting part isn’t just the look—it’s how fast a small tool stack can turn an absurd idea into a finished sequence.
You have reached the end of the archive
All of qwen38