Inspired by these comparisons, I decided to do the same but using different Harness and my local Qwen 3.8 27B.
Very surprised at what the model is capable of with so little. And it is very clear that the harness matters a lot. I think droid, opencode, omp and hermes gave the
This was one shorted with qwen 3.8 max Alibaba_Qwen
Here's Qwen 3.8 27b in action using the DeepSeek harness.
This is self-hosted from our facility. 55-65 tokens per second. Text go woosh. This is FP16 precision 262k context vision enabled. We will be adding more open weight models soon, next up will be Qwen 3.8 Flash Next
I put MCDMA through its paces today and tested Qwen 3.8 Flash Next on my dual Spark / Mac Studio cluster.
My first test was disaggregated prefill across my two Sparks, then decode onto my Studio. PP 2,100 tok/s Decode 80 tok/s at 20k context, higher on short replies. That's
You have reached the end of the archive
All of qwen38