Ok. Qwen 3.8 Next Q3, 7900xtx 64gbddr5.
Linux outpaces Windows by .5 toks on best configs running on my Chonk Buffer llama. 19.8 t/s Linux vs 19.34 best recipe. Full offload linux 15.4 vs 15.5. Was kind of surprised. Expected more but my zero overhead protocols 🦾 14 toks at 262k
Every inference engine kind of runs on vibes.
I got Qwen 3.8 already running on this modest tab BTW, however I'm trying to pool as many ops/fixes to the OpenCL backend as possible. has the fixes in opencl-fixes and qwen38-fixes cc: ggerganov
Qwen 3.8 flash next built this landing page in 29 minutes and served it on my tailnet on its own.
I am running official fp8 on 2x dgx spark, full 256k context loaded, 45 tok/s with mtp on. i have run deepseek and glm on these boxes and it was good, this one just feels right.
Inspired by these comparisons, I decided to do the same but using different Harness and my local Qwen 3.8 27B.
Very surprised at what the model is capable of with so little. And it is very clear that the harness matters a lot. I think droid, opencode, omp and hermes gave the
You have reached the end of the archive
All of qwen38