Every inference engine kind of runs on vibes.
I got Qwen 3.8 already running on this modest tab BTW, however I'm trying to pool as many ops/fixes to the OpenCL backend as possible. has the fixes in opencl-fixes and qwen38-fixes cc: ggerganov
Qwen 3.8 flash next built this landing page in 29 minutes and served it on my tailnet on its own.
I am running official fp8 on 2x dgx spark, full 256k context loaded, 45 tok/s with mtp on. i have run deepseek and glm on these boxes and it was good, this one just feels right.
Inspired by these comparisons, I decided to do the same but using different Harness and my local Qwen 3.8 27B.
Very surprised at what the model is capable of with so little. And it is very clear that the harness matters a lot. I think droid, opencode, omp and hermes gave the
This was one shorted with qwen 3.8 max Alibaba_Qwen
You have reached the end of the archive
All of qwen38