I plugged Qwen-3.8-Max-0902 API directly into my coding workflow.
Don't worry... There's No fancy setup. Cuz You can use it with any AI Coding Agent This is my setup 👇 VS Code + Cline + Qwen. Because I have been using VS Code from the start 😅 I gave it real coding tasks…
Ok. Qwen 3.8 Next Q3, 7900xtx 64gbddr5.
Linux outpaces Windows by .5 toks on best configs running on my Chonk Buffer llama. 19.8 t/s Linux vs 19.34 best recipe. Full offload linux 15.4 vs 15.5. Was kind of surprised. Expected more but my zero overhead protocols 🦾 14 toks at 262k
Every inference engine kind of runs on vibes.
I got Qwen 3.8 already running on this modest tab BTW, however I'm trying to pool as many ops/fixes to the OpenCL backend as possible. has the fixes in opencl-fixes and qwen38-fixes cc: ggerganov
Qwen 3.8 flash next built this landing page in 29 minutes and served it on my tailnet on its own.
I am running official fp8 on 2x dgx spark, full 256k context loaded, 45 tok/s with mtp on. i have run deepseek and glm on these boxes and it was good, this one just feels right.
You have reached the end of the archive
All of qwen38