FRAMEWIREIndonesiaUpdated Sep 16Live wire
0:00 / 0:00

How much faster can Qwen 3.8 Flash Next run with its target weights pinned?

We opened cuda(.)fast and mlx(.)fast on yukonresearch: CUDA on DGX Spark, MLX on Apple Silicon Optimize the engine, pass correctness, move the verified speed frontier. Here’s how to get started

ZeeshanSep 169
0:00 / 0:00

I plugged Qwen-3.8-Max-0902 API directly into my coding workflow.

Don't worry... There's No fancy setup. Cuz You can use it with any AI Coding Agent This is my setup 👇 VS Code + Cline + Qwen. Because I have been using VS Code from the start 😅 I gave it real coding tasks…

Faith MudashiSep 16
0:00 / 0:00

Ok. Qwen 3.8 Next Q3, 7900xtx 64gbddr5.

Linux outpaces Windows by .5 toks on best configs running on my Chonk Buffer llama. 19.8 t/s Linux vs 19.34 best recipe. Full offload linux 15.4 vs 15.5. Was kind of surprised. Expected more but my zero overhead protocols 🦾 14 toks at 262k

ChonkESep 166
0:00 / 0:00

Every inference engine kind of runs on vibes.

I got Qwen 3.8 already running on this modest tab BTW, however I'm trying to pool as many ops/fixes to the OpenCL backend as possible. has the fixes in opencl-fixes and qwen38-fixes cc: ggerganov

Karthik Kumar ViswanathanSep 163