Yeah, it literally feels like AMD R9700 AI Pros are FLYING!
🔥 Using Radiance Quant: Qwen 3.8 27B MXFP4 by Just1
It's because Voice mode doesn't use reasoning, see my local qwen 3.8 27b example
First answer with no reasoning, second answer with reasoning set to low
Qwen 3.8 Next Flash 4bit MTP on my M3U Studio is absolutely flying now.
My latest OMP run measured 97.1 tok/s overall decode at 23k context, with ~1,131 tok/s uncached PP. PR opened on oMLX for this.
Holy smokes qwen-3.8-27b on cerebras with Hermes Agent feels so good and fast it's almost unreal.
Testing with the upcoming Cadu app by fl_rn_st. Tool calls are now the main slowdown 😅
You have reached the end of the archive
All of qwen38