Qwen 3.8 Next Flash 4bit MTP on my M3U Studio is absolutely flying now.
My latest OMP run measured 97.1 tok/s overall decode at 23k context, with ~1,131 tok/s uncached PP. PR opened on oMLX for this.
Holy smokes qwen-3.8-27b on cerebras with Hermes Agent feels so good and fast it's almost unreal.
Testing with the upcoming Cadu app by fl_rn_st. Tool calls are now the main slowdown 😅
DeepSeek V4.1 Flash is now listed with free access on Apinex—and the catalog includes several other models
DeepSeek V4 Pro • DeepSeek V4 Flash • Gemini 3.8 Flash • GLM 5.3 Flash • Muse Spark 1.3 • GPT-5.6 Luna • Qwen 3.8 Max Browse the models:
Operating System powered by Qwen 3.8 27B at 1950 tokens/sec!
Here is what 1,950 tokens/second Alibaba_Qwen's 3.8 27b actually looks like on cerebras: i wrote a minimal python web server that turns cerebras inference into a live operating system. zero apps on disk. when you
You have reached the end of the archive
All of qwen38