Qwen 3.8 Flash is now live in closed beta.
Join the waitlist:
The VRAM barrier is officially dead.
I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer
Qwen 3.8 went to zero on two platforms today and one of them takes the word free as the API key
K2S is the only one in this whole free-model wave who put the base URL and the key next to the model names instead of a screenshot of a pricing page. The strongest thing in that post
SITUATION EXPLAINED: Alibaba’s Qwen 3.8 Flash beats DeepSeek V4 Flash and Opus 4.6 Max on most tasks, at just $0.16/M input tokens.
It leads DeepSeek V4 Flash and Opus 4.6 Max on most tested tasks, at $0.16 input and $0.47 output per million tokens • 125B total parameters
You have reached the end of the archive
All of qwen38