The VRAM barrier is officially dead.
I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer
Qwen 3.8 went to zero on two platforms today and one of them takes the word free as the API key
K2S is the only one in this whole free-model wave who put the base URL and the key next to the model names instead of a screenshot of a pricing page. The strongest thing in that post
SITUATION EXPLAINED: Alibaba’s Qwen 3.8 Flash beats DeepSeek V4 Flash and Opus 4.6 Max on most tasks, at just $0.16/M input tokens.
It leads DeepSeek V4 Flash and Opus 4.6 Max on most tested tasks, at $0.16 input and $0.47 output per million tokens • 125B total parameters
We still don't have any open-source model anywhere close to Kimi K3
I tested the exact same Mini Militia-style game with Qwen 3.8 Max after trying it with Kimi and 0x Alpha the results were surprisingly bad for Qwen the controls were actually good, but the UI was nowhere close
You have reached the end of the archive
All of qwen38