Qwen 3.8 Next Flash... this is actually cool af.
Prompt below
Alibaba's Qwen team just dropped Qwen 3.8 Max, a multimodal model with open weights.
It boasts 2.4 trillion parameters, a 1M token context window, and can create apps from screenshots. Best part? It's 80% cheaper than GPT-4.5 and 88% cheaper than Claude 3. | Peter Diamandis
Qwen 3.8 Flash is now live in closed beta.
Join the waitlist:
The VRAM barrier is officially dead.
I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer
You have reached the end of the archive
All of qwen38