I'm using qwen 3.8 q8. I've used two 3090s and a context of 260k. The speed feels pretty good.
What a crazy week in AI!
🚀 Ox Alpha GLM 5.3 Qwen 3.8 Flash Next Tencent Hy4 Minimax FastH3 World Humanoid Games Gemini 3.5 Transcribe Gemini Omni 1.1 Flash Block 3D One Video One World FixAnything Google PPE Code World Model VoiceMem Orbit++ DiffusionOPSD &
The model underneath it is Qwen 3.8 Max.
Alibaba says it has: 2.4T total parameters ~95B active parameters per task 1M-token context The idea is simple: Huge model capacity without activating the entire model for every job.
Qwen 3.8 flash next is built different.
125B parameters. Only 6B active per token. And Alibaba says it trained at roughly 1/10th the cost of its previous flagship. The architecture: → 125B main parameters with just 6B activated per token → Another 51B parameters via Engram
You have reached the end of the archive
All of qwen38