The model underneath it is Qwen 3.8 Max.
Alibaba says it has: 2.4T total parameters ~95B active parameters per task 1M-token context The idea is simple: Huge model capacity without activating the entire model for every job.
Qwen 3.8 flash next is built different.
125B parameters. Only 6B active per token. And Alibaba says it trained at roughly 1/10th the cost of its previous flagship. The architecture: → 125B main parameters with just 6B activated per token → Another 51B parameters via Engram
Here's a Qwen 3.8 27b NVFP4 w/ dflash2 result, medium thinking, 32k reasoning budget, 11 minute run.
Not as intricate obv. but decent
Qwen 3.8 Flash Next has some wild numbers.
Here are the ones to remember: → 125B total parameters. → Only 6B active per token. → 51B engram embeddings. → 262K context out of the box. → Up to 1M tokens with YAN. → 62.5 on SWE-Bench Pro. → 73.9 on agentic office
You have reached the end of the archive
All of qwen38