Uncensored version of Qwen3.8-27B on 5090
I wanted to run it on NVFP4 with 200K context. The prefill seems fast and it seems easy to use. There were no ready-made ones, so I baked it myself. ▫️ 2 types in NVFP4 version All 4 bits, The MLP is 4 bits and the attention and output layers are 8 bits.
Qwen 3.8 max just beat GPT-5.6 sol max at actually using a computer.
And the benchmark isn't even the most interesting part. The numbers: → 2.4 trillion parameter Mixture-of-Experts model → Only ~95B parameters activate per token → 1M-token context window across text,
Designing with Qwen 3.8 27B in I am in a bit of a shock.
It's a big deal. These one-shot designs are marginally different from Opus, there are still nuances to be fair, but it's dangerously close. Once cerebras runs this model north of 1000 tps 👉we
Qwen 3.8 27b rebuilt the water physics with a video analysis skill.
It's still going, too. This is another intermediate result. Unreal!
You have reached the end of the archive
All of qwen38