Qwen 3.8 27B is insanely powerful.
I've been running it locally on my PC and I have to say it's AT least as good as Opus 4.6 max and I can run it 100% free with zero limits, restrictions, and completely privately. Open source AI is catching up to SOTA models fast.
Kimi K3, Opus 5, Grok 4.6, Qwen 3.8 27B, GPT 5.6 SOL: Recreate Ghost of Tsushima
I put 5 AIs against each other to see which can build the prettiest experience of a samurai roaming an ancient island of Japan (all one shotted except for one model) Kimi K3 - ran for 6+ hours and
448 GB/s divided by 4.22 GB is 106 tok/s.
That is the hard decode ceiling on my 8GB 3070, and yesterday I measured 203. That is not cheating. I went and checked where the line actually is. LocalMaxxing auto-flags any submitted run faster than 10x the most generous decode
I tested 0xWhiteMage's recipe: Qwen3.8-27B Kearuga on a single DGX Spark.
One of the most interesting builds I've run this year. Why it's interesting Most quant work is compression engineering: shrink the model, keep it fast, accept the loss. Kearuga treats the same problem as
You have reached the end of the archive
All of qwen38