Alibaba dropped Qwen 3.8 last week, and it's now on DigitalOcean Serverless Inference.
The thing that got my attention is the context window, hundreds of thousands of tokens in one pass. Basically a whole codebase. So I threw an entire repo at it.
Japanese Tea Garden Bench - Qwen 3.8 Max
In the Text Arena, Gemini 3.7 Flash (High) landed #9 with 1490 pts, another improvement from Gemini 3.6 Flash (High) (#16 -> #9).
By category, Gemini 3.7 Flash (High) also strongly improved: - #3 Creative Writing (#12 -> #3) - #4 Math (#7 -> #4) - #9 Instruction Following (
Meta's Muse-Glimmer-30B is the first model running on my 3090 able to one shot this.
Not super complicated at all, all other models couldn't one-shot and needed more guidance to get to result. Not only this, but Muse is able to one-shot almost everything i have thrown
You have reached the end of the archive
All of qwen38