Alibaba just built an AI agent that beat GPT-5.6, gemini and claude at using a computer.
But the benchmark isn't even the wildest part. Qwen 3.8 Max: → 2.4 trillion total parameters → 95B active parameters per task → 1 MILLION token context window Computer-use results:
GLM 5.3 Max vs Qwen 3.8 Max vs Grok 4.6 vs Opus 5
GLM 5.3 Max is actually holding up really well here. My ranking for this run: Opus 5 > Qwen 3.8 Max > GLM 5.3 Max > Grok 4.6 Definitely competitive now. GLM is getting scary close.
All models are so creative now
I gave Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, and Claude Opus 5 an ace of clubs card. And told them to be creative and draw inside the card, using the card as inspo and here are their results
DeepSeek V4 Pro (0813) vs Qwen 3.8 Max
You have reached the end of the archive
All of qwen38