Alibaba just released Qwen 3.8 Max, an open model that beats GPT and Gemini at using a computer.
The test is called OSWorld. It checks if an AI can click buttons and finish tasks on a real desktop like a person. Qwen 3.8 Max scored 86.1. GPT scored 83.2. Gemini scored 76.2.
I asked Hermes Studio to create a desert island using threejs 🏝️ For me who can't travel abroad during summer vacation lol
Go's Qwen-3.8-MAX used up the 4-hour limit twice in no time and finished halfway, so I had Codex's GPT 5.6 Luna help me with the final adjustments.
GLM 5.3 Max vs Qwen 3.8 Max vs Grok 4.6 vs Opus 5
GLM 5.3 Max is actually holding up really well here. My ranking for this run: Opus 5 > Qwen 3.8 Max > GLM 5.3 Max > Grok 4.6 Definitely competitive now. GLM is getting scary close.
All models are so creative now
I gave Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, and Claude Opus 5 an ace of clubs card. And told them to be creative and draw inside the card, using the card as inspo and here are their results
You have reached the end of the archive
All of qwen38