FRAMEWIREIndonesiaUpdated Aug 14Live wire
0:00 / 0:00

Alibaba just built an AI agent that beat GPT-5.6, gemini and claude at using a computer.

But the benchmark isn't even the wildest part. Qwen 3.8 Max: → 2.4 trillion total parameters → 95B active parameters per task → 1 MILLION token context window Computer-use results:

Julian Goldie SEOAug 1013
0:00 / 0:00

GLM 5.3 Max vs Qwen 3.8 Max vs Grok 4.6 vs Opus 5

GLM 5.3 Max is actually holding up really well here. My ranking for this run: Opus 5 > Qwen 3.8 Max > GLM 5.3 Max > Grok 4.6 Definitely competitive now. GLM is getting scary close.

OmedTheVibeCoderAug 141
0:00 / 0:00

All models are so creative now

I gave Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, and Claude Opus 5 an ace of clubs card. And told them to be creative and draw inside the card, using the card as inspo and here are their results

Ann NguyenAug 148
0:00 / 0:00

DeepSeek V4 Pro (0813) vs Qwen 3.8 Max

FoodTruck BenchAug 141