Qwen 3.8 Max just beat GPT and Gemini at using a real computer.
The test: OSWorld. Can an AI click buttons, open apps, and finish tasks like a person? The scores: → Qwen 3.8 Max: 86.1 → Claude Fable 5: 85.0 → GPT 5.6: 83.2 → Gemini 3.1 Pro: 76.2 An OPEN model is now
Qwen 3.6-27B just embarrassed a model 14X bigger.
Alibaba’s 27B model is beating its own 397B flagship on coding. But the benchmark score isn’t even the most useful part. Why this matters: → 27B dense model — all parameters fire on every token → 77.2% on SWE-bench Verified
Gemini Flash 3.7's UI design is sooo good considering its speed and pricing
It built this whole 3D skate game for me with just $3.4 and the speed is super fast compared to other models in the same tier like Kimi 3, Qwen 3.8 Max, or Grok 4.6
Advanced Frontier LLM Coding Benchmark 13.08.2026
Modeller: - Grok 4.6 Extra High - Cursor - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - Kimi K3 Max - Kimi-Code (App) - GLM 5.2 Max - Zcode (App) - GPT 5.6 Sol Very High - Codex (App) -
You have reached the end of the archive
All of qwen38