FRAMEWIREIndonesiaUpdated Aug 14Live wire
0:00 / 0:00

Qwen 3.8 Max: An open model just beat GPT and Claude at using a computer.

OSWorld tests if an AI can click buttons and finish real desktop tasks like a person. The scores: → Qwen 3.8 Max: 86.1 → Claude Fable 5: 85.0 → GPT 5.6 Soul Max: 83.2 → Gemini 3.1 Pro: 76.2 The

Julian Goldie SEOAug 111
0:00 / 0:00

DeepSeek V4 Pro (0813) vs Qwen 3.8 Max

FoodTruck BenchAug 141
0:00 / 0:00

GLM 5.3 Max vs Fable 5 vs Qwen 3.8 Max vs Grok 4.6

GLM 5.3 is a massive jump over 5.2. Still not beating Fable 5 for me, but it’s way closer now. And vs Grok 4.6? I’d take GLM 5.3 pretty easily on this test.

OmedTheVibeCoderAug 146
0:00 / 0:00

Advanced Frontier LLM Coding Benchmark 2 14.08.2026

Models: - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - GLM 5.3 Max - Zcode (App) - GPT 5.6 Sol Ultra - Codex (App) Task: Ferrofluid: Rising towards metaball surface + magnet cursor

Alican KirazAug 1418