Qwen 3.8 Max: An open model just beat GPT and Claude at using a computer.
OSWorld tests if an AI can click buttons and finish real desktop tasks like a person.
OSWorld tests if an AI can click buttons and finish real desktop tasks like a person.
The few-shot image prompting is amazing: you can annotate an image with one bounding box, and it will find the remaining objects 🤯 I tried it out on…
And you can run it yourself. It's one of the biggest models ever released. Open. Not locked behind one company.
Muse puts up a good show but it's not really close here. However like always Qwen uses way too many tokens in general, I wonder how 3.8 is going to…
This was made with MiniMax_AI M3 model i was testing the harness and thought i should try the boss fight benchmark and I'm really impressed, this is…
With <200GB of vram, i'm getting ~800K context @ 200 tokens per second. It's not quite Cerebras fast, but damn.
Your always -on workmate for: - Coding - Reasoning - Researching - Writing - AI Agent With Qwen 3.8-Max: Just live your life to the fullest.
From memory, this result is probably a little better than stock Qwen 3.5 out of the box. It's not perfect, but it's pleasant!
🔥 Without D flash it was pushing 40-60 tokens/sec. With it? We're hitting 160-180 on the peaks before it settles.
It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config.
Nome estranho, benchmark forte. Alguém já testou isso?
2.4T total parameters. 95B active during inference. 1M token context. But parameter count isn't the important number.
The test: click buttons, open apps, finish tasks on a real desktop. → Qwen 3.8 Max scored 86.1 → GPT 5.6 scored 83.2. Gemini scored 76.2.
Aimlapi ran the test: same prompt, same one-shot job, no tricks Cost per figure: Grok 4.5 → $0.15 GPT 5.6 Sol → $0.60 Qwen 3.8 Max → $0.67 Kimi K3 →…
It just landed, and most people have no idea it exists. → Reads text, images, and video → Holds 1M tokens of context → Works on projects for days…
It's shared memory. Every Hermes conversation gets written into an Obsidian vault.
China ran out of GPUs. Microsoft is losing relevance. The American AI monopoly just ended.