Qwen 3.8 now runs inside Agent OS, turning a 2.4 trillion parameter brain into a full worker.
A chat tab gives Qwen a mouth. Agent OS gives it a body: → Layer 1: The brain. Qwen, holding 1 million tokens of context → Layer 2: The hands.
A chat tab gives Qwen a mouth. Agent OS gives it a body: → Layer 1: The brain. Qwen, holding 1 million tokens of context → Layer 2: The hands.
Parameters are knobs in the AI's brain. This has 2.4 trillion. More than almost any AI on Earth.
The receipts are public on GitHub: 265 commits. 127 pull requests. 151 issues closed.
Not 10 minutes. Not 10 hours. Ten DAYS. It starts with an empty folder and ends with finished software.
The honest verdict: it lost almost every round. The tests: RPGs, racing games, flight sims, real builds you can click and play.
The scattered agent problem: You hired a brilliant assistant. Then locked them in a different building for every job. Chat in one building.
I mentioned it to add sound effects and voice its much better than kimi k3 and GPT 5.6 Sol here SpaceX Falcon 9 launch, stage separation, boostback,…
Same screen. Same prompts. The real winner isn't who you'd guess. The 3 jobs: a full landing page, a working calculator, and a Flappy Bird game in…
3D games. A full website. A working operating system. A complete promo video. All built in hours.
Same prompts. One shot each. Here's who actually wins. The test: Flappy Bird, Tetris, and a real landing page. One HTML file each. No do-overs.
Qwen 3.8 Max: 86.1 Claude Fable 5: 85.0 GPT-5.6 Sol Max: 83.2 Gemini 3.1 Pro: 76.2 That's clicking, opening apps and completing desktop tasks.
The test is called OSWorld. It checks if an AI can click buttons and finish tasks on a real desktop like a person. Qwen 3.8 Max scored 86.1.
This is the first test as usual a ballister test for comparison i have quote the post with output from opus 5 and sol 5.6 You can also play with it…
I will prefer using V4 over GPT-5.6, Claude Opus 5, Kimi K3 Max or Qwen 3.8 Max, and Claude Fable 5
DeepSeek 4 flash, Sonnet 5, kimi k3, Qwen 3.8. All made a similar game all work pretty well.
Qwen 3.8 Max, Wan 3.0 Video, Qwen Image 3.0 and Pro, and Doubao Seedance 2.5.
สรุป AI สัปดาห์นี้ EP.14 1. Gemini Spark เปิดให้ไทยใช้แล้ว (AI Pro/Ultra ขึ้นไป) สั่งงานยาก ๆ แล้วรันบนคลาวด์ต่อเนื่องได้ 2.
On cost per Intelligence Index task it's $1.13 against Kimi's $0.84 - it burns more tokens to finish the same job.
It’s interesting because it can actually DO things on a computer. A practical Agent OS setup looks like this: → Web search finds current info → Web…