Grok 4.6 seems like a big improvement
I gave it, along with Claude Opus 5, GPT 5.6 Sol, and Qwen 3.8 Max, a banana.
I gave it, along with Claude Opus 5, GPT 5.6 Sol, and Qwen 3.8 Max, a banana.
The biggest upgrade isn’t another model. It’s what happens when every agent finally shares the same memory.
Qwen 3.8 2.4T—a full-on game-building battle royale. DeepSeek was the only model to get the game right on its first attempt.
61 on the Index, two points behind Claude Opus 5, at a quarter of the output price.
This one is actually pretty close. Grok 4.6 and DeepSeek V4 Pro are very similar in raw quality here.
Qwen 3.8 2.4T A95B Seed 2.0 Code Seed 2.1 Turbo
Dive into the comparison to see which AI model delivers more value for your coding needs.
The difference is actually pretty obvious. Sadly, Grok 4.6 still doesn’t really reach the top models in my tests.
Python offers rapid development, but C++ wins for trading latency. The choice depends on your priorities.
Grok 4.6 Deepseek v4 Pro GA Qwen 3.8 2.4T aka 3.8 Max Qwen 3.8 27B What a great day for LocalAI & humans!
Python offers efficiency, and AI like Qwen 3.8 Max is revolutionizing development insights. Rethink your tech stack.
I tested it across games, coding-style builds and agent tasks — and the surprising part wasn’t where it failed.
Some speculate it could be an early Qwen 4.0 candidate, but Alibaba has not confirmed that.
The test: OSWorld. Can an AI click buttons, open apps, and finish tasks like a person?
I've been loving this 1-bit dithered style lately. Qwen 3.8 max is going to be a nice win for everyone tomorrow I'm excited to see what size people…