Grok 4.6 vs Grok 4.5 vs Kimi K3 vs Qwen 3.8 Max Preview.
The difference is actually pretty obvious. Sadly, Grok 4.6 still doesn’t really reach the top models in my tests. I’d put it more in the mid-tier range right now.
Qwen 3.6-27B just embarrassed a model 14X bigger.
Alibaba’s 27B model is beating its own 397B flagship on coding. But the benchmark score isn’t even the most useful part. Why this matters: → 27B dense model — all parameters fire on every token → 77.2% on SWE-bench Verified
Gemini Flash 3.7's UI design is sooo good considering its speed and pricing
It built this whole 3D skate game for me with just $3.4 and the speed is super fast compared to other models in the same tier like Kimi 3, Qwen 3.8 Max, or Grok 4.6
Advanced Frontier LLM Coding Benchmark 13.08.2026
Modeller: - Grok 4.6 Extra High - Cursor - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - Kimi K3 Max - Kimi-Code (App) - GLM 5.2 Max - Zcode (App) - GPT 5.6 Sol Very High - Codex (App) -
You have reached the end of the archive
All of qwen38