I used grok with Grok Build to continue developing my hobby game in threejs almost no errors
Very fast but being fast also makes me prompt more meaning spend more as well, it is both good and bad in a way :) Almost no errors as well, flawless…
Modeller: - Grok 4.6 Extra High - Cursor - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - Kimi K3 Max - Kimi-Code (App) - GLM 5.2 Max - Zcode (App) - GPT 5.6 Sol Very High - Codex (App) -
Very fast but being fast also makes me prompt more meaning spend more as well, it is both good and bad in a way :) Almost no errors as well, flawless…
Told it to write a song. it did. minimax-music3 made the track. first output.
Grizzly bear riding a scooter in a coastal environment. Qwen 3.8 Max/Gemini 3.5 Flash Lite.
Here's the first song it wrote ☆*: .。. o(≧▽≦)o .。.:*☆
Sin tarjeta. Sin trial de 7 días. Gratis. ¡Y sin marca de agua! He visto la prueba completa de Qwen 3.8 Max y hay 3 cosas que no esperaba. Hilo 🧵
The thing that got my attention is the context window, hundreds of thousands of tokens in one pass. Basically a whole codebase.
By category, Gemini 3.7 Flash (High) also strongly improved: - #3 Creative Writing (#12 -> #3) - #4 Math (#7 -> #4) - #9 Instruction Following (
Not super complicated at all, all other models couldn't one-shot and needed more guidance to get to result.
I’m running Meta’s new 30B Muse Glimmer Q6_K_XL with a massive 130k context window on just 26GB VRAM FREE compute on Kaggle.
And a chat tab is the worst place to use it. The model is a monster: → Near the top of the largest models ever publicly released → Reads text,…
So we gave it the same traditional Chinese paper-cut style animation prompt alongside GPT-5.6 Sol, Claude Opus 5, and Qwen 3.8 Max.
Grok 4.6 - gpt 5.6 sol - qwen 3.8-max - Claude opus 5 who won ?
There's now a free AI model that beats one 14 times its size. Alibaba's own 397 billion parameter flagship just lost to their new 27 billion one.
I gave it, along with Claude Opus 5, GPT 5.6 Sol, and Qwen 3.8 Max, a banana.
The biggest upgrade isn’t another model. It’s what happens when every agent finally shares the same memory.
Qwen 3.8 2.4T—a full-on game-building battle royale. DeepSeek was the only model to get the game right on its first attempt.