Qwen 3.8 just beat Grok 4.6, Gemini, and GLM 5.3 on this one.
Same prompt. Four models. Gemini 3.7 Flash vs Qwen 3.8 vs Grok 4.6 vs GLM 5.3 asked each one to build two scenes in Three.js: a harvester and farmers working. Qwen 3.8 was the surprise for me. It understood the
Meme → ChatGPT → Prompt → Krea 2 → Image → Qwen 3.8 → Image + Prompt → MiniMax H3 → 🎬
Somewhere along the way, it became magic. ✨
Hey gang, here is my full evaluation/comparison of Qwen 3.8 27B against everyone's favorite, Qwen 3.6 27B.
I also found the first case of NVFP4 precision loss against BF16 that I could actually demonstrate. There was really no way I can do the full comparison justice in single
I tested four frontier models with the exact same prompt
Gemini 3.7 Flash, Qwen 3.8, Grok 4.6, and GLM 5.3. The task included two scenes: a working harvester and farmers in the field. Honestly, Qwen 3.8 impressed me the most—it simply understood the assignment and
You have reached the end of the archive
All of qwen38