I've released the benchmark for 7 different models
GPT 4.6, GPT 5.6 Sol, GLM 5.2 & 5.3, Laguna S2.1, Qwen 3.8 27B and DeepSeek Flash V4 0731 ! some local, some cloud with one single prompt and no steering, using OMP (for the first time). Yotubue Link :
Qwen 3.8 27b firework show creation using my local build MARCUS.
Prompt: "Create a 4th of July fireworks show over NYC skyline. Make it detailed, polished, with eye popping effects. Do any web search needed to gather more information about the NYC skyline and the locations of
Qwen 3.8 27B just dropped, so we put it head-to-head against GLM 5.3, Kimi K3, and Opus 5 on the same prompt.
All four managed to ship something playable: movement, shooting, hit detection, enemy spawning, etc. But Qwen 3.8 27B stands out for a few reasons: • Dense 27B model
I had Qwen 3.8 27B make the game.
Completely offline development in a local LLM. The basic implementation was completed by leaving my MacBook alone for about 50 minutes.
You have reached the end of the archive
All of qwen38