GLM 5.3 Flash (320B) vs. Grok 4.6 (1.5T) vs. GPT 5.6 Sol vs. Qwen 3.8 Flash (125B)
GLM genuinely outperformed massive multi-trillion flagship architectures in side-by-side tests. The efficiency gap is wild when an 18B active model beats 1.5T monsters.
Benchmarked Hy4 Preview vs.
GLM 5.3 Flash vs. GLM 5.3 and Qwen 3.8 Flash. Expected Hy4 Preview on WorkBuddy to be impressive, but it completely flopped compared to the rest. GLM 5.3 Flash and Qwen 3.8 Flash run circles around it without breaking a sweat.
Qwen 3.8 Flash vs Tencent Hy4 Preview
Qwen 3.8 27B works very stably under 900K, occupying 120G of video memory throughout the process
In order to clearly show you the speed of this model optimization (125 TPS), I asked it to memorize the Three-Character Sutra. Can you see how fast it memorizes it😂
You have reached the end of the archive
All of qwen38