The GLM 5.3 API just went live, and it got so good at finding weak spots in code that the makers paused the download.
Here's the strange part. It uses the exact same base model as GLM 5.2. Not one bit smarter under the hood. All they did was more post-training. More practice.
I ran Qwen 3.8 27B on a single 8GB GPU.
IQ4_XS + Q4 KV cache at 70k context. The token speed hovers around 4-5 t/s. Slow but it's still amazing that you could get a result that can outperform even Claude Opus 4.6 on some of my test. All in one-shot with zero agent looping.
GLM 5.3 is free right now, and it just scored the highest number ever recorded on an independent coding benchmark.
91.25% on Kingbench 3. It beat Fable 5. It beat Opus 5. It beat Kimi K3. New users get it on the free tier inside Zcode. No card. No subscription. 5 million
I made a new benchmark for AI models, the "ramen test".
I wanted a way to measure initial vibes as someone who isn't really technical The goal is to create a cozy experience of eating a bowl of ramen, first person POV. Then they have pretty much free reign In this short video
You have reached the end of the archive
All of Claude Opus 5