GLM 5.3 just dropped - and it's sitting at the top of the benchmarks.
But benchmarks are a lab result, not a reality check. So we ran it against every frontier model, head to head: 1. GPT-5.6 Sol vs GLM 5.3 2. Opus 5 vs GLM 5.3 3. Gemini 3.7 Flash vs GLM 5.3 4. Grok 4.6 vs GLM
Every model we've tested answers the surface question.
Claude Opus 5 was the first to tell us the question was self-contradictory - then hand us the working envelope instead of a compromise. Two requirements. Only one was survivable. New clip from the review series 👇
Town number 3 - a Wild West Frontier town.
This one was a lot more work (though also a lot farther along) than the previous 2, definitely not a one shot. A lot of lessons learned though, so I feel more optimistic the pipeline will do better with Town number 4. The goal with each
Anthropic dropped Opus 5, GPT-5.6 Luna gets an 80% price cut, DeepSeek V4 Flash is basically free
Meta releases spark and even xAI is throwing punches.
You have reached the end of the archive
All of Claude Opus 5