I ran Qwen 3.8 27B on a single 8GB GPU.
IQ4_XS + Q4 KV cache at 70k context. The token speed hovers around 4-5 t/s. Slow but it's still amazing that you could get a result that can outperform even Claude Opus 4.6 on some of my test. All in one-shot with zero agent looping.
GLM 5.3 is free right now, and it just scored the highest number ever recorded on an independent coding benchmark.
91.25% on Kingbench 3. It beat Fable 5. It beat Opus 5. It beat Kimi K3. New users get it on the free tier inside Zcode. No card. No subscription. 5 million
I made a new benchmark for AI models, the "ramen test".
I wanted a way to measure initial vibes as someone who isn't really technical The goal is to create a cozy experience of eating a bowl of ramen, first person POV. Then they have pretty much free reign In this short video
Hermes + Kimmy K3 turns a chat model into a worker that runs for HOURS.
Kimmy K3 just hit #1 on the front-end code arena. Beat Fable 5, GPT 5.6, and Opus 4.8. Most people type into a chat box. That's the weakest way to use it. Make it the BRAIN of an agent instead: → /learn
You have reached the end of the archive
All of Claude Opus 5