I gave 4 "Flash" models the same job: redact the PII in an HR letter using an agentic mask tool.
One of them cost 630× more than another. And it wasn't 630× better. Results 🧵 🥇 GLM 5.3 Flash $0.005, 15 tool calls Clean sweep. Overshot a bit (also redacted employee ID +
I have some MTP dense benchmarks Qwen 3.8 27B Q4 MTP vs Qwen 3.6 35B Q4 MTP....
Poor 3.8 27b lol
Qwen 3.8 27B is amazing. They developed the game in no time.
Environment is RTX 3090 RAM 32GB I was able to do it by simply issuing a few instructions remotely using Tailscalse from my smartphone in the OpenHnads container. Please give it a try!
We put Fable 5.1, Muse Spark 1.3, Gemini 3.8 Flash, and Qwen 3.8 Max through the same bank heist FPS shooter prompt.
The results were honestly surprising. 👀 Fable 5.1 had my favorite map design, but it also cost me the most. Gemini 3.8 Flash was the speed demon - it generated
You have reached the end of the archive
All of qwen38