Thinking models are not getting any smarter when they ponder for two minutes
They are simply taking sequential steps in a loop with notes while draining your API credits Prompt Engineering demonstrated this in a head-to-head test against Jev on shuffled company hierarchies Jev
Someone wired an iphone 17 pro max into a 24gb m4 pro macbook over usb-c and split qwen 3.8 27b across both.
Mac alone at 16k context: 109 tok/s prefill. with the phone: 157 tok/s. +44%. at 32k: 101 to 130 tok/s. +29%. the phone also holds part of the kv cache past 64k so the
Benchmarks say these 4 models are almost equal
Then you ask them to build a simple 3D farm They are NOT equal Gemini 3.7 Flash vs Qwen 3.8 vs Grok 4.6 vs GLM 5.3 ↓
Stained Glass Projector - Qwen 3.8 27B
So this one took 11 hours and a couple of restarts. At about 9 hours, all three images worked. Looks like it got confused and broke the color quantization math. Bummer.
You have reached the end of the archive
All of qwen38