Interesting results after testing ponytail and caveman with deepseek-v4-flash-vision-exp to build a zoo
I had two similar sessions, one with ponytail and caveman and one without. The results clearly show WITHOUT is much better. What's even crazier, token usage: With: 16.9M
Deepseek open sourced their agent harness and raised their api prices in the same week.
Bold combo Because the harness treats the model as a plugin. So mid-task I unplugged their api and plugged in qwen3.8 27b running through rcli on a 24gb gpu. It finished the job like nothing
Not touching the context won the evals.
That's the surprise at the center of "Context Engineering in 2026," where Whats_AI, omar_solano1, and samridhivaid of Towards AI walk through what they measured while trying to fix their AI tutor. aiDotEngineer has it on YouTube. If you
Qwen3.8-27B vs DeepSeek V4 Flash Same Eiffel Tower voxel prompt, 3 thinking levels each.
You have reached the end of the archive
All of deepseek