Interesting, I’m using my own harness too, renaming the tools to make their purpose more explicit helped a lot.
Just ran your exact same prompt, on Qwen 3.8 27B Q4, with reasoning set to xhigh. 20m10s 1.93M input tokens 66.9K output tokens 24 model calls Here’s the result 👇
Running Qwen 3.8 27b via Cerebras on a Letta Code fork.
So fast it parallelized too quick and hit a ton of rate limits lol
Battle of the Flashes!
Three are free on your DGX Spark One is a low cost API. DeepSeek V4 Flash GLM 5.3 Flash Gemini 3.8 Flash Qwen 3.8 Flash Next Choose your champion… vote in the comments!
Musk says it too. China will completely destroy the American AI industry.
The Chinese LLMs are much more efficient and much cheaper. There are now free models (such as Qwen 3.8-27B from Alibaba) that run on a $1500 video card (Intel B70 32GB VRAM).
You have reached the end of the archive
All of qwen38