I ran the same todo-app prompt through the same agent harness a 3rd time.
Same model, same prompt. Only the process changed — and it changed everything. Run 2: 3h17m😲 (85 minutes of it burned by one subagent chasing a bug it introduced). Run 3: 43 minutes. 102 model calls,
This is wild, never thought I'd be running a 27B-parameter model on my phone.
Running Alibaba_Qwen's Qwen 3.8 27B (1-bit) on my iPhone 17 Pro. gonna share the benchmarks very soon. Stay tuned! You can try running it on your own phone using the RunAnywhereAI apps available on
Is DeepSeek V4 Flash still the one to beat?
TonyD2Wild thinks so. Pound for pound he still has it over Qwen 3.8 Flash on the benchmarks. DeepSeek stays a big dog.
All made local with qwen 3.8 27b researching and generating lyrics and minimax music generating the song
You have reached the end of the archive
All of qwen38