Gemini 3.7 Flash vs DeepSeek V4 Pro 0318 vs Muse Spark 1.2 vs Grok 4.6
The task: Create an eagle using triangles only. 🦅 Same prompt. Gemini 3.7 Flash absolutely nailed the challenge. The result is way better than I expected.
1 turns · 92 steps| LLM 70m21s · Tool call 4m55s| TTFT avg 1.5s · 35 tok/s| Cache hit 99%| Input 10.2M tok| Output 143K tok
Using the official harness + deepseek-v4-pro high, it took more than an hour to run the rough-like game: 1️⃣ 70m comes out in 1 turn, autonomous & stable for long time operation. The game is very complete 2️⃣ Cache hit
Since opencode go's deepseek-v4-flash has shrunk significantly
Since opencode go's deepseek-v4-flash has shrunk significantly, the current price-performance ratio should be to subscribe to the chatgpt plus package, and then use gpt-5.6-luna max in the codex+deepseek harness👀 I have slowly started to move from codex to dsh
Qwen 3.8 27b
DeepSeek V4 Flash 0731 GPT 5.6 Luna Sub-models rarely pass this test.
You have reached the end of the archive
All of deepseek