Same programming error. Same project.
And the same prompt. But the slowest run took 3.2 times as long as the fastest: Claude Code — 50 seconds OpenCode + DeepSeek — 74 seconds Codex — 160 seconds Kimi is excluded from the comparison due to Rate Limit The three who completed the task passed 79/79 tests. Just one test, not a general ranking…
New frontier release alert!
StepFun takes the lead with Step 5 Preview, out today taking over the latest model release position on Shaping Pareto frontier. Surpassing Kimi K3, DeepSeek v4 pro and GLM 5.3 in benchmark. Delivering frontier-level
I just ran claude opus 4.8 and deepseek V4.1 through an API for $0
ShareLLM accepted both requests from a fresh account with no deposit and a $0 balance The free limits: 120 calls every 5h 600 calls per week Setup: Sign up at sharellm Create an API key Set the base
Best way to test Step-5-preview is with comparison.
Here's it compared with Deepseek-v4.1-flash. Left: DS-v4.1-flash (high) Right: Step-5-preview (high) prompt: "A Three.js scene of giant red mesas in a desert, long shadows, a huge empty sky." one-shot result, single html
You have reached the end of the archive
All of deepseek