I ran the same prompt 5 times using Deepseek v4 flash, and added a little bit of playability with additional instructions.
I never thought that just throwing the same prompt multiple times could dramatically change the results... From here on, if you give detailed instructions to improve the quality, I wonder if it will improve further...
The harness multiplier: same AI.
Same Tasks. 20-Point Swing. A public test just proved something huge: Your SETUP is quietly deciding half your AI results. The test: one model — DeepSeek V4 Flash. Eight different harnesses. 30 hard agentic tasks. The scores: Pi Agent: 20/30
一块 4090 跑通 AI 视频通话全栈:Qwen3.8-27B 本地部署,对话时延 2 秒
Qwen3.8-27B刚开源,我用了一晚上把它部署到本地 4090,把自研开源框架 VoxEMW 的 LLM 大脑从 DeepSeek API 换成了这个本地 27B,实测时延从 2.4 秒降到 2.0 秒。
Kimi K3 seems to be cooked for me, and honestly, Grok 4.6, Gemini 3.7, DeepSeek V4 are way better.
Built Football Frenzy, a playable mini soccer game demo, and saved the version (ID: db774d2). What's in the game: • 5-a-side match on a striped canvas pitch with full
You have reached the end of the archive
All of deepseek