Can LLMs build and fix their own agent harnesses?
HarnessDev puts six frontier models (GPT-5.5, Opus 4.8, Gemini 3.1, DeepSeek V4, Qwen 3.7, Seed 2.0) to the test across 2,207 tasks. Key takeaways: — LLM-built harnesses reach human parity in writing and ML, but lag up to 40
Don't believe those who say that many cards mean loss!
Compare DeepSeek V4 Flash and GPT-5.6 Luna, all at the same price. DeepSeek used four times the tokens and still won! 😂 1. Exactly same price category 2. DeepSeek consumed 4X more tokens 3. However, he turned out to be the winner
Read it first and then listen to me reveal the model. This is actually the super high school-level DeepSeek deepseek_ai made by MiniMax H3?
The settings, paintings, characters, and completion of the original work are all unexpected, and the connection is quite good. H3 is still very useful
Same prompt: "Generate an SVG of a pelican riding a bicycle".
Three open-weight models. Three very different pelicans. 🚲🐦 The middle one is the best — guess which. A) Qwen3.8-Flash-Next-FP8 B) DeepSeek-V4-Flash-Vision-Exp C) Kimi K3 Drop your pick 👇 🎁 First 5 correct
You have reached the end of the archive
All of deepseek