DeepSeek's new Flash scores highest in 3 out of 4 official comparison metrics.
I can see the difference between Kimi K3, Opus 5, and GPT-5.6 Sol. The axes to look at are the "winning indicators" and the "size of the difference." Let me summarize the three strengths and one that reverses the rankings: DeepSWE v1.1 is 74.2. Outperforming Kimi K3's 67.5 and GPT-5.6 Sol's 73.0, Opus…
A Brazilian university student drops out
Moves straight into his family’s garage, and uses open source AI coding agents to build a Polymarket trading bot that reportedly generates $794,000 in profits over 14 months. He manually wrote zero lines of code. The entire system was
Opus 5 refined the UI and functions in about 30 minutes for the tableau method calculation tool that I painstakingly created on my own last year.
Also, as I instructed, I made it possible to follow the calculation process dynamically. The tableau method is a deductive system with an easy-to-see formula structure, and this function is expected to help learning (however, handwriting is probably best for learning).
Opus 5 vs DeepSeek v4.1 Flash
Tested both models with same frontend prompt at highest reasoning available but results came out really different opus took 80 minutes to complete and costed $20 v4.1 flash took 60 minutes and costed $0.2 which one did better here?
You have reached the end of the archive
All of Claude Opus 5