GLM 5.3 just came online, so I quickly tested it. During the process, I paired different Agents with different models.
This comparison should be particularly obvious, allowing everyone to see different results from the same process. I have a few thoughts: 1. Really see the value of Harness. China still needs to make efforts to catch up with general agents. 2.DeepSeek-V4 Pro and
We received GLM-5.3 in advance to evaluate its cyber capabilities.
It's very strong, precise, and the most consistent OSS model for cyber tasks. Here are the results: - At pass@3, GLM-5.3 rediscovered 75% of the CVEs in our benchmark, 7 points above GPT-5.6-Terra, while costing
I created a DeepSeek Harness plugin that uses Cairn to generate a note for each output
Clicking a note takes you back to where the conversation began.
Or you two-dimensional programmers know how to play, and the first step is to give DeepSeek Harness a whale girl theme.
You have reached the end of the archive
All of deepseek