Actual testing of more than a dozen models
Actual testing of more than a dozen models: Who is better at writing human words? Today's large models, agents and programming capabilities are getting stronger and stronger, but the things written may become less and less human-like. Recently, I was adjusting my GEO writing skills and tested more than a dozen models in a row.
You can use GPT-5.6 Luna and DeepSeek V4 Pro for free.
Freebuff gives you access to: → DeepSeek V4 Pro - smartest → DeepSeek V4 Flash - smart + fast → GPT-5.6 Luna - all-around → MiniMax M3 - fastest → MiMo 2.5 - balanced → GLM 5.2 - available through earned sessions and
Deepseek-4-flash (0731 version)
Duration: 1h 25m Total tokens: 38.9M Observations: It got stuck a lot in the trap that the prompt had (Prompt below) that's why so many tokens but acceptable result
The gray testing of the model has disappeared since the peak period, and I suspect it is true that…
The gray testing of the model has disappeared since the peak period, and I suspect it is true that DeepSeek employees can use special models for work (below are a few examples of gray testing performance).
You have reached the end of the archive
All of deepseek