Towards AI spent about $600 on evals racing 6 context-trimming techniques against doing nothing
Doing nothing won Keeping the full history cost $0.11 per turn against $0.24 for their tuned production setup, answered 4 seconds faster, and recalled 92% of earlier facts against
Deepseek-v4.1-flash made this, with help from LTX 2.5
First look at today's value board: GLM 5.3 Flash leads with a 42 score at just 9 cents
While DeepSeek V4 Flash offers the cheapest entry at 3 cents.
You have reached the end of the archive
All of deepseek