For tool calling, GLM5.3-Flash gains the edge over Deepseek V4 flash 0731.
This is the expected result, and both models are extremely relevant when it comes to tool calling. Either can be enjoyed reliably. But my daily driver? GLM.
GLM5.3-Flash NVFP4 DFlash2 scores 86 on tool-eval bench, on the the full 69 scenario test⚡️
My strongest result as of yet. Comparison against Deepseek V4 Flash 0731 coming up next! Tool being used is Tool-eval bench at
Qoder is giving away 2 weeks of access to a bunch of top AI models 😳
You can try them all without setting up a single API key available models: - Qwen3.8-Max / 3.7-Max / Plus - Kimi-K3 / K2.7-Code - GLM-5.3 / 5.2 - DeepSeek-V4-Pro / Flash - MiniMax-M3 getting started: 1. go
DeepSeek worked well only 107 seconds with all the working functions 👍
You have reached the end of the archive
All of deepseek