Tool use & function calling
Toolathlon-Verified
Verified subset of Toolathlon, a tool-use benchmark measuring selecting, sequencing, and completing multi-step tasks with external tools; keep distinct from unverified Toolathlon runs.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Tool use & function calling
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 27 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1GLM-5.3 Flash78.4%
Zhipu AIopen
- 2Kimi K376.5%
Moonshotopen
- 3Claude Opus 4.876.2%
Anthropicclosed
- 4Muse Spark 1.275.9%
Metaclosed
- 5Muse Spark 1.175.6%
Metaclosed
- 6DeepSeek V4 Pro 081374.1%
DeepSeekclosed
- 7Hy4 Preview74.1%
Tencentopen
- 8GPT-5.573.5%
OpenAIclosed
- 9GLM-5.373.0%
Zhipu AIopen
- 10Qwen3.8 Max72.5%
Alibabaopen
- 11Claude Sonnet 571.6%
Anthropicclosed
- 12DeepSeek V4 Flash 073170.3%
DeepSeekopen
- 13Gemini 3.5 Flash67.3%
Googleclosed
- 14Gemini 3.1 Pro61.1%
Googleclosed
- 15GLM-5.259.9%
Zhipu AIopen
- 16Kimi K2.7 Code58.0%
Moonshotopen
- 17Kimi K2.658.0%
Moonshotclosed
- 18Gemini 3.5 Flash-Lite57.1%
Googleclosed
- 19Hunyuan 3 Preview56.8%
Tencentclosed
- 20DeepSeek V4 Pro55.9%
DeepSeekopen
- 21Inkling-Small54.4%
Thinking Machinesopen
- 22DeepSeek V4 Flash49.7%
DeepSeekopen
- 23MiniMax M2.747.5%
MiniMaxclosed
- 24Inkling45.5%
Thinking Machinesopen
- 25Qwen3.5 397B-A17B40.7%
Alibabaopen
- 26Nemotron 3 Ultra34.3%
NVIDIAopen
- 27Kimi K2.533.0%
Moonshotopen