Tool use & function calling
Tau-bench
Realistic tool-agent user interactions in airline, retail, and banking domains; scores are agent-system results, not model-only capability.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Tool use & function calling
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 3 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Qwen3.8 Max55.2%
Alibabaopen
- 2Claude Opus 548.7%
Anthropicclosed
- 3Grok 4.547.9%
xAIclosed