LLMcompare

Search

Search for a command to run...

Tool use & function calling

Tau-bench

Realistic tool-agent user interactions in airline, retail, and banking domains; scores are agent-system results, not model-only capability.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Tool use & function calling
Metric
Percent
Direction
Higher is better
Coverage
3 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Qwen3.8 Max

    Alibabaopen

    55.2%
  2. 2
    Claude Opus 5

    Anthropicclosed

    48.7%
  3. 3
    Grok 4.5

    xAIclosed

    47.9%

Back to all benchmarks