Benchgen

TAU-bench Airline — Results

RankModelScore
1claude-sonnet-4-50.7
2minimax-m1-80k0.62
3glm-4-5-air0.608
4glm-4-50.604
5claude-sonnet-40.6
6minimax-m1-40k0.6
7qwen3-coder-480b-a35b0.6
8claude-opus-40.596
9claude-opus-4-10.56
10gpt-4-50.5
11o10.5
12gpt-4-10.494
13o4-mini0.492
14qwen3-next-80b-a3b-thinking0.49
15claude-3-5-sonnet0.46
16qwen3-235b-a22b-thinking-25070.46
17qwen3-next-80b-a3b-instruct0.44
18gpt-4o0.428
19gpt-4-1-mini0.36
20o3-mini0.324
21claude-3-5-haiku0.228

TAU-bench Airline

1 phaseActive

Evaluates language agents on airline domain conversations — booking changes, rule-following, and dynamic dialogue via API tools. Metric: success rate.

Overview

TAU-bench Airline

Category Metric Saturation Domain

Quick answer: TAU-bench Airline is the airline domain of τ-bench (TAU-bench), evaluating language agents' ability to interact with users through dynamic conversations while following airline-specific rules and using API tools for tasks like booking changes and cancellations. Claude Sonnet 4.5 leads with 70.0% across 23 evaluated models — notably harder than the retail domain.


What Does TAU-bench Airline Test?

TAU-bench Airline evaluates tool-agent-user interaction in an airline customer-service setting, requiring agents to interpret user requests, apply strict airline policy rules (e.g., fare rules, baggage allowances, rebooking windows), and invoke the correct sequence of API tools — all while maintaining a coherent multi-turn conversation with a simulated user.

Focus areaExamples
Multi-turn dialogueSustaining context across a realistic airline support conversation
Tool callingInvoking airline APIs (flight search, rebooking, cancellation)
Policy adherenceFollowing complex fare, baggage, and rebooking rules
User simulationAgent is evaluated against an LLM-simulated user with hidden intents

How Is TAU-bench Airline Scored?

Each task is graded by comparing the final environment/database state after the conversation to a ground-truth expected state, plus verifying required user communications occurred. A task is scored as fully successful (1) or failed (0); the reported score is the average success rate (pass^1) across the task suite, normalized to 0–1.


Key Facts

PropertyValue
ReleasedJune 2024
Created bySierra Research
DomainAirline customer service
MetricSuccess rate
Score range0–1
Top modelClaude Sonnet 4.5 (0.700)
Models evaluated23

FAQ

What is TAU-bench Airline? TAU-bench Airline is the airline domain of τ-bench, a benchmark for tool-agent-user interaction that evaluates language agents on realistic airline support conversations requiring strict rule-following and tool use.

Who created TAU-bench Airline? TAU-bench (τ-bench) was created by Sierra Research — Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan — published in June 2024.

Why is TAU-bench Airline harder than TAU-bench Retail? Airline policies involve more complex, interacting rules (fare classes, baggage, rebooking windows, multi-passenger itineraries), and even top models plateau well below saturation — the current average score across all models is just 0.5, versus 0.7 for retail.

What score does the best model achieve on TAU-bench Airline? Claude Sonnet 4.5 leads with 0.700 (70.0%), followed by MiniMax M1 80K at 0.620 and Zhipu AI's GLM-4.5-Air at 0.608.

Is TAU-bench Airline saturated? No. With a leader at just 0.700 and an average of 0.5 across 23 models, this remains one of the harder agentic benchmarks, with substantial room for improvement.