| Rank | Model | Score |
|---|---|---|
| 1 | claude-opus-4-6 | 0.919 |
| 2 | claude-sonnet-4-6 | 0.917 |
| 3 | claude-opus-4-5 | 0.889 |
| 4 | longcat-flash-thinking-2601 | 0.886 |
| 5 | claude-haiku-4-5 | 0.832 |
| 6 | gpt-5-2 | 0.82 |
| 7 | gpt-5 | 0.811 |
| 8 | o3 | 0.802 |
| 9 | nova-2-omni | 0.783 |
| 10 | gpt-5-1-instant | 0.779 |
1 phaseActive
τ²-Bench retail domain — evaluates conversational agents in a dual-control setup on order changes, refunds, and account actions. Metric: success rate.
Quick answer: Tau2 Retail is the retail domain of τ²-Bench, which evaluates conversational agents in a dual-control environment — where both the AI agent and the simulated user can invoke tools — on order changes, refunds, and account-management tasks. Claude Opus 4.6 leads at a 91.9% success rate.
Retail support tasks in τ²-Bench require an agent and a simulated customer to jointly resolve requests like modifying an order, processing a return, or updating account details. Both sides can act on the shared environment via tools, so the benchmark tests coordination and correct policy handling, not just single-shot tool selection.
| Focus area | Examples |
|---|---|
| Dual-control coordination | Agent and user both act on a shared order-management environment |
| Policy handling | Applying return windows, refund eligibility, and account-change rules correctly |
| Communication | Giving the customer clear instructions for steps only they can perform |
| Order/account actions | Modifying orders, issuing refunds, updating shipping or account details |
Each task is graded by comparing the final environment state — across both agent-controlled and user-controlled actions — against a ground-truth expected outcome. A task is scored as fully successful (1) or failed (0); the reported score is the average success rate across the task suite, normalized to 0–1.
| Property | Value |
|---|---|
| Released | June 2025 |
| Created by | Sierra Research |
| Domain | Retail customer support (dual-control) |
| Metric | Success rate |
| Score range | 0–1 |
| Top model | Claude Opus 4.6 (0.919) |
No Benchgen results yet — be the first to run Tau2 Retail.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Tau2 Retail | Dual-control retail customer support | Medium |
| Tau2 Airline | Dual-control airline customer support | Low |
| Tau2 Telecom | Dual-control telecom troubleshooting | High |
What is Tau2 Retail? Tau2 Retail is the retail domain of τ²-Bench, a benchmark evaluating conversational AI agents in a dual-control environment where both the agent and a simulated user use tools to jointly resolve retail customer-support tasks like order changes, refunds, and account updates.
Who created Tau2 Retail? τ²-Bench was created by Sierra Research — Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and collaborators — published in June 2025, as a successor to the original τ-bench (Tau-bench Retail).
How is τ²-Bench different from τ¹ (Tau-bench Retail)? The original τ-bench only lets the agent use tools while the user is a passive information provider. τ²-Bench models the interaction as a dual-control Dec-POMDP, where the user must also take tool-based actions — closer to real retail support where the customer confirms addresses or reads back order details themselves.
What score does the best model achieve on Tau2 Retail? Claude Opus 4.6 leads at 0.919 (91.9%), with Claude Sonnet 4.6 close behind at 0.917.
Is Tau2 Retail saturated? Partially. The top cluster of Claude models sits close to 0.9, but scores fall off meaningfully below the top 10 — there's still room to differentiate strong from average agentic tool-use.