Benchgen

Tau2 Telecom — Results

RankModelScore
1claude-opus-4-60.993
2gpt-5-40.989
3gpt-5-20.987
4claude-opus-4-50.982
5gpt-5-50.98
6claude-sonnet-4-60.979
7mimo-v2-pro0.968
8gpt-50.967
9gpt-5-1-instant0.956
10gpt-5-1-thinking0.956
11gpt-5-10.956
12nova-2-pro0.927
13muse-spark0.915
14minimax-m2-10.87
15minimax-m20.87
16longcat-flash-thinking0.831
17claude-haiku-4-50.83
18nova-2-omni0.8
19nova-2-lite0.76
20longcat-flash-chat0.737
21longcat-flash-lite0.728
22kimi-k2-instruct-09050.658
23kimi-k2-instruct0.658
24nemotron-3-super-120b-a12b0.644
25o30.582

Tau2 Telecom

1 phaseActive

τ²-Bench telecom domain — evaluates conversational agents in a dual-control environment where both agent and user use tools for shared troubleshooting. Metric: success rate.

Overview

Tau2 Telecom

Category Metric Saturation Domain

Quick answer: Tau2 Telecom is the telecom domain of τ²-Bench, which evaluates conversational agents in a dual-control environment modeled as a Dec-POMDP — where both the AI agent and the simulated user can use tools — on shared telecommunications troubleshooting scenarios. LongCat-Flash-Thinking-2601 and Claude Opus 4.6 tie for the lead at 99.3% across 35 evaluated models.


What Does Tau2 Telecom Test?

Unlike the original τ-bench, where only the agent uses tools, τ²-Bench introduces a dual-control setting: both the agent and the (simulated) user can invoke tools to resolve a shared telecom troubleshooting task, such as diagnosing a connectivity issue that requires actions on both sides. This tests coordination, communication, and joint problem-solving rather than just one-sided tool execution.

Focus areaExamples
Dual-control coordinationAgent and user both act on a shared environment via tools
CommunicationClearly conveying instructions the user must execute themselves
TroubleshootingDiagnosing and resolving telecom connectivity/service issues
Tool callingInvoking diagnostic and account-management APIs

How Is Tau2 Telecom Scored?

Each task is graded by comparing the final environment state (across both agent-controlled and user-controlled actions) to a ground-truth expected outcome. A task is scored as fully successful (1) or failed (0); the reported score is the average success rate (pass^1) across the task suite, normalized to 0–1.


Key Facts

PropertyValue
ReleasedJune 2025
Created bySierra Research
DomainTelecom troubleshooting (dual-control)
MetricSuccess rate
Score range0–1
Top modelLongCat-Flash-Thinking-2601 / Claude Opus 4.6 (0.993)
Models evaluated35

FAQ

What is Tau2 Telecom? Tau2 Telecom is the telecom domain of τ²-Bench, a benchmark evaluating conversational AI agents in a dual-control environment where both the agent and a simulated user use tools to jointly resolve telecommunications troubleshooting tasks.

Who created Tau2 Telecom? τ²-Bench was created by Sierra Research — Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and collaborators — published in June 2025, as a successor to the original τ-bench.

How is τ²-Bench different from τ-bench (TAU-bench)? The original τ-bench only lets the agent use tools while the user is a passive information provider. τ²-Bench models the interaction as a dual-control Dec-POMDP, where the user must also take tool-based actions — a more realistic simulation of scenarios like telecom troubleshooting where the customer performs steps on their own device.

What score does the best model achieve on Tau2 Telecom? LongCat-Flash-Thinking-2601 (Meituan) and Claude Opus 4.6 (Anthropic) tie for the lead at 0.993 (99.3%), with GPT-5.4 close behind at 0.989.

Is Tau2 Telecom saturated? Near the top, yes — the leading models cluster above 0.95. But scores drop off sharply further down the leaderboard, with several models below 0.5 and the weakest at 0.132, so the benchmark still separates strong from weak agentic tool-use.