Τ³-Bench is out
can agents handle complex docs and live calls?
Details
- External ID
- 47520448
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- March 25, 2026
- Cohort
- —
- Upvotes
- 12
- Upvotes percentile
- 0.6439114391143912
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
τ-Bench is an open benchmark for evaluating AI agents on grounded, multi-turn customer service tasks with verifiable outcomes. It's been great to see the community adopt it since launch — this is now the third iteration. With τ³-Bench, we're extending it to two new settings: knowledge-intensive retrieval and full-duplex voice.τ-Knowledge: agents must navigate ~700 interconnected policy documents to complete multi-step tasks. Best frontier model (GPT-5.2, high reasoning) hits ~25%. The surprising part: even when you hand the model the exact documents it needs, performance only reaches ~40%. We found that the bottleneck isn't retrieval — it's reasoning over complex, interlinked policies and executing the right actions in the right order.τ-Voice: same grounded tasks, but over live full-duplex voice with realistic audio — accents, background noise, interruptions, compressed phone lines. Voice agents score 31–51% in clean audio conditions and 26–38% in realistic ones. A consistent failure pattern across providers (OpenAI, Gemini, xAI): agent mishears a name or email during authentication, and everything downstream fails.We also incorporated 75+ task fixes to the original airline, retail, and telecom domains — many based on community audits and PRs (including contributions from Amazon and Anthropic). We believe a benchmark is only as good as its maintenance, and we're grateful for the community's help improving it.Code and leaderboard are open — we'd welcome community submissions and feedback.Blog post (papers, code, leaderboard): https://sierra.ai/blog/bench-advancing-agent-benchmarking-to...
Enrichment
- Theme
- voice AI agents and infrastructure
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark for agents handling complex documents and calls
- Manually corrected
- False
Could you build this?
No Tau-bench is an academic and industrial benchmark requiring novel evaluation methodologies, simulated interactive environments, verifiable state checks, and scientific dataset curation.
What it would actually take: Creating this requires deep ML research expertise to design deterministic multi-turn simulation environments, ground-truth database states, and call-transcript simulators. The infrastructure needs mock API servers, state validation suites, and rigorous statistical evaluation harnesses to test complex reasoning and error recovery across LLM models without data leakage.
Discussion
1 comment analyzed.
Competitors mentioned: tau bench, tau voice
Concerns raised: knowledge and context rot in multimodal models, audio quality and duplex capabilities
Competitors
Other products that read as similar to this one — 479 launches clear the similarity bar, closest 8 shown.
Attention rank: #155 of 480 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 146 days after the earliest competitor.
- Cua-Bench · hn · 2026-01-26 · 40 upvotes · similarity 0.51
- Dialogus: Building the infra for enterprise voice agents · yc · 2026-07-24 · 10 upvotes · similarity 0.47
- benchboard · github · 2026-09-18 · 18 upvotes · similarity 0.46
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.45
- Callab AI - AI voice agents for legacy phone systems. · yc · 2026-05-22 · 14 upvotes · similarity 0.45
- Simreal-MLBench · github · 2026-09-21 · 71 upvotes · similarity 0.44
- Speko: OpenRouter for Voice · yc · 2026-07-29 · 14 upvotes · similarity 0.44
- 🏦 Monumint: Voice AI for financial services · yc · 2026-07-09 · 18 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.