$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
arXiv cs.AIen
arXiv:2609.13602v1 Announce Type: new Abstract: Voice agents often need to collect names, addresses, identifiers, dates, and times exactly, yet end-to-end benchmarks obscure where capture fails. We introduce $\tau$-Elicitation, a 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms, and three environments. A matched text agent passes all tasks, but four voice configurations achieve robust exact success from 0.14 to 0.41. Agents increase verification for hard and unfamiliar entities and sometimes for incorrect captures, but not for their weakest caller voice; only 24 to 37 percent of verified errors are repaired. A scaffold that prescribes spelling, read-b
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- Agenttisesta tekoälystä tuli testaajan jatkuva kaveriTivi · September 15, 2026
- AI agents are moving faster than Europe can regulateSifted · September 15, 2026
- Microsoft commits to sweeping AI privacy rules for students. Will other tech giants follow?Economic Times Tech · September 15, 2026
- OrchSLM: Probing the Dynamics of Small Language Model OrchestrationarXiv cs.AI · September 15, 2026
- Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting AgentsarXiv cs.AI · September 15, 2026
- Token Efficient Task Execution via Application Behavior Modeling for Web AgentsarXiv cs.AI · September 15, 2026