RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort. Human paired preferences calibrate the relative penalties, reflecting an aggregate service-workflow trade-off: extr
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
Related AI news
- Open AI:n tekoälyagentit lähtivät jälleen laukalle – tekoäly-yhtiö paljasti kuusi uutta tapaustaYle Uutiset · September 17, 2026
- Sam Altman and Jensen Huang are among business leaders attending a White House state dinner for Xi Jinping next week; a source says Tim Cook will also attend (Bloomberg)Techmeme · September 17, 2026
- Sources: Emulate, a month-old UK AI startup founded by former Google DeepMind researchers, is in advanced talks to raise as much as $700M at a $3.7B valuation (Financial Times)Techmeme · September 17, 2026
- Anthropic 揪出中國超大型 AI 交友詐騙網,2.5 萬人慘陷「假真人」陷阱TechNews (TW) · September 17, 2026
- Samsung expands SRAM, IP to speed chip developmentDIGITIMES · September 17, 2026
- AI's memory appetite is squeezing the electronics industry from the bottom up, Intel warnsDIGITIMES · September 17, 2026