VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets
arXiv cs.AIen
arXiv:2609.12404v1 Announce Type: new Abstract: Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that usi
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- Hong Kong IPO boom and Chinese AI race fuel frenzy of follow-on dealsNikkei Asia · September 14, 2026
- SoftBank Group plunges 11% after OpenAI says no IPO this yearNikkei Asia · September 14, 2026
- Filing: San Jose-based Chinese optical module maker Ligent seeks ~$727M in a Hong Kong IPO, after its rival Zhongji Innolight's $6.8B Hong Kong debut in July (Sangmi Cha/Bloomberg)Techmeme · September 14, 2026
- New warnings about the risks of AI to humanity revive a long-running debateEconomic Times Tech · September 14, 2026
- China state newspaper blasts Anthropic's calls to slow AI as 'Cold War' tacticEconomic Times Tech · September 14, 2026
- Anthropic tells investors it will be profitable for second straight quarterEconomic Times Tech · September 14, 2026