When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
arXiv cs.AIen
arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting. MERIT provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check; a difficulty ladder ending in updated-fact recall; controlled memory corruption; and
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Företag
Related AI news
- Sources: China Securities Regulatory Commission is informally tightening IPO approvals for humanoid startups after a volatile debut by industry leader Unitree (The Information)Techmeme · September 9, 2026
- LG Innotek tackles glass substrate microcrack issue as 2028 production race heats upDIGITIMES · September 9, 2026
- Apple yields to memory suppliers in historic strategy shift, with Kioxia tipped as NAND long-term agreement recipientDIGITIMES · September 9, 2026
- Mittwoch: Huawei-Verstöße gegen US-Sanktionen, Metas privater KI-Agent für alleheise online – KI · September 9, 2026
- 6 av 10 anställda saknar tiden innan AI fannsComputer Sweden · September 9, 2026
- Damage-Aware Bandit Pruning for Vision and Language TransformersarXiv cs.AI · September 9, 2026