Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
arXiv cs.AIen
arXiv:2609.03493v1 Announce Type: new Abstract: Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropri
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Bild
Related AI news
- A profile of Hugging Face, which started in 2016 to build a sassy chatbot for teens; CEO Clément Delangue says he approached Nvidia this summer to pursue a deal (Wall Street Journal)Techmeme · September 4, 2026
- The sameness problem behind those unappetizing AI-generated menusTechCrunch AI · September 4, 2026
- Lite-On makes US$170 million strategic investment in DCX for AI cooling pushDIGITIMES · September 4, 2026
- K&S targets CPO, CoPoS growth with expanded TCB advanced packaging roadmapDIGITIMES · September 4, 2026
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency PenaltyarXiv cs.AI · September 4, 2026
- Analysis of Prompt Engineering for Drug Toxicity PredictionarXiv cs.AI · September 4, 2026