Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees
arXiv cs.AIen
arXiv:2609.35874v1 Announce Type: new Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment TestingarXiv cs.AI · September 30, 2026
- Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge DevicesarXiv cs.AI · September 30, 2026
- More Programs or More Rolls? Separating Coverage from Specialization in LLM HarnessesarXiv cs.AI · September 30, 2026
- Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language ModelsarXiv cs.AI · September 30, 2026
- Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial TrainingarXiv cs.AI · September 30, 2026
- Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled ObjectivesarXiv cs.AI · September 30, 2026