Fragility of Value under Imperfect Alignment
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent w
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous DocumentsarXiv cs.AI · August 3, 2026
- EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter DiagnosesarXiv cs.AI · August 3, 2026
- LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann HypothesisarXiv cs.AI · August 3, 2026
- Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic DiscoveryarXiv cs.AI · August 3, 2026
- MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping AgentsarXiv cs.AI · August 3, 2026
- On the Generalization of Steering Vectors for Chain-of-Thought FaithfulnessarXiv cs.AI · August 3, 2026