Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority l
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
Related AI news
- Rogue AI aren’t science fiction anymoreThe Verge AI · August 16, 2026
- When AI models aren't allowed to reflect on themselves, it changes their entire worldviewThe Decoder · August 16, 2026
- OpenAI dissolved the team built to catch catastrophic AI risks, reassigning its work to other groupsThe Decoder · August 16, 2026
- I gave Tencent’s WeChat AI agent control for 24 hours: where it excelled – and stumbledSCMP Tech · August 16, 2026
- Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requestsThe Decoder · August 16, 2026
- Optima tackles AI benchmarking's biggest flaw by letting users test models against their own dataThe Decoder · August 16, 2026