How Language Models Choose Sides: Internal Representations of Instruction Hierarchy
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.28648v1 Announce Type: new Abstract: We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, conflict, and same-channel control conditions. Behaviourally, the models split into three regimes by System Authority Delta: hierarchy-respecting models use the system channel as an authority signal, anti-hierarchy models follow the system less often than their same-channel baseline predicts, and no-effect models show little channel sensitivity. Llama-3.1-8B is the strongest anti-hierarchy case in our suite, following
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Meta
- Forskning
Related AI news
- Anthropic launches Claude Fable 5.1 and Mythos 5.1, cuts agentic-task costs by up to 45%DIGITIMES · September 2, 2026
- AI 代理生態成關鍵考量,Meta 內部通訊工具棄 Google Chat 改用 SlackTechNews (TW) · September 2, 2026
- AI 代理生態成關鍵考量,Meta 通訊工具棄 Google Chat 改用 SlackTechNews (TW) · September 2, 2026
- 先進封裝邁向「化圓為方」!美商 ACM Research 卡位 FOPLP,電鍍、清洗、濕式蝕刻「三箭齊發」TechNews (TW) · September 2, 2026
- CrowdStrike builds security frontier models with Nvidia and opens an AI labSiliconANGLE · September 1, 2026
- Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and othersThe Decoder · September 1, 2026