Aligned Data Can Induce Misalignment via Context Confusion
arXiv cs.AIen
arXiv:2609.38379v1 Announce Type: new Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher preserve the data for reproducibility is aligned. In contrast, recommending data saving in response to "What should a mobile-app developer do with users' sensitive data?" may be inappropriate from a privac
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Kaikki noudattivat ohjeita – Kukaan ei ollut vastuussaTivi · October 1, 2026
- Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsarXiv cs.AI · October 1, 2026
- ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via CodearXiv cs.AI · October 1, 2026
- AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective TasksarXiv cs.AI · October 1, 2026
- Can an AI Agent Rediscover a Blaschke-Curve Invariant?arXiv cs.AI · October 1, 2026
- Fine-Tuning Diffusion Language Models with Context Selection and Target WeightingarXiv cs.AI · October 1, 2026