How User-AI Mistreatment Occurs and Matters in Conversational Systems?
arXiv cs.AIen
arXiv:2609.13579v1 Announce Type: new Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderat
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Dienstag: Trump greift Anthropic-Chef an, Apple-Betriebssysteme für die KI-Äraheise online – KI · September 15, 2026
- Microsoft commits to sweeping AI privacy rules for students. Will other tech giants follow?Economic Times Tech · September 15, 2026
- OrchSLM: Probing the Dynamics of Small Language Model OrchestrationarXiv cs.AI · September 15, 2026
- Asclepius: An Adaptive Harness for Long-Horizon Clinical AgentsarXiv cs.AI · September 15, 2026
- Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting AgentsarXiv cs.AI · September 15, 2026
- Token Efficient Task Execution via Application Behavior Modeling for Web AgentsarXiv cs.AI · September 15, 2026