The Assistant's Ideal Self
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five published self-concept instruments are compared exhaustively in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Results show that models prioritize moral qualities, reflecting their alignment to 3H principles. Following, a desire for self-understanding emerges, as models p
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- UI-Venus-2 Technical ReportarXiv cs.AI · September 2, 2026
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal EstimationarXiv cs.AI · September 2, 2026
- MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical TextsarXiv cs.AI · September 2, 2026
- Asymmetries in Spontaneous and Instructed DeceptionarXiv cs.AI · September 2, 2026
- ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide GenerationarXiv cs.AI · September 2, 2026
- ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User FeedbackarXiv cs.AI · September 2, 2026