Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
arXiv cs.AIen
arXiv:2609.13637v1 Announce Type: new Abstract: Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates while keeping scoring oracles outside the target process. Two frozen campaigns cover sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations, yielding 1,536 retained responses. A judge-independent literal audit finds direct-parent identifiers in 48/
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- AI agents are moving faster than Europe can regulateSifted · September 15, 2026
- Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting AgentsarXiv cs.AI · September 15, 2026
- Token Efficient Task Execution via Application Behavior Modeling for Web AgentsarXiv cs.AI · September 15, 2026
- Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment GenerationarXiv cs.AI · September 15, 2026
- Asclepius: An Adaptive Harness for Long-Horizon Clinical AgentsarXiv cs.AI · September 15, 2026
- AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web AgentsarXiv cs.AI · September 15, 2026