LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
arXiv cs.AIen
arXiv:2608.05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded i
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Företag
Related AI news
- 不用關掉 ChatGPT,Adobe 一口氣整合 70 項自家工具進 AITechNews (TW) · August 7, 2026
- Chinese AI firms push Hong Kong data center leasingTech in Asia · August 7, 2026
- Backed by DeepSeek, Unitree IPO tests investor appetite for China’s AI robotics boomSCMP Tech · August 7, 2026
- The Ignition Index: Measuring Global Workspace Dynamics in Language ModelsarXiv cs.AI · August 7, 2026
- Otter: A Time-Aware, History-Conditioned Human Chess AIarXiv cs.AI · August 7, 2026
- Project2Task: Graph-Guided Project-Level Planning for Autonomous ResearcharXiv cs.AI · August 7, 2026