Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exa
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN OptimizationarXiv cs.AI · September 23, 2026
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian SplattingarXiv cs.AI · September 23, 2026
- An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival PredictionarXiv cs.AI · September 23, 2026
- Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal EmbeddingsarXiv cs.AI · September 23, 2026
- X-Planner: Event-Structured Task Planning for Embodied IntelligencearXiv cs.AI · September 23, 2026
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026