MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation item
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN OptimizationarXiv cs.AI · September 23, 2026
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian SplattingarXiv cs.AI · September 23, 2026
- An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival PredictionarXiv cs.AI · September 23, 2026
- Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal EmbeddingsarXiv cs.AI · September 23, 2026
- X-Planner: Event-Structured Task Planning for Embodied IntelligencearXiv cs.AI · September 23, 2026
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026