Benchmarking Prompt Optimization of Large Language Models With Chess
arXiv cs.AIen
arXiv:2610.00416v1 Announce Type: new Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution infrastructure. These challenges are amplified in automatic prompt optimization (APO), where evaluation is repeated throughout the search for better prompts. Studying APO therefore requires a benchmark that is cheap and deterministic to score, hard enough to leave room for improvement, and renewable as models evolve. We introduce a chess benchmark built from 1,118 Lichess puzzles to study APO for frozen LLMs: we optimiz
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMsarXiv cs.AI · October 2, 2026
- Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?arXiv cs.AI · October 2, 2026
- Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational HumorarXiv cs.AI · October 2, 2026
- Before Agents Decide: Epistemic Action in LLM-Based SystemsarXiv cs.AI · October 2, 2026
- Heavy-Tailed Memory Traces in Long-Horizon Language AgentsarXiv cs.AI · October 2, 2026
- From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent ExecutionarXiv cs.AI · October 2, 2026