Benchy: towards a universal language for task-oriented AI benchmarks
arXiv cs.AIen
arXiv:2609.30550v1 Announce Type: new Abstract: Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine executes. Compilation changes representation, not meaning: it does not repair invalid definitions or inject hidden defaults. Programs use fixed schemas
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Bringing AI to Autonomous Systems -- From Cognition to Collective IntelligencearXiv cs.AI · September 28, 2026
- Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context ProtocolarXiv cs.AI · September 28, 2026
- Predicting Transmembrane Protein Topology from 3D StructurearXiv cs.AI · September 28, 2026
- Spectral Feedback for Test-Time Alignment of Protein Diffusion ModelsarXiv cs.AI · September 28, 2026
- Pretrained ASR Pseudo-labeling for Noisy Police AudioarXiv cs.AI · September 28, 2026
- Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search StudyarXiv cs.AI · September 28, 2026