FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.20574v1 Announce Type: new Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating dif
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Source: AI researcher Luke Metz, who returned to OpenAI from TML earlier this year, joins Meta's Superintelligence Labs and will report to Alexandr Wang (Ina Fried/Axios)Techmeme · August 24, 2026
- Terminal Agents: A Survey of AI Agents in Command-Line EnvironmentsarXiv cs.AI · August 24, 2026
- Difficulty-Aware Semantic-ID Optimization for Generative RecommendationarXiv cs.AI · August 24, 2026
- Environmental Slow AI: Design Principles for Generative SystemsarXiv cs.AI · August 24, 2026
- World models of environment, agent and joint agent-environment systemsarXiv cs.AI · August 24, 2026
- StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language ModelsarXiv cs.AI · August 24, 2026