K-Bench: measuring model performance on real scientific agent requests
arXiv cs.AIen
arXiv:2608.21601v1 Announce Type: new Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-ancho
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- AI chipmaker Enflame sets subscription date for near $900 million Shanghai IPOEconomic Times Tech · August 25, 2026
- Measuring Activation Control in Large Language ModelsarXiv cs.AI · August 25, 2026
- KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model InferencearXiv cs.AI · August 25, 2026
- AIREP: A Protocol for Per-Decision Evidence in AI Runtime GovernancearXiv cs.AI · August 25, 2026
- LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review PlatformarXiv cs.AI · August 25, 2026
- SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAGarXiv cs.AI · August 25, 2026