EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

arXiv cs.AIen

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

arXiv:2609.31906v1 Announce Type: new Abstract: Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol co

This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.

Read the full story at arXiv cs.AI
  • Forskning
  • Agenter
  • Företag

Related AI news