ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
arXiv cs.AIen
arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers. We constructed and released six populated schemas (465 tables, 164,682 rows, zero empty tables) with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Ser
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Peak XV invests $15m in Indian voice AI startup Ringg AITech in Asia · August 26, 2026
- Indian crypto exchange WazirX unveils AI trading assistantTech in Asia · August 26, 2026
- Digs, which is building AI software for residential construction, raised a $25.3M Series A led by building materials giant Builders FirstSource (Kurt Schlosser/GeekWire)Techmeme · August 26, 2026
- A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust CertificationarXiv cs.AI · August 26, 2026
- A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshiftsarXiv cs.AI · August 26, 2026
- AI Agents Push Humans Out of the LooparXiv cs.AI · August 26, 2026