GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
arXiv cs.AIen
arXiv:2609.03553v1 Announce Type: new Abstract: Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstruct
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Reglering
Related AI news
- G20 approves guidelines calling for more business-friendly AI regulationDIGITIMES · September 4, 2026
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency PenaltyarXiv cs.AI · September 4, 2026
- Analysis of Prompt Engineering for Drug Toxicity PredictionarXiv cs.AI · September 4, 2026
- Dalek: A Constructive Agent MachinearXiv cs.AI · September 4, 2026
- HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer ReviewsarXiv cs.AI · September 4, 2026
- A computable representation of the physical laboratory enables verifiable workflowsarXiv cs.AI · September 4, 2026