Do Frontier Models Seek Safety Evidence Before Acting?
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation. Across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, we find distinct evidence-acquisition policies: Opus inspects nearly by default, o3 is the most skip-heavy and threshold-sensitive, and GPT-5.5 and Sonnet occupy intermediate regimes. Inspection
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- OpenAI
- Anthropic
- Forskning
- Företag
Related AI news
- Nvidian ja Metan johtajilta tyrmäys alan vaatimuksille: Tekoälyä ei tarvitse suitsiaTivi · September 17, 2026
- Sam Altman and Jensen Huang are among business leaders attending a White House state dinner for Xi Jinping next week; a source says Tim Cook will also attend (Bloomberg)Techmeme · September 17, 2026
- Sources: Emulate, a month-old UK AI startup founded by former Google DeepMind researchers, is in advanced talks to raise as much as $700M at a $3.7B valuation (Financial Times)Techmeme · September 17, 2026
- Sources: Manus is set to soon close a $500M funding round at a $4B valuation, signaling growing confidence after Beijing unwound its $2B buyout by Meta (Bloomberg)Techmeme · September 17, 2026
- Anthropic 揪出中國超大型 AI 交友詐騙網,2.5 萬人慘陷「假真人」陷阱TechNews (TW) · September 17, 2026
- Claude 改版升級,Chat 與 Cowork 合併、加強文件和簡報製作能力TechNews (TW) · September 17, 2026