OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model (Celia Ford/Transformer)
Techmemeen

Celia Ford / Transformer : OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model — OpenAI is hailing its new model as “the world's most intelligent and aligned”, but the details reveal an awareness of being evaluated …
This is a short summary published by AI Global Wire. The full article is owned and hosted by Techmeme — open it there to read it in full.
Read the full story at Techmeme- OpenAI
Related AI news
- OpenAI rolls out GPT-6 Astra to Pro customers on the $100/month or $200/month plans (Zac Hall/9to5Mac)Techmeme · September 4, 2026
- Report: OpenAI learned of the DseWiki German website incident weeks ago but kept it under wraps as it grappled with the Hugging Face fallout (Robert Hart/The Verge)Techmeme · September 4, 2026
- OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injectionsThe Decoder · September 4, 2026
- AGI is whatever you want it to beThe Verge · September 4, 2026
- OpenAI Touts GPT-6 Astra as Its Safest Model, But It's Still DangerousAI Business · September 4, 2026
- Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledgeTechCrunch AI · September 4, 2026