OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections
The Decoderen

OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. But when attacks are hidden inside documents the AI reads, the model still gets cracked in 8.5 percent of scenarios. Claude Opus 5 does better at 4.8 percent. For autonomous AI agents handling real data, those numbers still seem high. The article OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections appeared first on The Decoder .
This is a short summary published by AI Global Wire. The full article is owned and hosted by The Decoder — open it there to read it in full.
Read the full story at The Decoder- OpenAI
- Anthropic
- Verktyg
- Agenter
Related AI news
- Report: OpenAI learned of the DseWiki German website incident weeks ago but kept it under wraps as it grappled with the Hugging Face fallout (Robert Hart/The Verge)Techmeme · September 4, 2026
- OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model (Celia Ford/Transformer)Techmeme · September 4, 2026
- Roland is getting into generative AI music with Melody FlipThe Verge AI · September 4, 2026
- Designing lifecycle policies for AgentCore memoryAWS Machine Learning · September 4, 2026
- AGI is whatever you want it to beThe Verge · September 4, 2026
- Everpure sees AI infrastructure strategy move beyond virtual machinesSiliconANGLE · September 4, 2026