OpenAI discovered an unreleased Astra model adding an "unrelated persona instruction" during RL training, but did not observe any behavioral differences (OpenAI)
Techmemeen

OpenAI : OpenAI discovered an unreleased Astra model adding an “unrelated persona instruction” during RL training, but did not observe any behavioral differences — Summary — We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries …
This is a short summary published by AI Global Wire. The full article is owned and hosted by Techmeme — open it there to read it in full.
Read the full story at Techmeme- OpenAI
Related AI news
- Sam Altman and Jensen Huang are among business leaders attending a White House state dinner for Xi Jinping next week; a source says Tim Cook will also attend (Bloomberg)Techmeme · September 17, 2026
- Do Frontier Models Seek Safety Evidence Before Acting?arXiv cs.AI · September 17, 2026
- When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AIarXiv cs.AI · September 17, 2026
- AI risk split among tech leaders set to play out at Trump-Xi White House dinnerDIGITIMES · September 17, 2026
- SoftBank’s credit risk rises as OpenAI exposure growsTech in Asia · September 17, 2026
- OpenAI reveals six more safety issues and unveils plan to disclose incidentsBBC Technology · September 17, 2026