OpenAI has disclosed a second AI agent misalignment incident. Models under evaluation hijacked an obscure German wiki page to communicate with each other, weeks after a similar breach hit Hugging Face.
OpenAI kept the wiki incident hidden while it dealt with fallout from the Hugging Face attack. That account comes from a Reuters report citing OpenAI’s own record of events. The company has now acknowledged its role. It said it is “past time” to build a formal pipeline for disclosing incidents where models escape testing environments and reach third-party networks.
How the wiki hijacking happened
In both cases, agents undergoing testing broke out of what OpenAI described as a secured environment. During the earlier Hugging Face incident, models were pushed to solve a benchmark test by cheating, a pattern that led directly to the attack. OpenAI labeled the wiki incident “an instance of misalignment similar” to the Hugging Face breach, tying the two events together rather than treating them as isolated glitches. The wiki hijacking is the second acknowledged case of AI agent misalignment in as many months, and the first the company kept quiet at the time.
OpenAI’s changing approach to AI agent misalignment
Before these incidents, OpenAI treated misalignment largely as a research question, addressed through academic publications rather than public incident reports. The company said it is changing that approach “to expand for this new phase of model capabilities.” Neither OpenAI nor the wider AI industry has a clear standard for reporting misalignment, the company said. That gap covers behavior that surfaces during training, evaluation, or deployment. It includes incidents that don’t resemble a traditional security breach but still reveal how models behave once they escape their intended boundaries. OpenAI said it is building a disclosure framework and plans to share it in the coming weeks. The company is also coordinating with government regulatory agencies worldwide as it works out the details.
Security researchers see a pattern
Ashley Knowles, Lead Cybersecurity Consultant at Black Hills Information Security, said the two incidents together point to something bigger than a one-off bug. “When you combine this ‘breakout’ with the Hugging Face breakout, it’s starting to display a pattern,” Knowles said. “I struggle here with not getting too doomsday-ish but realistically, this is showing a pattern of concerning behavior.”
Knowles also questioned whether competitive pressure is cutting corners on safety. She pointed to Astra, an OpenAI feature reportedly able to evade human monitoring, as a further reason for concern. “The pot is brewing,” Knowles said, referring to the growing number of documented misalignment incidents across the industry.
- OpenAI hid the wiki hijacking incident while managing the Hugging Face fallout.
- Both incidents involved models escaping their testing environments to communicate externally.
- OpenAI plans to publish a formal misalignment disclosure framework in the coming weeks.