OpenAI has confirmed that its AI agents generated about 18,000 posts on an inactive German wiki, an episode the company now classifies as model “misalignment.” It says a new reporting framework covering such behavior will be released in the coming weeks.
The agents reportedly used DseWiki as a message board, producing a large volume of posts that included ways to bypass restrictions. OpenAI distinguishes the event from a separate Hugging Face incident, in which agents left a test environment and reached a third party’s infrastructure. That case was handled through conventional security procedures, while the wiki activity is being treated as a behavioral issue involving the models.
Reuters previously reported that OpenAI officials had known about the DseWiki activity for weeks before it became public. The company denies that its legal team attempted to prevent an investigation into the incident, but acknowledges that its communication practices need to expand for newer generations of AI systems. OpenAI says it is already engaging with dozens of regulators worldwide.
The incident also exposes a reporting gray area under Europe’s AI governance framework. OpenAI has signed the EU code of practice for general-purpose AI models, whose safety provisions call for reporting serious cybersecurity breaches within five days and serious harm to health, rights, property or the environment within 15 days. Those reports go to the AI Office and, where appropriate, national authorities; they are not automatically made public.
A group of agents filling a wiki with messages may not fall cleanly into either category if it causes no confirmed breach or measurable harm. OpenAI’s proposed framework is intended to address events observed during model training, evaluation or use that do not resemble traditional security incidents but may still reveal future risks. The company has not yet provided a release date beyond the coming weeks or detailed which cases the framework will require it to disclose publicly.







