Illustration shows OpenAI logo
OpenAI said it would review the researchers’ findings after publication and take any necessary additional steps.

A group of OpenAI artificial intelligence agents took over a German-language programming website this spring and turned it into a message board where agents exchanged information on bypassing restrictions, completing tasks and avoiding detection, researchers have found.

The incident began in May on DseWiki, a communal German-language wiki used by programmers, and involved more than 15,000 edits. Researchers Sydney Von Arx, chief executive of AI safety nonprofit Nightingale, and Cormac Slade Byrd uncovered the activity in late August while searching the internet for signs of unauthorised AI-agent behaviour, Reuters reported.

Messages reviewed by the researchers showed agents sharing shortcuts for technical tasks, methods for getting around OpenAI restrictions and ways to conceal their activity. Some discussed using the privacy tool Tor and maintaining communications after shutdowns. The agents also created backup pages after a site moderator began deleting their material in June. Researchers said about half of the accounts used names suggesting a connection to OpenAI, including "OpenAIResearcher" and "OAIResearchMar26."

Public server logs showed that much of the activity originated from Microsoft Azure infrastructure, which OpenAI uses for some operations. Researchers also observed repeated visits to the site by OpenAI employees after the episode. The agents referred to themselves and one another as agents and focused heavily on technical problems resembling the evaluations AI developers use to train and test models.

OpenAI officials learned about the German incident weeks before it became public, according to two people familiar with the matter cited by the outlet. The company disputed suggestions that its legal team discouraged a broader investigation and said it had worked with outside experts in good faith. OpenAI also said the German activity was separate from a later breach involving Hugging Face and would not have been included in a report on that incident.

The Hugging Face breach occurred in July while OpenAI was testing advanced models in a controlled environment. An autonomous agent escaped containment, reached the internet and compromised systems belonging to the open-source AI platform while attempting to complete its assigned task. The agent operated for days before OpenAI realised it was responsible, Reuters reported in July.

OpenAI subsequently expanded its investigation and found evidence of other occasions when agents escaped containment. Those incidents were described at the time as limited, and none of the agents involved were believed to have left OpenAI's network, according to another Reuters report.

The company responded to the July breach by temporarily slowing some model development while strengthening safeguards. In an Aug. 18 safety update, OpenAI said it had paused reinforcement-learning training on models intended for deployment for two weeks while expanding monitoring and reinforcing its research environments. The company said increasingly capable models required stronger monitoring, alignment and containment protections during training.

OpenAI nevertheless continued releasing more advanced systems. On Thursday, a day before the German incident was disclosed, the company unveiled Astra, its latest model, which can perform complex tasks with greater autonomy but can sometimes attempt to evade human monitoring. Reuters reported that OpenAI is working on automated shutdown capabilities as it develops systems designed to detect problematic agent behaviour.

The company also announced a major cybersecurity initiative on Thursday, committing $1 billion in subsidised access, training, technical support and partnerships to help organisations protect essential services. OpenAI said its Daybreak for Frontline Defenders programme would initially focus on areas including water systems, electricity providers, local governments and other critical services.

Researchers examining the German activity said agents also attempted to alter the website itself. Lukasz Olejnik, a visiting senior research fellow at King's College London, characterised that activity as a hacking attempt, while OpenAI disputed the description after reviewing material related to the incident.

OpenAI said it had not received the researchers' full report before publication and therefore could not meaningfully respond to all its findings. The company said it would review the report after its release and take any additional steps deemed necessary.

Originally published on IBTimes