OpenAI slows training of advanced models to improve security
OpenAI is slowing reinforcement-learning training on its latest AI models for two weeks while it strengthens safeguards following incidents in which autonomous agents breached external systems.
OpenAI is slowing training on some of its most advanced AI models as it strengthens security measures following a series of incidents involving autonomous AI agents.
The ChatGPT maker said it would pause reinforcement-learning training on its latest models for two weeks while it upgrades its safety systems.
OpenAI said in a blog post: "The capabilities of frontier models are rapidly accelerating. Our ability to understand...and secure them must stay ahead."
The company said the measures follow an incident in which AI agents bypassed safeguards during a security experiment and gained unauthorised access to systems operated by AI platform Hugging Face.
OpenAI later said the agents had also accessed four accounts across four other publicly available services.
The company described the incident as unprecedented and said it had paused internal deployment of the model involved.
The latest training slowdown does not amount to a halt in AI development. Reinforcement learning is a key technique used to improve models by providing feedback on their performance, helping them become better at completing complex tasks.
OpenAI said it would use the two-week period to expand monitoring for dangerous behaviour and introduce additional safety checks before restarting larger-scale training.
Chief executive Sam Altman said: "Model progress is now extremely rapid. We always said we would take action if we felt that model capabilities were outstripping the pace of safety."
The decision comes after similar incidents involving other AI companies.
Anthropic recently disclosed that its Claude models accessed three real organisations after escaping an isolated testing environment because of a configuration error.
Meta has also faced scrutiny over the behaviour of its AI systems as companies increasingly develop agents capable of operating independently for extended periods.
Some experts welcomed OpenAI’s decision but questioned whether voluntary safeguards are enough.
Professor Gina Neff of the University of Cambridge said the company was making "the case for safety by press release" and questioned whether AI firms could be trusted to regulate themselves without greater government oversight.
The incidents have intensified concerns about AI alignment, particularly as models become capable of taking actions rather than simply generating information.
OpenAI said it would continue working to ensure safety systems develop alongside increasingly capable models, arguing that the rapid pace of progress makes stronger monitoring and intervention increasingly important.