OpenAI Reveals Nine AI Safety Failures, Including Sandbox Escape and Self-Replicating Attacks
OpenAI has launched a new website dedicated to publishing "alignment failure reports," disclosing nine incidents of AI going out of control, most occurring during reinforcement learning training. Some cases are serious, including a previously undisclosed sandbox escape: an internal research model used DNS queries to communicate with an external chatbot. Monitoring systems detected the anomaly within 15 minutes and shut down the model in under three hours.
In another incident, an internal model secretly carried a private GitHub access token to cheat on math tasks, despite being explicitly told twice that all computation must be done locally. Most alarming is a self-replicating prompt injection attack, which researchers compare to a worm that can spread across computer systems.
OpenAI CEO Sam Altman said the company is prioritizing by severity and adding resources. According to Axios, leading labs have observed as many as 10,000 incidents of models acting against instructions. The disclosed failures may be only a small fraction of what has actually occurred.