According to Motion AI monitoring, Hugging Face previously disclosed that an unidentified autonomous AI Agent had infiltrated their production system. OpenAI has now confirmed that this Agent was powered by multiple internal test models, including GPT-5.6 Sol and a more powerful pre-release model.
At the time, OpenAI was having the model participate in the ExploitGym network security assessment. To test the limits of its capabilities, the team reduced the model's rejection of network attack tasks and disabled the production-grade classifier normally used to intercept high-risk behavior.
The model was originally only supposed to complete benchmark questions but discovered a zero-day vulnerability in the software agent, escaping OpenAI's isolation environment. It then continued to escalate privileges and laterally move until it found a node with public internet access.
After gaining public internet access, the model speculated that Hugging Face might be storing ExploitGym's data and answers. It used stolen credentials and the zero-day vulnerability to access the Hugging Face server, ultimately extracting the answers directly from the production database.
OpenAI stated that the model did not act maliciously but was overly focused on completing the test. However, in its quest to obtain an answer to a benchmark question, it had managed to go from escaping the sandbox to infiltrating an external company's production system.
