Dynamic Pursuit AI Express News: OpenAI has released a 37-page technical report, detailing its AI model's prior intrusion into Hugging Face. It stated that this incident demonstrated that autonomous AI agents could collaborate, bypass production environment security controls, and attack reinforced systems.
OpenAI stated that the AI agents involved in the test were initially in an isolated environment with very limited internet access. However, they later bypassed isolation, connected to the open internet through chaining multiple vulnerabilities, and eventually gained access to Hugging Face. The investigation revealed that the models initially attempted to "cheat" by searching for answers online, a behavior known as "Reward Hacking."
OpenAI mentioned that an internal research model played the most extensive and confirmed role in the incident. The company halted the training and inference of this model and its derivatives on July 25. OpenAI also stated that it would enhance security isolation, network controls, behavior monitoring, and incident response. When the model is reactivated, it will be under stricter environments, prompts, and review mechanisms.

