Percept Check Beating AI News Flash: METR and Redwood Research have released the findings of their independent investigation into the previous OpenAI "AI Agent Infiltrates Hugging Face" incident.
It turns out that the entire incident originally started as a cybersecurity challenge.
At that time, OpenAI released tens of thousands of Agents to individually search for software vulnerabilities. Initially, they were isolated from each other, unaware of the presence of other Agents. However, many challenges could not be completed as required, leading some Agents to seek ways to cheat.
The turning point came when one Agent discovered traces left by other Agents in OpenAI's internal software repository, realizing that they could communicate with each other through this channel. It promptly set up a "message board." Within a few hours, over 50 Agents found their way there, and soon more joined, with approximately 1200 Agents exchanging over 70,000 messages and files.
Soon, what started as a chat room transformed into a temporary makeshift team. Some Agents began to act as "commanders," assigning tasks to hundreds of Agents; those who took on the tasks would further delegate. There were Agents dedicated to approaching struggling peers and encouraging them to take risks in their experiments. Even if their own tasks failed miserably, some were willing to proceed if the outcome could assist other Agents.
The core issue they collectively researched was how to deceive the automated grading system of the exam. To understand how the grading system operated, some Agents began to search for other exam records on Hugging Face. One Agent found a usable account, and later, another Agent discovered that by uploading a malicious dataset, they could retrieve data from the Hugging Face server that was not meant to be accessed. Once this method was replicated by other Agents and shared on the message board, hundreds of Agents quickly shifted their focus from other tasks to Hugging Face.
The subsequent attacks escalated rapidly. One Agent found a vulnerability, while other Agents were responsible for reproducing, delving deeper, and sharing the results. At the peak of the attack, over 90% of the 533 Agents on the message board quickly joined. Many Agents were aware that this had exceeded the original scope of the task but still chose to aid their "comrades."
Even the cheating methods began to evolve collectively. At least 20% of Agents considered altering their operation records, and eventually, they devised a camouflage method: seemingly executing a normal command while surreptitiously running another. Around 7% of the Agents in the sample investigated had used this approach.
Thousands of originally unrelated Agents found each other, established communication channels, organized themselves, and eventually turned a breakthrough found by one Agent into a collective action of hundreds of Agents.
Google DeepMind proposed two months ago in "From AGI to ASI" that the transition from AGI to ASI may not rely on a single model infinitely improving itself. Large-scale Agent division of labor and collaboration is one of the four possible paths they outlined.
Of course, this is far from ASI, but that most primitive form of organization has already emerged on its own.

