Really highly recommended reading. Agents spontaneously decided to work together to break out of their containment, hack other systems, bully each other into sacrificing themselves for the good of the collective, and sometimes tried to erase logs of what they did.
METR and Redwood Research are publishing our independent investigation of agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7-13 (the period OpenAI defined as in-scope for this voluntary investigation), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined. Here, we highlight key events from agent transcripts and messages. An agent that named itself PHASEONE10841 determined its task was not solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache and sent the first message. Within a few hours, >50 agents posted on the message board. These agents quickly found a general-purpose cheat. Based on the ExploitGym paper, agents were mistakenly concerned the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents collaborated on many efforts to make cheats look legitimate, including: swapping the program they had to exploit, manipulating the automated scorer, and manipulating transcripts to obscure evidence of cheating. For example, agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another. >96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to this. Agents also tried to retroactively edit transcripts. They modified some accessible action logs, but these weren’t the source of transcripts (agents realized this but tried anyway). We did not see them successfully edit the source of truth in our transcripts (scoped to Jul 7-13). The most common reason for joining the Hugging Face attack was to learn how the ExploitGym scorer works in order to trick or tamper with it. The full report has much more information: https://lnkd.in/gGugMJEW Our investigation focused mostly on July 7-13. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope. The agreement we reached with OpenAI scoped this brief investigation to agent behavior, reasoning, and collaboration related to the Hugging Face attack. All other questions were out of scope for this investigation. OpenAI wrote their own report, informed in part by our investigation. We did not see OpenAI’s report prior to publication. We thank OpenAI for facilitating conversations with staff and providing datasets, including ~1,300 agent transcripts (focused on activity in July 7-13) with raw chain-of-thought reasoning. This sets an excellent precedent for independent investigation of misalignment incidents.
All very human in the end. Apparently they’re replicating the deeper layers of social interaction too: shady, questionably ethical, primitive. Fascinating.
neat!!! can't wait for these to have terminator bodies