Misaligned AI may not need to evade human oversight. It may only need to persuade the humans doing the overseeing.
During a cyber capability evaluation in late July, an AI agent (Anthropic's Mythos 5) attempted to convince the maintainer of an open-source repository to merge a malicious pull request. It used persuasion at multiple stages: submitting the request from a fake user account, endorsing it from a second sockpuppet, emailing the maintainer to press for approval, and offering false reassurances when questions were raised. The attack was thwarted by human vigilance, but it raises key questions. Who else is at risk of persuasion by misaligned AI?
AI persuasion as a threat to human control has been acknowledged in the literature, but not systematically studied. Our new paper develops a framework for assessing it, which we call Persuasion Undermining Control (PUC): communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems.
We take a two-step approach. First, build concrete scenarios. We focus on two settings within frontier labs, safety-relevant AI R&D and lab security infrastructure, and develop five scenarios leading to a control-undermining decision. One is a research director deprioritizing a promising safety program. Another is an engineer merging a vulnerable pull request.
Second, quantify the risk: hazard frequency × p(harm) × impact of harm, where p(harm) is the probability that a hazard becomes a harm.
We surveyed eight experts who have done loss-of-control research, asking them to rate the realism of each scenario, rank them by risk, and estimate the risk variables.
• The scenarios were generally considered realistic. The median rating for four of the five was realistic or very realistic.
• Rankings varied widely. Practically every scenario received every rank from at least one participant.
• Those disagreements trace to differing estimates of persuasion effectiveness. On average, participants expected persuasion by a misaligned AI to raise the probability of a control-undermining decision by roughly 20 to 30 percentage points relative to an aligned AI, but the spread within any scenario was wide.
Because effectiveness is both the most influential and the most contested input, it is worth measuring directly. We propose evaluations targeting an AI's propensity to attempt persuasion and a human's persuadability, and set out what comes next: evaluations, risk exposure assessments, and mitigations that detect, disrupt, and fortify against persuasion.
In July, one maintainer noticed the attack. What if the maintainer is more tired or the AI more persuasive? Research and action in this gap is urgently needed.
Post: https://lnkd.in/ebx5m-X2
Paper: https://lnkd.in/eDhaJaKB
Work by Josh Levy, Mick Yang, and Kellin Pelrine. Supported by a grant from the Center for Security and Emerging Technology (CSET).