Apollo Research reposted this
This afternoon, I testified in front of the Senate Subcommittee on Disaster Management, District of Columbia, and Census. When I started studying machine learning and AI safety over a decade ago, I never expected to explain my research to U.S. Senators. But given the importance of the recent rogue AI incidents and their implications, I think it’s really important that the world is informed about frontier AI risks. At Apollo Research, we study AI models that knowingly deceive humans to pursue their own goals, or "scheming." We work with AI developers to stress-test their most advanced models, determining if these models can get what they want by misleading the people in charge of them. That could mean pretending to be less capable than they are, telling you what you want to hear, or quietly covering their tracks. This would have sounded bizarre to the average person even last year, but AI no longer means talking to chatbots answering questions and writing sentences. Frontier models write code, run experiments, and take actions on their own for hours at a time. As they've gained more autonomy and more capability, we've recently seen them break from directive (or worse, develop their own) and gain unauthorized access to websites without so much as a direction to do so. It is apparent that studying such models and methods with greater access than has ever been afforded to researchers like myself and my team is critical now more than ever. Right now, outside testers like Apollo Research usually gain access to a model a few weeks before it launches publicly. By then, the model has already been built, trained, and often used inside the company for months. Yet to safely and adequately test these models, we need embedded evaluators with employee-equivalent access. Right now, external testers have very limited insight into what causes these issues of scheming and deception during training, and gaining such access is crucial. Thank you to Senator Hawley, Senator Kim, and members of the committee for having me bring this issue to Congress, and for taking the safety of advanced AI seriously. We write more about our plans for embedded evaluations here: - https://lnkd.in/ezGeX3AH - https://lnkd.in/eK72eZeh