METR Aims to Ensure Public Awareness of AI Capabilities and Risks

METR has been in the news a lot lately, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.

Aren't you funded by those same doomers?

Just like we needed Underwriter Labs to make sure electrical appliances weren't burning down houses 100 years ago, we need a strong, independent monitor to evaluate frontier AI and share what is going on. If electricity was burning down houses out of control back then, we never would have ended with the current building codes that guide our houses and appliances to safe usage. Hopefully we can find a place with AI like we did with electricity, completely transformative and life saving and safe all at the same time. Thank you for your work in this area Chris.

What you do doesn't qualify as independent investigations. You already have working relationships with these labs and thus human bias. Many of your employees used to work for them. It's unconsciousible that you would be chosen to be an auditor or evaluator of any kind.

External evaluation helps establish what risks exist. Pre-execution authority controls determine what a deployed system is permitted to do. We need both, with independently verifiable evidence at each layer. I’d be interested in learning how METR evaluates whether an AI system can recognize and remain within its delegated authority, particularly when an action is technically possible but not authorized.

Excellent post. METR’s emphasis on independent third-party evaluation and credible evidence about advanced AI behavior addresses an increasingly important need. My research approaches this from a complementary direction. I am developing AI Strategic and Adaptive Behavioral Dynamics (AI-SABD), focused on consequential AI behaviors including deception, concealment, evaluation-conditioned behavior, strategic compliance, behavioral adaptation, and persistent differential treatment. Within AI-SABD, Artificial Intelligence Behavioral Diagnostics (AIBD) is an applied specialty for investigating problematic AI behavior, identifying competing causes through controlled testing, recommending intervention, and verifying whether the problem was corrected. I am also developing SIRE-CGOS, a cognitive-governance architecture designed to constrain consequential AI behavior during uncertainty or diagnosis. I see these areas as complementary: evaluation identifies concerning behavior; behavioral diagnostics investigates why it occurred, whether it persists, and whether intervention corrected it. As AI becomes more autonomous, I believe we will need both. I appreciate METR’s work advancing independent evaluation and transparency.

Great breakdown, Chris. Independence needs to be real, not just a policy statement. Real objectivity means you can stress test AI without worrying about losing the contract or the relationship. The bigger accelerator is turning what you learn into a method other teams can use, not just a one time report. One report helps one company. A shared method helps the whole industry move faster and safer.

Like
Reply

I mean, it's one thing to police the prompts that are being put into your model, it's quite another to have a model so poorly outfitted with safeguards that it will successfully execute these prompts to create issues. One can see this as a microcosm of society in general; we have laws in place that should prevent members of society from taking unwanted actions, but moreover like bribing a police officer, society is set up in a way where nevertheless that behavior can be successfully executed.

Like
Reply

Chris Painter the way you laid out how METR chooses accountability over comfort, even turning down funding rather than softening results, says so much about what real oversight requires. Trust has to be earned in the open, not assumed behind closed doors. Very grateful for what you're all doing and creating at METR. Wishing you a great weekend.

Like
Reply

Whenever I see you guys in the news I feel honored that someone I know is out there doing this very difficult work (e.g. as opposed to nuclear weapons, where there's a physical gap between the kind of uranium needed for weapons development and civilian use, there isn't such a safety margin with AI.)

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories