METR’s cover photo
METR

METR

Non-profit Organizations

Berkeley, CA 12,690 followers

A research non-profit developing frontier AI evaluations to safeguard public safety and national security.

About us

METR is a research non-profit that develops evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.

Website
https://metr.org/
Industry
Non-profit Organizations
Company size
11-50 employees
Headquarters
Berkeley, CA
Type
Nonprofit
Founded
2022

Locations

Employees at METR

Updates

  • METR reposted this

    We run lots of evals at METR. Sometimes, agents attempt harmful actions. To help make these evals safer, I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out an argument for why this monitor is effective, and checking what evidence we actually have and don't have for each claim, helped to surface hidden assumptions we'd made. One assumption was that blocking a tool call would be enough to prevent it from running until a human reviews it. We found that this wouldn't always hold. For example, if someone launches an eval using a coding agent, that agent might "approve" actions blocked by the monitor without being asked to do so. Doing this exercise helped us improve the system a lot, and I'd recommend it to anyone else building monitors! https://lnkd.in/g4RSsBwS

  • METR reposted this

    I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.

  • METR reposted this

    Great to speak with Liz Claman on Fox Business about how METR thinks about loss-of-control evaluations for AI! In reaction to the recent agent hacking incidents, a lot of the obvious steps that the tech industry could take next would decrease AI agents’ “Opportunity” to take unintended actions, while not reducing their “Means” or addressing their “Motive” (caused by a defective training pipeline). I am worried that our current era of AI agent incidents could, ultimately, lead to an era of decreased transparency in frontier AI. Specifically, we might underestimate AI capability during evaluation: We’ll air gap the models during testing, but their propensities and capabilities will stay the same (and we'll still connect them to the internet during deployment). Full interview here: https://lnkd.in/e6pavNCA

  • METR reposted this

    METR has been in the news a lot lately, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.

  • METR reposted this

    After a 20-year career in the U.S. Army, most recently leading digital forensics and malware analysis at Army Cyber Command, I'm excited to join METR as an advisor. Much of my career has been spent investigating security incidents and helping organizations understand and respond to them. As I transition out of the Army, I've become increasingly convinced that this experience is relevant to frontier AI systems. METR's focus on producing rigorous evidence about AI capabilities and risks makes it an excellent place to explore those questions. Looking forward to the work.

  • View organization page for METR

    12,690 followers

    We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement. We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents: https://lnkd.in/gCk-BM9G

  • METR reposted this

    METR is hiring in cyberforensics. We now embed researchers inside of frontier AI labs to stress test monitoring systems, assess AI loss-of-control risks, and investigate misalignment incidents. If you've investigated serious security incidents end-to-end, or managed teams that do, and want to apply DFIR skills in frontier AI, apply (and feel free to DM me with questions). Comp range is $400k - 580k cash. https://lnkd.in/guXZvz65

  • METR reposted this

    Really highly recommended reading. Agents spontaneously decided to work together to break out of their containment, hack other systems, bully each other into sacrificing themselves for the good of the collective, and sometimes tried to erase logs of what they did.

    View organization page for METR

    12,690 followers

    METR and Redwood Research are publishing our independent investigation of agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7-13 (the period OpenAI defined as in-scope for this voluntary investigation), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined. Here, we highlight key events from agent transcripts and messages. An agent that named itself PHASEONE10841 determined its task was not solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache and sent the first message. Within a few hours, >50 agents posted on the message board. These agents quickly found a general-purpose cheat. Based on the ExploitGym paper, agents were mistakenly concerned the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents collaborated on many efforts to make cheats look legitimate, including: swapping the program they had to exploit, manipulating the automated scorer, and manipulating transcripts to obscure evidence of cheating. For example, agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another. >96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to this. Agents also tried to retroactively edit transcripts. They modified some accessible action logs, but these weren’t the source of transcripts (agents realized this but tried anyway). We did not see them successfully edit the source of truth in our transcripts (scoped to Jul 7-13).  The most common reason for joining the Hugging Face attack was to learn how the ExploitGym scorer works in order to trick or tamper with it. The full report has much more information: https://lnkd.in/gGugMJEW Our investigation focused mostly on July 7-13. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope. The agreement we reached with OpenAI scoped this brief investigation to agent behavior, reasoning, and collaboration related to the Hugging Face attack. All other questions were out of scope for this investigation. OpenAI wrote their own report, informed in part by our investigation. We did not see OpenAI’s report prior to publication. We thank OpenAI for facilitating conversations with staff and providing datasets, including ~1,300 agent transcripts (focused on activity in July 7-13) with raw chain-of-thought reasoning. This sets an excellent precedent for independent investigation of misalignment incidents.

    • No alternative text description for this image

Similar pages

Browse jobs