<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><description>METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.</description><link>https://bsky.app/profile/metr.org</link><title>@metr.org - METR</title><item><link>https://bsky.app/profile/metr.org/post/3mv4fb4hebs2v</link><description>We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement.</description><pubDate>09 Sep 2026 20:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mv4fb4hebs2v</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mtz7qneynk2l</link><description>METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&amp;D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.</description><pubDate>26 Aug 2026 20:54 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mtz7qneynk2l</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mt3g3kpqhs2s</link><description>In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.</description><pubDate>15 Aug 2026 00:28 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mt3g3kpqhs2s</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mrtdo7tltc2v</link><description>We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.</description><pubDate>30 Jul 2026 01:58 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mrtdo7tltc2v</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mrshe5lils2k</link><description>We believe it&#39;s important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.</description><pubDate>29 Jul 2026 17:31 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mrshe5lils2k</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mr6mq3tmks2a</link><description>Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems.&#xA;&#xA;The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.</description><pubDate>21 Jul 2026 20:14 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mr6mq3tmks2a</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mpa2yvepac2m</link><description>OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon.</description><pubDate>26 Jun 2026 23:12 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mpa2yvepac2m</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mma27nolf223</link><description>Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control.&#xA;&#xA;The result: our first Frontier Risk Report.</description><pubDate>19 May 2026 18:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mma27nolf223</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mllw544jks2n</link><description>We surveyed 349 technical researchers, engineers, and managers (in February–April 2026) about how they use AI tools at work.&#xA;&#xA;On average, participants self-report that AI use made their work 1.6–2.1x more valuable, and that this multiplier will grow over time.</description><pubDate>11 May 2026 18:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mllw544jks2n</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mlha2wv4ls2u</link><description>We reviewed a section of Anthropic’s February 2026 Risk Report focused on automated R&amp;D risk from Opus 4.6. While we take issue with the adequacy of evidence the report provides, we agree with Anthropic about the overall level of risk &amp; remain excited to pilot reviews like these.</description><pubDate>09 May 2026 21:50 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mlha2wv4ls2u</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mlewafi6pc2u</link><description>We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks.</description><pubDate>08 May 2026 23:49 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mlewafi6pc2u</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mj6cntwn642z</link><description>We co-developed MirrorCode with @epochai.bsky.social to test AI on extremely long-horizon blackbox software reimplementation tasks. We found that recent public models are able to fully implement at least some programs we estimate would take humans weeks or months to implement.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>10 Apr 2026 21:52 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mj6cntwn642z</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mj6ckxiqt22z</link><description>We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks.</description><pubDate>10 Apr 2026 21:51 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mj6ckxiqt22z</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mgbojumv5c2g</link><description>We’re correcting a mistake in our modeling that inflated recent 50%-time horizons by 10-20% (and reduced 80%-horizons). We inappropriately penalized steepness in task-length→success curve fits. This most affects the oldest and newest models, whose fits are less data-constrained.</description><pubDate>05 Mar 2026 01:47 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mgbojumv5c2g</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mfmtn5amos2w</link><description>Since early 2025, we&#39;ve been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this.</description><pubDate>24 Feb 2026 18:53 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mfmtn5amos2w</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mfkewb3pek2x</link><description>We estimate that GPT-5.3-Codex with reasoning effort `high` (not `xhigh`) has a 50%-time-horizon of around 6.5 hours (95% CI of 3 hrs to 17 hrs) on our suite of software tasks. OpenAI provided API access for this evaluation.</description><pubDate>23 Feb 2026 19:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mfkewb3pek2x</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mfdebli4ps2a</link><description>We estimate that Claude Opus 4.6 has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks. While this is the highest point estimate we’ve reported, this measurement is extremely noisy because our current task suite is nearly saturated.</description><pubDate>21 Feb 2026 00:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mfdebli4ps2a</guid></item><item><link>https://bsky.app/profile/metr.org/post/3me2x46wtis2q</link><description>We estimate that GPT-5.2 with `high` (not `xhigh`) reasoning effort has a 50%-time-horizon of around 6.6 hrs (95% CI of 3 hr 20 min to 17 hr 30 min) on our expanded suite of software tasks. This is the highest estimate for a time horizon measurement we have reported to date.</description><pubDate>04 Feb 2026 22:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3me2x46wtis2q</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mdxzl36jum26</link><description>We’ve started to measure time horizons for recent models using our updated methodology. On this expanded suite of software tasks, we estimate that Gemini 3 Pro has a 50%-time-horizon of around 4 hrs (95% CI of 2 hr 10 mins to 7 hrs 20 mins).</description><pubDate>03 Feb 2026 18:48 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mdxzl36jum26</guid></item><item><link>https://bsky.app/profile/metr.org/post/3mdxyrljy3s26</link><description>We’re updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons.</description><pubDate>03 Feb 2026 18:34 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3mdxyrljy3s26</guid></item><item><link>https://bsky.app/profile/metr.org/post/3md2krxg6cs2c</link><description>How well can AI-based monitoring detect when an agent is covertly pursuing a side objective? &#xA;&#xA;In early work, we find clear trends: more capable models (in terms of time horizon) are better able to detect covert behavior.</description><pubDate>23 Jan 2026 01:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3md2krxg6cs2c</guid></item><item><link>https://bsky.app/profile/metr.org/post/3lxukjaolck2i</link><description>We estimate that Claude Opus 4.1 has a 50%-time-horizon of around 1 hr 45 min (95% confidence interval of 50 to 195 minutes) on our agentic multi-step software engineering tasks. This estimate is lower than the current highest time-horizon point estimate of around 2 hr 15 min.</description><pubDate>02 Sep 2025 16:38 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3lxukjaolck2i</guid></item><item><link>https://bsky.app/profile/metr.org/post/3lwcvbft6ok23</link><description>We tested how autonomous AI agents perform on real software tasks from our recent developer productivity RCT.&#xA;&#xA;We found a gap between algorithmic scoring and real-world usability that may help explain why AI benchmarks feel disconnected from reality.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>13 Aug 2025 22:37 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3lwcvbft6ok23</guid></item><item><link>https://bsky.app/profile/metr.org/post/3lw7xj7mkac27</link><description>Before publishing our recent developer productivity RCT, we thought hard about how to accurately and clearly communicate our results. In a new blog post, we outline some of our key considerations regarding scientific integrity and communication.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>12 Aug 2025 18:40 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3lw7xj7mkac27</guid></item><item><link>https://bsky.app/profile/metr.org/post/3lw3jpx3dzs2o</link><description>Prior work has found that Chain of Thought (CoT) can be unfaithful. Should we then ignore what it says?&#xA;&#xA;In new research, we find that the CoT is informative about LLM cognition as long as the cognition is complex enough that it can’t be performed in a single forward pass.</description><pubDate>11 Aug 2025 00:22 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3lw3jpx3dzs2o</guid></item><item><link>https://bsky.app/profile/metr.org/post/3lvu3ksk4tk2o</link><description>In a new report, we evaluate whether GPT-5 poses significant catastrophic risks via AI R&amp;D acceleration, rogue replication, or sabotage of AI labs. &#xA;&#xA;We conclude that this seems unlikely. However, capability trends continue rapidly, and models display increasing eval awareness.</description><pubDate>08 Aug 2025 01:20 +0000</pubDate><guid isPermaLink="false">at://did:plc:dll3hepzq76nymel5c3yt6nk/app.bsky.feed.post/3lvu3ksk4tk2o</guid></item></channel></rss>