METR Eval Monitor Improves Safety with Human Review

We run lots of evals at METR. Sometimes, agents attempt harmful actions. To help make these evals safer, I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out an argument for why this monitor is effective, and checking what evidence we actually have and don't have for each claim, helped to surface hidden assumptions we'd made. One assumption was that blocking a tool call would be enough to prevent it from running until a human reviews it. We found that this wouldn't always hold. For example, if someone launches an eval using a coding agent, that agent might "approve" actions blocked by the monitor without being asked to do so. Doing this exercise helped us improve the system a lot, and I'd recommend it to anyone else building monitors! https://lnkd.in/g4RSsBwS

Thank you, Reilly Haskins. This reminds me of work I did with the Software Engineering Institute | Carnegie Mellon University some years ago using their Assurance Case Confidence methods (https://www.sei.cmu.edu/blog/assurance-cases-and-confidence/). Having confidence in your arguments can be extraordinarily tricky! Chuck Weinstock John Goodenough

Like
Reply

Reilly Haskins have you guys tested encoded, split, or polygot payloads?

Like
Reply

for the approval gap in 4a, i'd take the approve button off the keyboard: if approving needs a security key touch (WebAuthn user presence), the coding agent can open the review panel and send keystrokes all it likes, it still can't produce the tap

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories