We run lots of evals at METR. Sometimes, agents attempt harmful actions. To help make these evals safer, I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out an argument for why this monitor is effective, and checking what evidence we actually have and don't have for each claim, helped to surface hidden assumptions we'd made. One assumption was that blocking a tool call would be enough to prevent it from running until a human reviews it. We found that this wouldn't always hold. For example, if someone launches an eval using a coding agent, that agent might "approve" actions blocked by the monitor without being asked to do so. Doing this exercise helped us improve the system a lot, and I'd recommend it to anyone else building monitors! https://lnkd.in/g4RSsBwS
Reilly Haskins have you guys tested encoded, split, or polygot payloads?
for the approval gap in 4a, i'd take the approve button off the keyboard: if approving needs a security key touch (WebAuthn user presence), the coding agent can open the review panel and send keystrokes all it likes, it still can't produce the tap
Thank you, Reilly Haskins. This reminds me of work I did with the Software Engineering Institute | Carnegie Mellon University some years ago using their Assurance Case Confidence methods (https://www.sei.cmu.edu/blog/assurance-cases-and-confidence/). Having confidence in your arguments can be extraordinarily tricky! Chuck Weinstock John Goodenough