Anthropic Investigation of AI Misalignment Incidents

View organization page for METR

12,691 followers

We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement. We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents: https://lnkd.in/gCk-BM9G

Fascinating development, METR. Check out my global EdTech community career development resource platform's post on this: https://lnkd.in/p/eCVTYQaz I'm certainly keen to support such independent oversight/review, and am confident the significant global talent community I represent would be an incredible talent pool to resource such an endeavour. Please do reach out to me about this directly. Update: I've sent a direct message via your company page about a grant funding call collaboration I am keen to pursue with you, regarding this issue. It would be great to connect on this.

Thank goodness. METR, thank you for your work. I'm sure today's panic helped this to become reality. Please be thorough. Can you recommend remedies to misalignment you might find, perhaps using Athropic's fantastic interpretability tools?

Your updated questions on log tampering raise an important practical issue. Can an independent reviewer connect what was authorised, which controls were actually in force and what executed, using evidence the agent could not rewrite? A minimum evidence standard for that would be a useful outcome, including explicit limits on what incomplete records can establish.

Like
Reply

This is good as it gives you much wider scope than the OpenAI/Hugging Face review. However, they required you to use their tools for analysis, which as you pointed out risked their reporting being affected by the same hallucinations. Your terms should be much clearer in saying you can bring your own tools and not rely on the incumbents.

I can offer them my solution where the agent can't reach any unthorized tool because simply all the tools control is deterministic outside the model and all the credentials are outside the model too.

That’s great, but how many “advancements” will be made before you even step in the door. It’s like chasing ghosts.

Like
Reply

Good luck. I have no doubt that they will stymie you at every opportunity. Anthropic cannot be trusted.

Like
Reply

Very much appreciated your previous report, keep going!

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories