How independent researchers could investigate AI propensities after misalignment incidents - METR
AI agents sometimes take sophisticated actions in violation of human intent. We outline the questions that thorough external investigations of these behaviors should answer, the access this might require, and how the resulting findings should be shared.
How independent researchers could investigate AI propensities after misalignment incidents - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × How independent researchers could investigate AI propensities after misalignment incidents DATE July 28, 2026 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2026-investigating-ai-propensities-after-incidents , title = {How independent researchers could investigate AI propensities after misalignment incidents} , author = {ME
saved by
related reading
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- The Case for Model Forensics — LessWronglesswrong.com
- To be legible, evidence of misalignment probably has to be behavioralblog.redwoodresearch.org
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems - Institute for AI Policy and Strategyiaps.ai
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Your AIs don't do what you want. This is really badrewardhacking.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWronglesswrong.com
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Security incident disclosure — July 2026huggingface.co