How independent researchers could investigate AI propensities after misalignment incidents - METR
AI agents sometimes take sophisticated actions in violation of human intent. We outline the questions that thorough external investigations of these behaviors should answer, the access this might require, and how the resulting findings should be shared.
How independent researchers could investigate AI propensities after misalignment incidents - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × How independent researchers could investigate AI propensities after misalignment incidents DATE July 28, 2026 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2026-investigating-ai-propensities-after-incidents , title = {How independent researchers could investigate AI propensities after misalignment incidents} , author = {ME
Explore this link on the map →related reading
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Security incident disclosure — July 2026huggingface.co
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWronglesswrong.com
- The Case for Model Forensics — LessWronglesswrong.com
- Spring 2026 Projects - SPARsparai.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com