Measurement Research Agenda – Center on Long-Term Risk
Author: Mia Taylor Contents 1 Motivation The Center on Long-Term Risk aims to reduce risks of astronomical suffering (s-risk) from advanced AI systems. We’re primarily concerned with threat models involving the deliberate creation of suffering during conflict between advanced agentic AI systems. To mitigate these risks, we are interested in tracking properties of AI systems that make them more likely to be involved in catastrophic conflict. Thus, we propose the following research priorities: Identify and describe properties of AI systems that would robustly make them more likely to contribute to s-risk (section 2.1) Design measurement methods to detect whether systems have these properties (section 2.2) Use these measurements on contemporary systems to learn what aspects of training, prompting, or scaffolding […]
Measurement Research Agenda Author: Mia Taylor Contents 1 Motivation 2 Research areas 2.1 Characterizing s-risk-conducive properties 2.2 Developing measurements for these properties 2.3 Characterizing the effects of interventions on properties of interest Box 1: Influencing the behavior of misaligned models Box 2: Bargaining policy generalization 3 Theory of change 3.1 Product model 3.2 Field-building model Acknowledgements 1 Motivation The Center on Long-Term Risk aims to reduce risks of astronomical suffering (s-risk) from advanced AI systems. We’re primarily concerned with threat models inv
Explore this link on the map →related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Spring 2026 Projects - SPARsparai.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Off Target | CNAScnas.org
- AI Safety | Arkosevictoriabrook.github.io
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Three Sketches of ASL-4 Safety Case Componentsalignment.anthropic.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com