flâneur — a map of the web's best reading

Measurement Research Agenda – Center on Long-Term Risk

longtermrisk.org · 2,945 words · saved by 1 readers

Author: Mia Taylor Contents 1 Motivation The Center on Long-Term Risk aims to reduce risks of astronomical suffering (s-risk) from advanced AI systems. We’re primarily concerned with threat models involving the deliberate creation of suffering during conflict between advanced agentic AI systems. To mitigate these risks, we are interested in tracking properties of AI systems that make them more likely to be involved in catastrophic conflict. Thus, we propose the following research priorities: Identify and describe properties of AI systems that would robustly make them more likely to contribute to s-risk (section 2.1) Design measurement methods to detect whether systems have these properties (section 2.2) Use these measurements on contemporary systems to learn what aspects of training, prompting, or scaffolding […]

Measurement Research Agenda Author: Mia Taylor Contents 1 Motivation 2 Research areas 2.1 Characterizing s-risk-conducive properties 2.2 Developing measurements for these properties 2.3 Characterizing the effects of interventions on properties of interest Box 1: Influencing the behavior of misaligned models Box 2: Bargaining policy generalization 3 Theory of change 3.1 Product model 3.2 Field-building model Acknowledgements 1 Motivation The Center on Long-Term Risk aims to reduce risks of astronomical suffering (s-risk) from advanced AI systems. We’re primarily concerned with threat models inv

Explore this link on the map →

related reading