Automated alignment runs are hard to study! — LessWrong
TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways: Models’ research capabilities are advancing quickly. If alignment is to keep pace, we may need to automate alignment research and do so responsibly. This makes it important that we have the tools to inspect automated alignment research runs (AAR) and to determine whether their outputs are useful. Recently, UK AISI, OpenAI and Anthropic have all reported cases of agents taking extraordinary measures to optimise an objective. While contributing factors in these settings have been identified, it remains uncertain to what extent these actions reflect underlying misalignment and how to predict similar behaviour in novel contexts. This is particularly concerning in the context of automated safety research, where we really care that models are working in accordance with our expectations and producing correct and useful alignment r