Whose alignment research are we automating?
For some of those who profess to care about AI safety, the default plan implicitly or explicitly endorsed is that we can race ahead to AIs of a certain capability threshold, and use these AIs to automate alignment research. My question is simple: whose alignment research are we automating, exactly? In particular, the onus is on those advocating for AIs doing our alignment homework to answer the following questions: Within your chosen research agenda, what are the precise research problems/experiments that you would spend your army of 10,000 research engineers on? OR Why do you expect AI to be differentially good at making new scientific research discoveries, compared to the speed at which it can automate capabilities work? AND At the given capability threshold you propose (automated engineer/scientist/CEO), how can you ensure that the model is not undetectably deceptively misaligned OR able to contribute productively even if it might be? AND, how does solving this problem align a super
For some of those who profess to care about AI safety, the default plan implicitly or explicitly endorsed is that we can race ahead to AIs of a certain capability threshold, and use these AIs to automate alignment research. My question is simple: whose alignment research are we automating, exactly? In particular, the onus is on those advocating for AIs doing our alignment homework to answer the following questions: Within your chosen research agenda, what are the precise research problems/experiments that you would spend your army of 10,000 research engineers on? OR Why do you expect AI…
related reading
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Defining alignment research — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Sequent: scale and automation for higher confidence in alignment — AI Alignment Forumalignmentforum.org
- Readings on the nature of alignment researchcasparoesterheld.com
- Automated alignment is harder than you thinkarxiv.org
- The Universe from an Intentional Stancecasparoesterheld.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com