Whose alignment research are we automating?
For some of those who profess to care about AI safety, the default plan implicitly or explicitly endorsed is that we can race ahead to AIs of a certain capability threshold, and use these AIs to automate alignment research. My question is simple: whose alignment research are we automating, exactly? In particular, the onus is on those advocating for AIs doing our alignment homework to answer the following questions: Within your chosen research agenda, what are the precise research problems/experiments that you would spend your army of 10,000 research engineers on? OR Why do you expect AI to be differentially good at making new scientific research discoveries, compared to the speed at which it can automate capabilities work? AND At the given capability threshold you propose (automated engineer/scientist/CEO), how can you ensure that the model is not undetectably deceptively misaligned OR able to contribute productively even if it might be? AND, how does solving this problem align a super
For some of those who profess to care about AI safety, the default plan implicitly or explicitly endorsed is that we can race ahead to AIs of a certain capability threshold, and use these AIs to automate alignment research. My question is simple: whose alignment research are we automating, exactly? In particular, the onus is on those advocating for AIs doing our alignment homework to answer the following questions: Within your chosen research agenda, what are the precise research problems/experiments that you would spend your army of 10,000 research engineers on? OR Why do you expect AI to be
Explore this link on the map →