AI Oversight + Control — ML Alignment & Theory Scholars
matsprogram.org · saved by 1 readers
As model develop potential dangerous behaviors, can we develop and evaluate methods to monitor and regulate AI systems, ensuring they adhere to desired behaviors while minimally undermining their efficiency or performance?