flâneur — a map of the web's best reading

How can we solve diffuse threats like research sabotage with AI control? — LessWrong

lesswrong.com · 2,659 words · saved by 1 readers

Misaligned AIs might engage in research sabotage: making safety research go poorly by doing things like withholding their best ideas or putting subtl…

x How can we solve diffuse threats like research sabotage with AI control? — LessWrong AI Control AI Frontpage 54 How can we solve diffuse threats like research sabotage with AI control? by Vivek Hebbar 30th Apr 2025 AI Alignment Forum 10 min read 1 54 Ω 34 Misaligned AIs might engage in research sabotage: making safety research go poorly by doing things like withholding their best ideas or putting subtle bugs in experiments. To mitigate this risk with AI control, we need very different techniques than those in the original control paper or Ctrl-Z , because research sabotage differs from those

Explore this link on the map →

saved by

related reading