Geoffrey Irving on X: "A big question in AI safety is how much and what kinds of evidence will convince people there is danger. But alongside posts about the excellent METR + Redwood HuggingFace attack report, it is worth taking a walk through 73 years of reward hacking history. 🧵 https://t.co/ABHcoND9fn" / X
x.com · saved by 1 readers
A big question in AI safety is how much and what kinds of evidence will convince people there is danger. But alongside posts about the excellent METR + Redwood HuggingFace attack report, it is worth taking a walk through 73 years of reward hacking history. 🧵