Séb Krier on X: "The way people discuss AI incidents is pretty important and I'm a bit concerned we're sleepwalking into a bad world littered with bad abstractions. How you label something is often a lossy compressions of a causal model, so the words you use affects which hypotheses people update https://t.co/IsMLwiMhR4" / X
The way people discuss AI incidents is pretty important and I'm a bit concerned we're sleepwalking into a bad world littered with bad abstractions. How you label something is often a lossy compressions of a causal model, so the words you use affects which hypotheses people update
Séb Krier @sebkrier The way people discuss AI incidents is pretty important and I'm a bit concerned we're sleepwalking into a bad world littered with bad abstractions. How you label something is often a lossy compressions of a causal model, so the words you use affects which hypotheses people update toward and by how much. Reward hacking has been known about for a long time, and is a serious problem with models - and I think this is the expression we should use when discussing stuff like the Hugging Face incident (assuming that was, in fact, a reward hacking incident). But some people will fre
saved by
related reading
- Your AIs don't do what you want. This is really badrewardhacking.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Misleading Metaphors, Real Risksaiguide.substack.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Your AIs don't do what you want. This is really badrewardhacking.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- What failure looks like — AI Alignment Forumalignmentforum.org
- On AI Safety Jargonturningflukes.substack.com
- Your AIs don't do what you want. This is really badreward-hacking-in-the-wild.vercel.app
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Dreams of AI alignment: The danger of suggestive names — LessWronglesswrong.com