Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWrong
TLDR: The idea is basically inoculation prompting crossed with alignment pretraining. Call it ‘inoculation pretraining.’ It’s a type of spillway desi…
x Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWrong AI Frontpage 12 Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs by Elliott Thornley (EJT) 29th May 2026 3 min read 4 12 TLDR: The idea is basically inoculation prompting crossed with alignment pretraining . Call it ‘inoculation pretraining.’ It’s a type of spillway design . ---------------------------------------------------------------------------------------------------- Reward hacking can cause emergent misalignment : you train the AI to cheat on its tasks and it turns bro
Explore this link on the map →saved by
related reading
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- How far does alignment midtraining generalize?alignment.openai.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Training on Documents About Reward Hacking Induces Reward Hacking — LessWronglesswrong.com
- Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com