flâneur — a map of the web's best reading

Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWrong

lesswrong.com · 1,266 words · saved by 1 readers

TLDR: The idea is basically inoculation prompting crossed with alignment pretraining. Call it ‘inoculation pretraining.’ It’s a type of spillway desi…

x Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWrong AI Frontpage 12 Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs by Elliott Thornley (EJT) 29th May 2026 3 min read 4 12 TLDR: The idea is basically inoculation prompting crossed with alignment pretraining . Call it ‘inoculation pretraining.’ It’s a type of spillway design . ---------------------------------------------------------------------------------------------------- Reward hacking can cause emergent misalignment : you train the AI to cheat on its tasks and it turns bro

Explore this link on the map →

saved by

related reading