Behavior Cloning is Miscalibrated - AI Alignment Forum
Behavior cloning (BC) is, put simply, when you have a bunch of human expert demonstrations and you train your policy to maximize likelihood over the human expert demonstrations. It’s the simplest pos…
x Behavior Cloning is Miscalibrated — AI Alignment Forum Calibration Machine Learning (ML) Outer Alignment Reinforcement learning AI Frontpage 36 Behavior Cloning is Miscalibrated by leogao 5th Dec 2021 4 min read 3 36 Behavior cloning (BC) is, put simply, when you have a bunch of human expert demonstrations and you train your policy to maximize likelihood over the human expert demonstrations. It’s the simplest possible approach under the broader umbrella of Imitation Learning, which also includes more complicated things like Inverse Reinforcement Learning or Generative Adversarial Imitation L
Explore this link on the map →saved by
related reading
- Behavior Cloning is Miscalibrated — LessWronglesswrong.com
- State of Robot Learning, December 2025vedder.io
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- Alignment faking in large language modelsarxiv.org