Behavior Cloning is Miscalibrated - AI Alignment Forum
Behavior cloning (BC) is, put simply, when you have a bunch of human expert demonstrations and you train your policy to maximize likelihood over the human expert demonstrations. It’s the simplest pos…
x Behavior Cloning is Miscalibrated — AI Alignment Forum Calibration Machine Learning (ML) Outer Alignment Reinforcement learning AI Frontpage 36 Behavior Cloning is Miscalibrated by leogao 5th Dec 2021 4 min read 3 36 Behavior cloning (BC) is, put simply, when you have a bunch of human expert demonstrations and you train your policy to maximize likelihood over the human expert demonstrations. It’s the simplest possible approach under the broader umbrella of Imitation Learning, which also includes more complicated things like Inverse Reinforcement Learning or Generative Adversarial Imitation L
saved by
related reading
- Behavior Cloning is Miscalibrated — LessWronglesswrong.com
- Behavioral cloning mysteryseohong.me
- State of Robot Learning, December 2025vedder.io
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com