A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
arxiv.org · 6,398 words · saved by 1 readers
N/A
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity Andrew Lee 1 Xiaoyan Bai 1 Itamar Pres 1 Martin Wattenberg 2 Jonathan K. Kummerfeld 3 Rada Mihalcea 1 Abstract our understanding of the mechanisms by which the unde- sirable behavior is…
saved by
related reading
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Preference Tuning LLMs with Direct Preference Optimization Methodshuggingface.co
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Studyarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Fine-tune Llama 2 with DPOhuggingface.co
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- Alignment Faking Mitigationsalignment.anthropic.com
- RLHF | John Lambertjohnwlambert.github.io
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org