Conditioning, Prompts, and Fine-Tuning — LessWrong
lesswrong.com · 2,319 words · saved by 1 readers
(Thanks to Evan Hubinger and Nicholas Schiefer for comments on these ideas.) …
x Conditioning, Prompts, and Fine-Tuning — LessWrong Language Models (LLMs) Outer Alignment Reinforcement learning AI Frontpage 38 Conditioning, Prompts, and Fine-Tuning by Adam Jermyn 17th Aug 2022 AI Alignment Forum 4 min read 9 38 Ω 17 (Thanks to Evan Hubinger and Nicholas Schiefer for comments on these ideas.) These are some notes on the relation between conditioning language models, prompting, and fine-tuning. The key takeaways are: Prompting and fine-tuning can both be used to condition language models. Prompting is quite restricted in the kinds of conditionals it can achieve. Fine-tunin
saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersarxiv.org
- Recent Advances in Language Model Fine-tuningruder.io
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimesarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- 2402.07927arxiv.org