✳flâneur — a map of the web's best reading
Conditioning, Prompts, and Fine-Tuning — LessWrong
lesswrong.com · 2,319 words · saved by 1 readers
(Thanks to Evan Hubinger and Nicholas Schiefer for comments on these ideas.) …
x Conditioning, Prompts, and Fine-Tuning — LessWrong Language Models (LLMs) Outer Alignment Reinforcement learning AI Frontpage 38 Conditioning, Prompts, and Fine-Tuning by Adam Jermyn 17th Aug 2022 AI Alignment Forum 4 min read 9 38 Ω 17 (Thanks to Evan Hubinger and Nicholas Schiefer for comments on these ideas.) These are some notes on the relation between conditioning language models, prompting, and fine-tuning. The key takeaways are: Prompting and fine-tuning can both be used to condition language models. Prompting is quite restricted in the kinds of conditionals it can achieve. Fine-tunin
Explore this link on the map →saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Recent Advances in Language Model Fine-tuningruder.io
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Few-Shot Prompting | Prompt Engineering Guidepromptingguide.ai
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Modelsarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)arxiv.org
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org