Attention-Guided-RL for Human-Like LMs | Manifund
Current base language models (LMs) do not have the ability to control the order in which they receive training data. This leads them to become a fairly inhuman kind of cognition, similar to what a human baby might become if they were raised on an appetite of random YouTube videos with no ability to influence the video order. This project is about training the LM to pick its training order, with the hope the it will lead to a more human cognition. Prior work uses heuristics such as entropy on sub-distributions of data in order to pick future samples. We propose a technique that is closer to the core mechanism of a transformer -- we can let the model's attention mechanism guide which training sample it receives, and fine-tune that attention mechanism using an RL algorithm, such as expert iteration with KL regularization, PPO, or GRPO. Suppose we randomly sample a selection of consecutive sentences from a Wikipedia article -- let's think of the sentences pairs as keys-values pairs in a
Attention-Guided-RL for Human-Like LMs | Manifund 4 Attention-Guided-RL for Human-Like LMs Technical AI safety 🐯 Scott Viteri Active Grant $3,100 raised $30,000 funding goal Donate Sign in to donate Project summary Current base language models (LMs) do not have the ability to control the order in which they receive training data. This leads them to become a fairly inhuman kind of cognition, similar to what a human baby might become if they were raised on an appetite of random YouTube videos with no ability to influence the video order. This project is about training the LM to pick its trainin
Explore this link on the map →related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- GenAI Handbookgenai-handbook.github.io
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- rlhfbook.com/book.pdfrlhfbook.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Proposal: Using Monte Carlo tree search instead of RLHF for alignment research — LessWronglesswrong.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com