Proposal: Using Monte Carlo tree search instead of RLHF for alignment research — LessWrong
Currently the most powerful techniques for getting a language model to act as an agent is via RLHF and similar approaches. For example, ChatGPT was t…
x Proposal: Using Monte Carlo tree search instead of RLHF for alignment research — LessWrong AI Risk AIXI Inner Alignment Language Models (LLMs) RLHF AI Frontpage 2 Proposal: Using Monte Carlo tree search instead of RLHF for alignment research by Christopher King 20th Apr 2023 4 min read 7 2 Currently the most powerful techniques for getting a language model to act as an agent is via RLHF and similar approaches. For example, ChatGPT was trained to be an agent that tries to give humans answers that they want. Another approach is taking a LLM and getting it to predict what the agent you want wou
Explore this link on the map →saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- Thoughts on the impact of RLHF research — AI Alignment Forumalignmentforum.org
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- RL & search is a terrifying way to build AGI (an FAQ) — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- rlhfbook.com/book.pdfrlhfbook.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- the void — LessWronglesswrong.com