flâneur — a map of the web's best reading

Proposal: Using Monte Carlo tree search instead of RLHF for alignment research — LessWrong

lesswrong.com · 1,582 words · saved by 1 readers

Currently the most powerful techniques for getting a language model to act as an agent is via RLHF and similar approaches. For example, ChatGPT was t…

x Proposal: Using Monte Carlo tree search instead of RLHF for alignment research — LessWrong AI Risk AIXI Inner Alignment Language Models (LLMs) RLHF AI Frontpage 2 Proposal: Using Monte Carlo tree search instead of RLHF for alignment research by Christopher King 20th Apr 2023 4 min read 7 2 Currently the most powerful techniques for getting a language model to act as an agent is via RLHF and similar approaches. For example, ChatGPT was trained to be an agent that tries to give humans answers that they want. Another approach is taking a LLM and getting it to predict what the agent you want wou

Explore this link on the map →

saved by

related reading