✳flâneur — a map of the web's best reading
The case for more ambitious language model evals — LessWrong
lesswrong.com · 6,820 words · saved by 1 readers
Here are some capabilities that I expect to be pretty hard to discover using an RLHF’d chat LLM[1]: …
x The case for more ambitious language model evals — LessWrong AI Evaluations GPT Language Models (LLMs) Practice & Philosophy of Science RLHF Simulator Theory AI Frontpage 121 The case for more ambitious language model evals by Jozdien 30th Jan 2024 AI Alignment Forum 6 min read 30 121 Ω 43 Here are some capabilities that I expect to be pretty hard to discover using an RLHF’d chat LLM [1] : Eric Drexler tried to use the GPT-4 base model as a writing assistant, and it [...] knew who he was from what he was writing. He tried to simulate a conversation to have the AI help him with some writing h
Explore this link on the map →related reading
- gpt-4.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The bitter lesson of LLM evalsparsed.com
- Quantifying Truesight With SAEs · Gwern.netgwern.net
- LLMs Know More Than What They Say - by Ruby Paiarjunbansal.substack.com
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com