The case for more ambitious language model evals — LessWrong
lesswrong.com · 6,820 words · saved by 1 readers
Here are some capabilities that I expect to be pretty hard to discover using an RLHF’d chat LLM[1]: …
x The case for more ambitious language model evals — LessWrong AI Evaluations GPT Language Models (LLMs) Practice & Philosophy of Science RLHF Simulator Theory AI Frontpage 121 The case for more ambitious language model evals by Jozdien 30th Jan 2024 AI Alignment Forum 6 min read 30 121 Ω 43 Here are some capabilities that I expect to be pretty hard to discover using an RLHF’d chat LLM [1] : Eric Drexler tried to use the GPT-4 base model as a writing assistant, and it [...] knew who he was from what he was writing. He tried to simulate a conversation to have the AI help him with some writing h
related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Successful language model evals - Jason Weijasonwei.net
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- The bitter lesson of LLM evalsparsed.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- User awareness in frontier modelstransluce.org