Model organisms researchers should check whether high LRs defeat their model organisms — LessWrong
lesswrong.com · 2,105 words · saved by 1 readers
Thanks to Buck Shlegeris for feedback on a draft of this post. …
x Model organisms researchers should check whether high LRs defeat their model organisms — LessWrong AI Frontpage 40 Model organisms researchers should check whether high LRs defeat their model organisms by Dylan Xu , SebastianP , Alek Westover , Vivek Hebbar , Julian Stastny 10th Apr 2026 6 min read 0 40 Thanks to Buck Shlegeris for feedback on a draft of this post. The goal-guarding hypothesis states that schemers will be able to preserve their goals during training by taking actions which are selected for by the training process. To investigate the goal-guarding hypothesis, we’ve been runni
saved by
related reading
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Composer2.pdfcursor.com
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalizationarxiv.org
- Discovering Undesired Rare Behaviors via Model Diff Amplification - Goodfiregoodfire.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Reinforcement learning towards broadly and persistently beneficial modelsalignment.openai.com