✳flâneur — a map of the web's best reading
Model organisms researchers should check whether high LRs defeat their model organisms — LessWrong
lesswrong.com · 2,105 words · saved by 1 readers
Thanks to Buck Shlegeris for feedback on a draft of this post. …
x Model organisms researchers should check whether high LRs defeat their model organisms — LessWrong AI Frontpage 40 Model organisms researchers should check whether high LRs defeat their model organisms by Dylan Xu , SebastianP , Alek Westover , Vivek Hebbar , Julian Stastny 10th Apr 2026 6 min read 0 40 Thanks to Buck Shlegeris for feedback on a draft of this post. The goal-guarding hypothesis states that schemers will be able to preserve their goals during training by taking actions which are selected for by the training process. To investigate the goal-guarding hypothesis, we’ve been runni
Explore this link on the map →saved by
related reading
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Composer2.pdfcursor.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org