flâneur — a map of the web's best reading

Model organisms researchers should check whether high LRs defeat their model organisms — LessWrong

lesswrong.com · 2,105 words · saved by 1 readers

Thanks to Buck Shlegeris for feedback on a draft of this post. …

x Model organisms researchers should check whether high LRs defeat their model organisms — LessWrong AI Frontpage 40 Model organisms researchers should check whether high LRs defeat their model organisms by Dylan Xu , SebastianP , Alek Westover , Vivek Hebbar , Julian Stastny 10th Apr 2026 6 min read 0 40 Thanks to Buck Shlegeris for feedback on a draft of this post. The goal-guarding hypothesis states that schemers will be able to preserve their goals during training by taking actions which are selected for by the training process. To investigate the goal-guarding hypothesis, we’ve been runni

Explore this link on the map →

saved by

related reading