flâneur — a map of the web's best reading

Advice for making robust-to-training model organisms

blog.redwoodresearch.org · 4,241 words · saved by 1 readers

We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like

Advice for making robust-to-training model organisms +1 Alek Westover , Sebastian Prasanna , Vivek Hebbar , and 2 others May 28, 2026 12 Share We’d like to develop training techniques that work when applied to future misaligned AI systems . One strategy for studying proposed techniques is to test them on model organisms . However, model organisms built with common techniques are often fragile: we (and other researchers like Roger et al. and Ryd et al. ) have observed them to stop misbehaving after untargeted training—training that doesn’t directly target the misbehavior. For example, we have o

Explore this link on the map →

saved by

related reading