flâneur — a map of the web's best reading

Recontextualization Mitigates Specification Gaming Without Modifying the Specification — LessWrong

lesswrong.com · 6,765 words · saved by 1 readers

Recontextualization distills good behavior into a context which allows bad behavior. More specifically, recontextualization is a modification to RL which generates completions from prompts that discourage misbehavior, appends those completions to prompts that are more tolerant of misbehavior, and finally reinforces the model on the recontextualized instruction-completion data. Due to the data generation and training prompts differing in their attitude towards misbehavior, recontextualization builds resistance to misbehaviors that the training signal mistakenly reinforces. For example, suppose our reward signal does not robustly penalize deception. Recontextualization generates completions while discouraging deception and then creates training data by updating those completions' prompts to encourage deception. That simple tweak can prevent the model from becoming dishonest! We developed recontextualization concurrently with recent work on inoculation prompting. Wichers et al. and Tan et

x Recontextualization Mitigates Specification Gaming Without Modifying the Specification — LessWrong MATS Program Prompt Engineering Reward Functions AI Frontpage 2025 Top Fifty: 13 % 144 Recontextualization Mitigates Specification Gaming Without Modifying the Specification by ariana_azarbal , Victor Gillioz , TurnTrout , cloud 14th Oct 2025 AI Alignment Forum 13 min read 15 144 Ω 57 Recontextualization distills good behavior into a context which allows bad behavior. More specifically, recontextualization is a modification to RL which generates completions from prompts that discourage misbehav

Explore this link on the map →

related reading