Turning off lights with model editing — AI Alignment Forum
This post advertises an illustrative example of model editing that David Bau is fond of, and which I think should be better known. As a reminder, David Bau is a professor at Northeastern; the vibe of his work is "interpretability + interventions on model internals"; examples include the well-known ROME, MEMIT, and Othello papers. Consider the problem of getting a generative image model to produce an image of a bedroom containing unlit lamps (i.e. lamps which are turned off). Doesn't sound particularly interesting. Let's try the obvious prompts on DALL-E-2. Doesn't work great: my first three attempts only got one borderline hit, at the expense of turning the entire bedroom dark (which isn't really what I had in mind). As David Bau tells things, even after putting further effort into engineering an appropriate prompt, they weren't able to get what they wanted: a normal picture of a bedroom with lamps which are not turned on. (Apparently the captions for images containing unlit lamps don'
x Turning off lights with model editing — AI Alignment Forum AI Frontpage 31 Turning off lights with model editing by Sam Marks 12th May 2023 3 min read 5 31 This is a linkpost for https://arxiv.org/abs/2207.02774 This post advertises an illustrative example of model editing that David Bau is fond of, and which I think should be better known. As a reminder, David Bau is a professor at Northeastern; the vibe of his work is "interpretability + interventions on model internals"; examples include the well-known ROME , MEMIT , and Othello papers. Consider the problem of getting a generative image m
related reading
- Uncensored Modelserichartford.com
- Coding Models Are Doing Too Much | whnrehiew.github.io
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Painting With Concepts Using Diffusion Model Latentsgoodfire.ai
- Editing Text in Images with AI | Towards Data Sciencetowardsdatascience.com
- But is it really in Rome? An investigation of the ROME model editing technique — AI Alignment Forumalignmentforum.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersarxiv.org
- Topicslearnmechinterp.com