Turning off lights with model editing — AI Alignment Forum
This post advertises an illustrative example of model editing that David Bau is fond of, and which I think should be better known. As a reminder, David Bau is a professor at Northeastern; the vibe of his work is "interpretability + interventions on model internals"; examples include the well-known ROME, MEMIT, and Othello papers. Consider the problem of getting a generative image model to produce an image of a bedroom containing unlit lamps (i.e. lamps which are turned off). Doesn't sound particularly interesting. Let's try the obvious prompts on DALL-E-2. Doesn't work great: my first three attempts only got one borderline hit, at the expense of turning the entire bedroom dark (which isn't really what I had in mind). As David Bau tells things, even after putting further effort into engineering an appropriate prompt, they weren't able to get what they wanted: a normal picture of a bedroom with lamps which are not turned on. (Apparently the captions for images containing unlit lamps don'
x Turning off lights with model editing — AI Alignment Forum AI Frontpage 31 Turning off lights with model editing by Sam Marks 12th May 2023 3 min read 5 31 This is a linkpost for https://arxiv.org/abs/2207.02774 This post advertises an illustrative example of model editing that David Bau is fond of, and which I think should be better known. As a reminder, David Bau is a professor at Northeastern; the vibe of his work is "interpretability + interventions on model internals"; examples include the well-known ROME , MEMIT , and Othello papers. Consider the problem of getting a generative image m
Explore this link on the map →related reading
- Coding Models Are Doing Too Much | whnrehiew.github.io
- Editing Text in Images with AI | Towards Data Sciencetowardsdatascience.com
- But is it really in Rome? An investigation of the ROME model editing technique — AI Alignment Forumalignmentforum.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Painting With Concepts Using Diffusion Model Latentsgoodfire.ai
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Nano Banana can be prompt engineered for extremely nuanced AI image generation | Max Woolf's Blogminimaxir.com