But is it really in Rome? An investigation of the ROME model editing technique — AI Alignment Forum
Thanks to Andrei Alexandru, Joe Collman, Michael Einhorn, Kyle McDonell, Daniel Paleka, and Neel Nanda for feedback on drafts and/or conversations which led to useful insights for this work. In addition, thank you to both William Saunders and Alex Gray for exceptional mentorship throughout this project. The majority of this work was carried out this summer. Many people in the community were surprised when I mentioned some of the limitations of ROME (Rank-One Model Editing), so I figured it was worth it to write a post about it as well as other insights I gained from looking into the paper. Most tests were done with GPT-2, some were done with GPT-J. The ROME paper (Locating and Editing Factual Associations in GPT) has been one of the most influential papers in the prosaic alignment community. It has several important insights. The main findings are: In this post, I show that the ROME edit has many limitations: One point I want to illustrate with this post is that the intervention is a b
x But is it really in Rome? An investigation of the ROME model editing technique — AI Alignment Forum Interpretability (ML & AI) MATS Program AI Frontpage 49 But is it really in Rome? An investigation of the ROME model editing technique by jacquesthibs 30th Dec 2022 21 min read 2 49 Thanks to Andrei Alexandru, Joe Collman, Michael Einhorn, Kyle McDonell, Daniel Paleka, and Neel Nanda for feedback on drafts and/or conversations which led to useful insights for this work. In addition, thank you to both William Saunders and Alex Gray for exceptional mentorship throughout this project. The majorit
Explore this link on the map →related reading
- Coding Models Are Doing Too Much | whnrehiew.github.io
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Turning off lights with model editing — AI Alignment Forumalignmentforum.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Alignment faking in large language modelsarxiv.org
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net