flâneur — a map of the web's best reading

Discovering Concept-Editing Algorithms With LLM Agents

dmodel.ai · 1,824 words · saved by 1 readers

Adam Scherlis*, Alex Bishka*, Ritesh Bhalerao, Kunal Bham, Tyra Burgess, David Cao, Jeff Coggshall, Dalton Combs, Alexana Dubois, Shaurya Jain, Zhanna Kaufman, Felix Michalak, Daniel Moon, Arjun Pagidi, Alex Sanchez-Stern, Autumn Sinclair, Ajit Sivakumar, Mohit Tekriwal, Aditya Vasantharao, Curry Winter, Yueru Yan, Anish Tondwalkar* * Core contributors 𝑑 model d model ​ Concept erasure is a technique that removes unwanted information from a model’s activations, but current erasure methods struggle to fully remove target concepts. In this study, we tasked LLM agents trained on our data with inventing concept erasure algorithms that outperform current methods under the same experimental constraints. We measure the performance of each algorithm family and explore the cause of why current methods fall short. To create safe models, we must be able to control what they know and use. Concept erasure is one tool for modifying a model’s representations. Rather than retraining or fine-tuning,

Discovering Concept-Editing Algorithms With LLM Agents 1 Introduction To create safe models, we must be able to control what they know and use. Concept erasure is one tool for modifying a model’s representations. Rather than retraining or fine-tuning, concept erasure reaches into a model’s activations and removes a target concept, rendering it unusable. The ability to erase concepts supports a range of AI-safety efforts. Concept erasure can help models unlearn knowledge needed to build bioweapons, remove sensitive attributes that enable discrimination, and ablate concepts that drive misalignme

Explore this link on the map →

related reading