Discovering Concept-Editing Algorithms With LLM Agents
Adam Scherlis*, Alex Bishka*, Ritesh Bhalerao, Kunal Bham, Tyra Burgess, David Cao, Jeff Coggshall, Dalton Combs, Alexana Dubois, Shaurya Jain, Zhanna Kaufman, Felix Michalak, Daniel Moon, Arjun Pagidi, Alex Sanchez-Stern, Autumn Sinclair, Ajit Sivakumar, Mohit Tekriwal, Aditya Vasantharao, Curry Winter, Yueru Yan, Anish Tondwalkar* * Core contributors 𝑑 model d model Concept erasure is a technique that removes unwanted information from a model’s activations, but current erasure methods struggle to fully remove target concepts. In this study, we tasked LLM agents trained on our data with inventing concept erasure algorithms that outperform current methods under the same experimental constraints. We measure the performance of each algorithm family and explore the cause of why current methods fall short. To create safe models, we must be able to control what they know and use. Concept erasure is one tool for modifying a model’s representations. Rather than retraining or fine-tuning,
Discovering Concept-Editing Algorithms With LLM Agents 1 Introduction To create safe models, we must be able to control what they know and use. Concept erasure is one tool for modifying a model’s representations. Rather than retraining or fine-tuning, concept erasure reaches into a model’s activations and removes a target concept, rendering it unusable. The ability to erase concepts supports a range of AI-safety efforts. Concept erasure can help models unlearn knowledge needed to build bioweapons, remove sensitive attributes that enable discrimination, and ablate concepts that drive misalignme
Explore this link on the map →related reading
- Large Language Models Relearn Removed Concepts - ACL Anthologyaclanthology.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- Painting With Concepts Using Diffusion Model Latentsgoodfire.ai
- GenAI Handbookgenai-handbook.github.io
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Understanding Memorization via Loss Curvaturegoodfire.ai
- Large Concept Models (LCMs) by Meta: The Era of AI After LLMs?aipapersacademy.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Competing with sampling — Alignment Research Centeralignment.org
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com