flâneur — a map of the web's best reading

Steering Along Manifolds to Control Neural Networks

goodfire.ai · 1,261 words · saved by 1 readers

Concept geometry provides a blueprint for controlling the behavior of neural networks—if you know how to look. Intervening on a model's internal representations to steer behavior, i.e., representation steering, promises lightweight, adaptable, and granular control of neural networks. That control can be leveraged during inference and training to design and align models. The typical approach is to steer in a straight line by adding a scaled "steering vector" to a hidden representation. This is successful when the concept being manipulated lies on a straight line in the model's representations, but what if this isn't the case? In our previous post, we posited that steering along manifolds, i.e., curved surfaces in representation space, provides better control than linear steering (see the mountain car demo). Here, we continue to make this case, using the cyclic concept geometry of the days of the week as a case study (see the paper for details and more complex tasks). Moreover, we show t

Steering Along Manifolds to Control Neural Networks Research ← The Neural Geometry Series → Steering Along Manifolds to Control Neural Networks Authors Daniel Wurgaft *,1,2 Noah D. Goodman †,2 Can Rager *,1,3 Thomas Fel †,1 Matthew Kowal *,1 Atticus Geiger †,1 Vasudev Shyam 1 Ekdeep Singh Lubana †,1 Sheridan Feucht 1,4 Usha Bhalla 1,5 * Equal contribution Tal Haklay 1,6 † Equal senior contribution Eric Bigelow 1,5 1 Goodfire Raphael Sarfati 1 2 Stanford University Thomas McGrath 1 3 University College London Owen Lewis 1 4 Northeastern University Jack Merullo 1 5 Harvard University 6 Technion

Explore this link on the map →

saved by

related reading