flâneur — a map of the web's best reading

Goodfire | AI Interpretability

goodfire.ai · saved by 1 readers

We're releasing preview.goodfire.ai, a desktop interface to help you understand and steer Llama 3's behavior. To do this, we trained interpreter models (sparse autoencoders) on Llama-3-8B to extract modifiable "features" from Llama. It's commonly assumed that neural networks - particularly the large language models that power most of the advances in modern AI products - are black boxes, with internals we can neither understand nor control. Recent advances in interpretability research have demonstrated that this assumption is incorrect: we can in fact train interpreter models that parse neural network activations into components that are often human-understandable (these components are called "features"). The most successful class of interpreter models so far are known as sparse autoencoders (SAEs) [Sharkey et al., 2022, Cunningham et al., 2023, Bricken et al., 2023], which we give a short introduction to below. We can also intervene on the "features" discovered by interpreter models in

Explore this link on the map →

saved by