Fluent dreaming for language models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. [2]label=. Feature visualization, also known as "dreaming", offers insights into vision models by optimizing the inputs to maximize a neuron’s activation or other internal component. However, dreaming has not been successfully applied to language models because the input space is discrete. We extend Greedy Coordinate Gradient, a method from the language model adversarial attack literature, to design the Evolutionary Prompt Optimization (EPO) algorithm. EPO optimizes the input prompt to simultaneously maximize the Pareto frontier between a chosen internal feature and prompt fluency, enabling fluent dreaming for language models. We demonstrate dreaming with neurons, output logits and arbitrary directions in activation space. We measure the fluency of the r
\setenumerate [2]label=. Fluent dreaming for language models T. Ben Thompson, Zygimantas Straznickas, Michael Sklar Confirm Labs t.ben.thompson@gmail.com (January 2024) Abstract Feature visualization, also known as "dreaming", offers insights into vision models by optimizing the inputs to maximize a neuron’s activation or other internal component. However, dreaming has not been successfully applied to language models because the input space is discrete. We extend Greedy Coordinate Gradient, a method from the language model adversarial attack literature, to design the Evolutionary Prompt Optimi
Explore this link on the map →related reading
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Transformer Circuits Threadtransformer-circuits.pub
- LLM Daydreaming · Gwern.netgwern.net
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Language Modelinglena-voita.github.io
- Feature Visualizationdistill.pub