goodfire.ai/blog/painting-with-concepts
We're launching Paint With Ember — a tool for generating and editing images by directly manipulating the neural activations of AI models. We're also open-sourcing the SAE model that powers the app, and sharing our findings on diffusion models and the features they learn. Nick Cammarata † Mark Bissell † Nam Nguyen † Max Loeffler Eric Ho Myra Deng Liv Gorton Daniel Balsam * May 27, 2025 Mechanistic interpretability techniques unlock new and powerful ways of interacting with generative models. By reverse engineering an image model to understand the visual features it's learned, we can then use these features to edit and create images in novel ways. Paint With Ember is a tool that replaces the traditional interface between creators and image models — the familiar prompt box — with a canvas that plugs directly into the "brain" of the model. While it is still possible to guide the model via text prompts (for example, to specify a style), the canvas offers a 2D interface for expressing creati
Painting With Concepts Using Diffusion Model Latents Research Painting with concepts using diffusion model latents We're launching Paint With Ember — a tool for generating and editing images by directly manipulating the neural activations of AI models. We're also open-sourcing the SAE model that powers the app, and sharing our findings on diffusion models and the features they learn. Authors Nick Cammarata * † Mark Bissell * † Nam Nguyen * † Max Loeffler * Eric Ho * Myra Deng * Liv Gorton * Daniel Balsam * Published May 27, 2025 * Goodfire † Core contributor Correspondence to dan@goodfire.ai M
Explore this link on the map →saved by
related reading
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Using Artificial Intelligence to Augment Human Intelligencedistill.pub
- Neuronpedianeuronpedia.org
- The Building Blocks of Interpretabilitydistill.pub
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Feature Visualizationdistill.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Introduction to Exemplar Partitioning for Mechanistic Interpretability — LessWronglesswrong.com
- pdfopenreview.net