Why Care About Natural Latents? — LessWrong
Suppose Alice and Bob are two Bayesian agents in the same environment. They both basically understand how their environment works, so they generally agree on predictions about any specific directly-observable thing in the world - e.g. whenever they try to operationalize a bet, they find that their odds are roughly the same. However, their two world models might have totally different internal structure, different “latent” structures which Alice and Bob model as generating the observable world around them. As a simple toy example: maybe Alice models a bunch of numbers as having been generated by independent rolls of the same biased die, and Bob models the same numbers using some big complicated neural net. Now suppose Alice goes poking around inside of her world model, and somewhere in there she finds a latent variable Λ A with two properties (the Natural Latent properties): In the die/net case, the die’s bias ( Λ A ) approximately mediates between e.g. the first 100 numbers ( X 1 ) a
x Why Care About Natural Latents? — LessWrong AI Frontpage 56 Why Care About Natural Latents? by johnswentworth , David Lorell 9th May 2024 5 min read 3 56 Suppose Alice and Bob are two Bayesian agents in the same environment. They both basically understand how their environment works, so they generally agree on predictions about any specific directly-observable thing in the world - e.g. whenever they try to operationalize a bet, they find that their odds are roughly the same. However, their two world models might have totally different internal structure, different “latent” structures which A
Explore this link on the map →related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Transformer Circuits Threadtransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- Natural Latents: The Math — LessWronglesswrong.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- The Building Blocks of Interpretabilitydistill.pub
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- (Approximately) Deterministic Natural Latents — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- World-Model Interpretability Is All We Need — AI Alignment Forumalignmentforum.org