(Approximately) Deterministic Natural Latents — LessWrong
Background: Natural Latents: The Math, Natural Latents: The Concepts, Why Care About Natural Latents?, the prototypical semantics use-case. This post does not assume that you’ve read all of those, or even any of them. Suppose I roll a biased die 1000 times, and then roll the same biased die another 1000 times. Then... The die’s bias is therefore a natural latent, which means it has various nice properties. Furthermore, the bias is a(n approximate) deterministic natural latent: the die’s bias (to reasonable precision) is approximately determined by[1] the first 1000 die rolls, and also approximately determined by the second 1000 die rolls. That implies one more nice property: We’ve proven all that before, mostly in Natural Latents: The Math (including the addendum added six months after the rest of the post). But it turns out that the math is a lot shorter and simpler, and easily yields better bounds, if we’re willing to assume (approximate) determinism up-front. That does lose us some
x (Approximately) Deterministic Natural Latents — LessWrong AI Frontpage 45 (Approximately) Deterministic Natural Latents by johnswentworth , David Lorell 19th Jul 2024 AI Alignment Forum 5 min read 1 45 Ω 23 Background: Natural Latents: The Math , Natural Latents: The Concepts , Why Care About Natural Latents? , the prototypical semantics use-case . This post does not assume that you’ve read all of those, or even any of them. Suppose I roll a biased die 1000 times, and then roll the same biased die another 1000 times. Then... Mediation: The first 1000 rolls are approximately independent of th
Explore this link on the map →related reading
- Natural Latents: The Math — LessWronglesswrong.com
- Why Care About Natural Latents? — LessWronglesswrong.com
- A Mathematical Theory of Communicationpeople.math.harvard.edu
- The "Minimal Latents" Approach to Natural Abstractions — AI Alignment Forumalignmentforum.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Laplace's demon - Wikipediaen.wikipedia.org
- Natural Abstractions: Key Claims, Theorems, and Critiques — LessWronglesswrong.com
- De Finetti's theorem - Wikipediaen.wikipedia.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Gregory Gundersengregorygundersen.com
- 1 An Informal Introduction to Intelligencema-lab-berkeley.github.io
- Towards a better circuit prior: Improving on ELK state-of-the-art — LessWronglesswrong.com