The "Minimal Latents" Approach to Natural Abstractions — AI Alignment Forum
How many functions are there from a 1 megabyte image to a yes/no answer to the question “does this image contain an apple?”. Well, there are 8M bits in the image, so 2 8000000 possible images. The function can assign “yes” or “no” independently to each of those 2 8000000 images, so 2 2 8000000 possible functions. Specifying one such function by brute force (i.e. not leveraging any strong prior information) would therefore require 2 8000000 bits of information - one bit specifying the output on each of the 2 8000000 possible images. Even if we allow for some wiggle room in ambiguous images, those ambiguous images will still be a very tiny proportion of image-space, so the number of bits required would still be exponentially huge. Empirically, human toddlers are able to recognize apples by sight after seeing maybe one to three examples. (Source: people with kids.) Point is: nearly-all the informational work done in a toddler’s mind of figuring out which pattern is referred to b
x The "Minimal Latents" Approach to Natural Abstractions — AI Alignment Forum Natural Abstraction AI Frontpage 16 The "Minimal Latents" Approach to Natural Abstractions by johnswentworth 20th Dec 2022 14 min read 24 16 Background: The Language-Learning Argument How many functions are there from a 1 megabyte image to a yes/no answer to the question “does this image contain an apple?”. Well, there are 8M bits in the image, so 2 8000000 possible images. The function can assign “yes” or “no” independently to each of those 2 8000000 images, so 2 2 8000000 possible functions. Specifying one such fun
Explore this link on the map →related reading
- Natural Abstractions: Key Claims, Theorems, and Critiques — LessWronglesswrong.com
- Natural Latents: The Math — LessWronglesswrong.com
- The Natural Abstraction Hypothesis: Implications and Evidence — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Generative modelling in latent space – Sander Dielemansander.ai
- (Approximately) Deterministic Natural Latents — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- Why Care About Natural Latents? — LessWronglesswrong.com
- What If We Had Bigger Brains? Imagining Minds beyond Ours-Stephen Wolfram Writingswritings.stephenwolfram.com
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- 2023 letter | Zhengdongzhengdongwang.com
- To Understand Language is to Understand Generalization | Eric Jangevjang.com