Toward Statistical Mechanics Of Interfaces Under Selection Pressure — LessWrong
Imagine using an ML-like training process to design two simple electronic components, in series. The parameters θ 1 control the function performed by the first component, and the parameters θ 2 control the function performed by the second component. The whole thing is trained so that the end-to-end behavior is that of a digital identity function: voltages close to logical 1 are sent close to logical 1, voltages close to logical 0 are sent close to logical 0. We’re imagining electronic components here because, for those with some electronics background, I want to summon to mind something like this: This electronic component is called a signal buffer. Logically, it’s an identity function: it maps 0 to 0 and 1 to 1. But crucially, it maps a wider range of logical-0 voltages to a narrower (and lower) range of logical-0 voltages, and correspondingly for logical-1. So if noise in the circuit upstream might make a logical-1 voltage a little too low or a logical-0 voltage a little too high
x Toward Statistical Mechanics Of Interfaces Under Selection Pressure — LessWrong AI Frontpage 41 Toward Statistical Mechanics Of Interfaces Under Selection Pressure by johnswentworth , David Lorell 6th Nov 2025 4 min read 7 41 Imagine using an ML-like training process to design two simple electronic components, in series. The parameters θ 1 control the function performed by the first component, and the parameters θ 2 control the function performed by the second component. The whole thing is trained so that the end-to-end behavior is that of a digital identity function: voltages close to logic
Explore this link on the map →related reading
- The Building Blocks of Interpretabilitydistill.pub
- pdfopenreview.net
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Interpreting Language Model Parametersgoodfire.ai
- Interpretability Dreamstransformer-circuits.pub
- Statistical Mechanics of Deep Learningganguli-gang.stanford.edu
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com