Interpretable by Design - Constraint Sets with Disjoint Limit Points — LessWrong
cart;horse: How can we constrain our models to be interpretable? Convex, linear sets make for more interpretable parameter spaces, and the simplex and the Birkhoff Polytope are great examples of this that have other desirable properties. An interpretation is something explicit, something discrete, something that compresses, something that summarizes. Our current paradigms do not lend themselves well to this. We may be able to fine-tune models and interpretations, via approaches built on Provable Guarantees for Model Performance via Mechanistic Interpretability, but in some sense we are fighting an uphill battle against "an uninterpretable base". In the same way we want to create models that are inherently not capable of deception rather than having to evaluate if an unknown model is deceptive, we should aim to create models that are interpretable by default rather than applying interpretability post-hoc. Building interpretable architectures and models from scratch with the explicit goa
x Interpretable by Design - Constraint Sets with Disjoint Limit Points — LessWrong Interpretability (ML & AI) Logic & Mathematics Machine Learning (ML) Optimization AI Frontpage 24 Interpretable by Design - Constraint Sets with Disjoint Limit Points by Ronak_Mehta 8th May 2025 Linkpost for ronakrm.github.io 11 min read 2 24 cart;horse : How can we constrain our models to be interpretable? Convex, linear sets make for more interpretable parameter spaces, and the simplex and the Birkhoff Polytope are great examples of this that have other desirable properties. An interpretation is something expl
Explore this link on the map →related reading
- Toy Models of Superpositiontransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Softmax Linear Unitstransformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- On Optimism for Interpretabilitygoodfire.ai
- Interpretability Dreamstransformer-circuits.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org