flâneur — a map of the web's best reading

Interpretable by Design - Constraint Sets with Disjoint Limit Points — LessWrong

lesswrong.com · 3,626 words · saved by 1 readers

cart;horse: How can we constrain our models to be interpretable? Convex, linear sets make for more interpretable parameter spaces, and the simplex and the Birkhoff Polytope are great examples of this that have other desirable properties. An interpretation is something explicit, something discrete, something that compresses, something that summarizes. Our current paradigms do not lend themselves well to this. We may be able to fine-tune models and interpretations, via approaches built on Provable Guarantees for Model Performance via Mechanistic Interpretability, but in some sense we are fighting an uphill battle against "an uninterpretable base". In the same way we want to create models that are inherently not capable of deception rather than having to evaluate if an unknown model is deceptive, we should aim to create models that are interpretable by default rather than applying interpretability post-hoc. Building interpretable architectures and models from scratch with the explicit goa

x Interpretable by Design - Constraint Sets with Disjoint Limit Points — LessWrong Interpretability (ML & AI) Logic & Mathematics Machine Learning (ML) Optimization AI Frontpage 24 Interpretable by Design - Constraint Sets with Disjoint Limit Points by Ronak_Mehta 8th May 2025 Linkpost for ronakrm.github.io 11 min read 2 24 cart;horse : How can we constrain our models to be interpretable? Convex, linear sets make for more interpretable parameter spaces, and the simplex and the Birkhoff Polytope are great examples of this that have other desirable properties. An interpretation is something expl

Explore this link on the map →

related reading