✳flâneur — a map of the web's best reading
Stochastic Parameter Decomposition
arxiv.org · 12,758 words · saved by 2 readers
N/A
# link_n87icxthf6.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Lucius Bushnaq; Dan Braun; Lee Sharkey - Creator=arXiv GenPDF (tex2pdf:) - Custom.DOI=https://doi.org/10.48550/arXiv.2506.20790 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Custom.arXivID=https://arxiv.org/abs/2506.20790v2 - Producer=pikepdf 8.15.1 - Title=Stochastic Parameter De
Explore this link on the map →saved by
related reading
- Towards Scalable Parameter Decompositiongoodfire.ai
- Attribution-based parameter decomposition — LessWronglesswrong.com
- Interpreting Language Model Parametersgoodfire.ai
- Towards Scalable Parameter Decompositiongoodfire.ai
- [2506.20790] Stochastic Parameter Decompositionarxiv.org
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Toy Models of Superpositiontransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- [2501.14926] Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org