✳flâneur — a map of the web's best reading
Towards a better circuit prior: Improving on ELK state-of-the-art - LessWrong
lesswrong.com · 6,583 words · saved by 1 readers
This post is the result of joint work with Kate Woolverton. Thanks to Paul Christiano for useful comments and feedback. …
x Towards a better circuit prior: Improving on ELK state-of-the-art — LessWrong Eliciting Latent Knowledge AI Frontpage 23 Towards a better circuit prior: Improving on ELK state-of-the-art by evhub , kcwoolverton 29th Mar 2022 AI Alignment Forum 17 min read 0 23 Ω 13 Thanks to Paul Christiano for useful comments and feedback. The basic circuit prior setup We’ll start with the basic setup that we’re trying to improve upon, which is trying to solve ELK via the use of a Boolean circuit size prior . Previously, Evan summarized Paul, Mark, and Ajeya’s argument for why this might work as follows : A
Explore this link on the map →saved by
related reading
- Musings on the Speed Prior — AI Alignment Forumalignmentforum.org
- Eliciting Latent Knowledge (ELK) - Distillation/Summary — AI Alignment Forumalignmentforum.org
- On the Biology of a Large Language Modeltransformer-circuits.pub
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- [2605.12671] All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMsarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- Shortform — LessWronglesswrong.com
- Better priors as a safety problem — LessWronglesswrong.com
- Lucius Bushnaq's Shortform — LessWronglesswrong.com