Towards a better circuit prior: Improving on ELK state-of-the-art - LessWrong
lesswrong.com · 6,583 words · saved by 1 readers
This post is the result of joint work with Kate Woolverton. Thanks to Paul Christiano for useful comments and feedback. …
x Towards a better circuit prior: Improving on ELK state-of-the-art — LessWrong Eliciting Latent Knowledge AI Frontpage 23 Towards a better circuit prior: Improving on ELK state-of-the-art by evhub , kcwoolverton 29th Mar 2022 AI Alignment Forum 17 min read 0 23 Ω 13 Thanks to Paul Christiano for useful comments and feedback. The basic circuit prior setup We’ll start with the basic setup that we’re trying to improve upon, which is trying to solve ELK via the use of a Boolean circuit size prior . Previously, Evan summarized Paul, Mark, and Ajeya’s argument for why this might work as follows : A
saved by
related reading
- Musings on the Speed Prior — AI Alignment Forumalignmentforum.org
- Eliciting Latent Knowledge (ELK) - Distillation/Summary — AI Alignment Forumalignmentforum.org
- As Rocks May Think | Eric Jangevjang.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- The generalization phase diagram — LessWronglesswrong.com
- [2605.12671] All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMsarxiv.org
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com