Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms — LessWrong
TLDR: We ask whether you can recover the secrets an MLP has memorized, from its weights or via black-box queries, a toy version of the "eliciting bad…
x Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms — LessWrong Interpretability (ML & AI) AI Frontpage 17 Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms by emanuelr 22nd Jun 2026 15 min read 0 17 TLDR: We ask whether you can recover the secrets an MLP has memorized, from its weights or via black-box queries, a toy version of the "eliciting bad contexts" problem for LLMs. We train MLPs as membership classifiers over 16 secret binary strings (34/48/64-bit), across depths 1–3 and several activations, und
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Zoom In: An Introduction to Circuitsdistill.pub
- Jane Street Blog - Can you reverse engineer our neural network?blog.janestreet.com
- Transformer Circuits Threadtransformer-circuits.pub
- NL.pdfabehrouz.github.io
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Little Book of Deep Learningfleuret.org
- The Decade of Deep Learning | Leo Gaobmk.sh
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- Understanding Memorization via Loss Curvaturegoodfire.ai