flâneur — a map of the web's best reading

Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms — LessWrong

lesswrong.com · 3,512 words · saved by 1 readers

TLDR: We ask whether you can recover the secrets an MLP has memorized, from its weights or via black-box queries, a toy version of the "eliciting bad…

x Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms — LessWrong Interpretability (ML & AI) AI Frontpage 17 Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms by emanuelr 22nd Jun 2026 15 min read 0 17 TLDR: We ask whether you can recover the secrets an MLP has memorized, from its weights or via black-box queries, a toy version of the "eliciting bad contexts" problem for LLMs. We train MLPs as membership classifiers over 16 secret binary strings (34/48/64-bit), across depths 1–3 and several activations, und

Explore this link on the map →

related reading