flâneur — a map of the web's best reading

SAE feature geometry is outside the superposition hypothesis — LessWrong

lesswrong.com · 6,958 words · saved by 1 readers

Summary: Superposition-based interpretations of neural network activation spaces are incomplete. The specific locations of feature vectors contain cr…

x SAE feature geometry is outside the superposition hypothesis — LessWrong Interpretability (ML & AI) Apollo Research (org) AI Frontpage 229 SAE feature geometry is outside the superposition hypothesis by jake_mendel 24th Jun 2024 AI Alignment Forum 14 min read 18 229 Ω 101 Written at Apollo Research Summary: Superposition-based interpretations of neural network activation spaces are incomplete. The specific locations of feature vectors contain crucial structural information beyond superposition, as seen in circular arrangements of day-of-the-week features and in the rich structures of feature

Explore this link on the map →

related reading