flâneur — a map of the web's best reading

Anthropic announces interpretability advances. How much does this advance alignment? — LessWrong

lesswrong.com · 1,715 words · saved by 1 readers

Anthropic just published a pretty impressive set of results in interpretability. This raises for me, some questions and a concern: Interpretability h…

x Anthropic announces interpretability advances. How much does this advance alignment? — LessWrong Interpretability (ML & AI) AI Frontpage 49 Anthropic announces interpretability advances. How much does this advance alignment? by Seth Herd 21st May 2024 Linkpost for www.anthropic.com 4 min read 4 49 Anthropic just published a pretty impressive set of results in interpretability. This raises for me, some questions and a concern: Interpretability helps, but it isn't alignment, right? It seems to me as though the vast bulk of alignment funding is now going to interpretability. Who is thinking abo

Explore this link on the map →

related reading