flâneur — a map of the web's best reading

A gentle introduction to sparse autoencoders

nickjiang.substack.com · 2,606 words · saved by 1 readers

Our biggest leap so far in deciphering how large language models work

A gentle introduction to sparse autoencoders Our biggest leap so far in deciphering how large language models work Nick Jiang Jul 03, 2024 6 Share DALLE3: a lab technician wielding a sledgehammer to break open a black box Sparse autoencoders (SAEs) are the current hot topic 🔥 in the interpretability world. In late May, Anthropic released a paper that shows how to use sparse autoencoders to effectively break down the internal reasoning of Claude 3 (Anthropic’s LLM) 1 . Shortly after, OpenAI published a paper successfully applying a similar procedure for GPT4. What’s exciting about SAEs is that

Explore this link on the map →

related reading