flâneur

A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models

arxiv.org · 6,328 words · saved by 2 readers

N/A

A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models Dong Shu1,† , Xuansheng Wu2,† , Haiyan Zhao3,† , Daking Rai4 , Ziyu Yao4 , Ninghao Liu2 , Mengnan Du3 1 Northwestern University 2 University of Georgia…

saved by

related reading