flâneur — a map of the web's best reading

kitft/natural_language_autoencoders ·

github.com · 1,251 words · saved by 1 readers

No description, website, or topics provided.

Natural Language Autoencoders (NLA) Open-source library accompanying the Anthropic Transformer Circuits post Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations . 📄 Blog post · ▶ Video walkthrough · 🔬 Try the released NLAs on Neuronpedia A Natural Language Autoencoder is a pair of fine-tuned LMs that map residual-stream activation vectors to natural language and back: direction mechanism AV (activation verbalizer) vector → text inject the vector as a single token embedding into a fixed prompt, autoregress a description AR (activation reconstructor) text → vecto

Explore this link on the map →

related reading