A Quick Tour of the Language Interpretability Tool
The Language Interpretability Tool (LIT) is a modular and extensible tool to interactively analyze and debug a variety of NLP models. LIT brings together common machine learning performance checks with interpretability methods specifically designed for NLP.
A Quick Tour of the Learning Interpretability Tool --> Learning Interpretability Tool Tutorials > Basics > Tour A Quick Tour of the Learning Interpretability Tool Follow along in the hosted demo. The Learning Interpretability Tool (LIT) is a modular and extensible tool to interactively analyze and debug a variety of NLP models. LIT brings together common machine learning performance checks with interpretability methods specifically designed for NLP. Building blocks - modules, groups, and workspaces Modules, groups, and workspaces form the building blocks of LIT. Modules are discrete windows in
Explore this link on the map →saved by
related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- Neuronpedianeuronpedia.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Langfuselangfuse.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com