✳flâneur — a map of the web's best reading
Understanding RL Vision
distill.pub · 9,098 words · saved by 1 readers
With diverse environments, we can analyze, diagnose and edit deep reinforcement learning models using attribution.
Understanding RL Vision Distill Understanding RL Vision With diverse environments, we can analyze, diagnose and edit deep reinforcement learning models using attribution. Observation (video game still) Positive attribution (good news) Negative attribution (bad news) Attribution from a hidden layer to the value function, showing what features of the observation (left) are used to predict success (middle) and failure (right). Applying dimensionality reduction (NMF) yields features that detect various in-game objects. Coin Enemy Buzzsaw Authors Affiliations Jacob Hilton OpenAI Nick Cammarata Open
Explore this link on the map →related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- Debugging Reinforcement Learning Systemsandyljones.com
- Distill — Latest articles about machine learningdistill.pub
- Feature Visualizationdistill.pub
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- RL Pet Peeves Part 1 · Aurielaurielws.github.io
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu