Actually, Othello-GPT Has A Linear Emergent World Representation — Neel Nanda
neelnanda.io · 19,725 words · saved by 4 readers
A write up of work extending and building on the paper Emergent World Representations
Actually, Othello-GPT Has A Linear Emergent World Representation Mar 28 Written By Neel Nanda Othello-GPT Epistemic Status : This is a write-up of an experiment in speedrunning research, and the core results represent ~20 hours/2.5 days of work (though the write-up took way longer). I'm confident in the main results to the level of " hot damn, check out this graph ", but likely have errors in some of the finer details. Disclaimer : This is a write-up of a personal project, and does not represent the opinions or work of my employer This post may get heavy on jargon. I recommend looking up unfam
saved by
related reading
- [2210.13382] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Structure and Interpretation of Deep Networkssidn.baulab.info
- Actually, Othello-GPT Has A Linear Emergent World Representation — AI Alignment Forumalignmentforum.org
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- Manipulating Chess-GPT’s World Model | Adam Karvonenadamkarvonen.github.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Large Language Model: world models or surface statistics?thegradient.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org