flâneur — a map of the web's best reading

OthelloGPT learned a bag of heuristics — LessWrong

lesswrong.com · 5,177 words · saved by 1 readers

Work performed as a part of Neel Nanda's MATS 6.0 (Summer 2024) training program. …

x OthelloGPT learned a bag of heuristics — LessWrong MATS Program Interpretability (ML & AI) AI Frontpage 111 OthelloGPT learned a bag of heuristics by jylin04 , JackS , Adam Karvonen , Can 2nd Jul 2024 AI Alignment Forum 11 min read 10 111 Ω 43 Work performed as a part of Neel Nanda's MATS 6.0 (Summer 2024) training program. TLDR This is an interim report on reverse-engineering Othello-GPT , an 8-layer transformer trained to take sequences of Othello moves and predict legal moves. We find evidence that Othello-GPT learns to compute the board state using many independent decision rules that ar

Explore this link on the map →

saved by

related reading