listend to a podcast on goodfire.ai. a few good resources on mechanistic interpretability Key Research Papers - Tracing the thoughts of a large language model (