✳flâneur — a map of the web's best reading
Aryaman Arora
aryaman.io · 774 words · saved by 1 readers
NLP Researcher
Ωhttps://github.com/aryamanarora/aryamanarora.github.io Aryaman Arora Aryaman Arora / Recruiting I'm Aryaman, a third-year Ph.D. student at Stanford University. I am looking for students who want to work on interpretability for language models! I'm most interested in coming up with new applications of interpretability to the entire LM training stack , towards the goal of improving models via better understanding. You might be a good fit for this if you are: excited about interpretability (regardless of prior experience in the area, or research in general!) have at least a little bit of backgro
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- ARENA - AI Safety Curriculumlearn.arena.education
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- On Optimism for Interpretabilitygoodfire.ai
- Neuronpedianeuronpedia.org