Rohan Kathuria
0 followers · 240 views
on the atlas — 5
- LLMs reward expertise8 savers
- A short note on interpretability and minds3 savers
- davidbau.com In Defense of Curiosity8 savers
- Dario Amodei — The Adolescence of Technology44 savers
- Julia10 savers
highlights — 5
into a mathematical science of Mind will be the scientific legacy of this century
A short note on interpretability and mindswill turn out, in the end, to have been studying the same object from different sides.
A short note on interpretability and mindsIf the computation that language models do when predicting our text can be decomposed into simpler parts, perhaps human minds and human thinking, the process that generated that text, can be similarly decomposed.
A short note on interpretability and mindsI wish that the field had some answer to whether there is any notion of a "right" decomposition for a network
A short note on interpretability and mindsHow is that computation different in an LLM that gets 1.8 nats of cross-entropy?
A short note on interpretability and minds