Confidence Regulation Neurons in Language Models
arxiv.org · 5,638 words · saved by 1 readers
N/A
Confidence Regulation Neurons in Language Models Alessandro Stolfo∗ Ben Wu∗ Wes Gurnee ETH Zürich University of Sheffield MIT Yonatan Belinkov Xingyi Song Mrinmaya Sachan Neel Nanda Technion University of Sheffield ETH Zürich arXiv:2406.16254v2 [cs.LG] 8…
saved by
related reading
- On Getting Confidence Estimates from Neural Networks | Bharath's notesbharathpbhat.github.io
- [1706.04599] On Calibration of Modern Neural Networksarxiv.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Language Modelinglena-voita.github.io
- interpreting GPT: the logit lens — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- A History of Large Language Modelsgregorygundersen.com
- How LLMs Actually Work | 0xkato0xkato.xyz
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub