2212.03827
arxiv.org · 8,240 words · saved by 1 readers
N/A
D ISCOVERING L ATENT K NOWLEDGE IN L ANGUAGE M ODELS W ITHOUT S UPERVISION Collin Burns∗ Haotian Ye∗ Dan Klein Jacob Steinhardt UC Berkeley Peking University UC Berkeley UC Berkeley A BSTRACT Existing techniques for training language models can be…
saved by
related reading
- [2212.03827] Discovering Latent Knowledge in Language Models Without Supervisionarxiv.org
- Unsupervised Elicitation of Language Modelsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Unsupervised Elicitationalignment.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- How well do truth probes generalise? — LessWronglesswrong.com
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc