✳flâneur — a map of the web's best reading
Position: It's Time to Optimize for Self-Consistency
time-for-consistency.github.io · 11,969 words · saved by 1 readers
N/A
# link_axau4agk3j.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas - CreationDate=D:20260305233138Z - Creator=LaTeX with hyperref - Keywords=Machine Learning, ICML - ModDate=D:20260305233138Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.26 (TeX Live 2024) kpathsea version 6.4.0 - Producer=pdfTeX-1.40.26 - Subject=Proceedings o
Explore this link on the map →saved by
related reading
- Self-Adapting Language Modelsarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- [2412.00543] Evaluating the Consistency of LLM Evaluatorsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- Alignment faking in large language modelsarxiv.org
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org