Position: It's Time to Optimize for Self-Consistency
time-for-consistency.github.io · 11,969 words · saved by 2 readers
N/A
# link_axau4agk3j.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas - CreationDate=D:20260305233138Z - Creator=LaTeX with hyperref - Keywords=Machine Learning, ICML - ModDate=D:20260305233138Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.26 (TeX Live 2024) kpathsea version 6.4.0 - Producer=pdfTeX-1.40.26 - Subject=Proceedings o
saved by
related reading
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- Self-Adapting Language Modelsarxiv.org
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- [2412.00543] Evaluating the Consistency of LLM Evaluatorsarxiv.org
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- 2401.10020.pdfarxiv.org
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com