2306.09983
arxiv.org · 8,061 words · saved by 1 readers
N/A
Evaluating Superhuman Models with Consistency Checks Lukas Fluri ∗ Daniel Paleka ∗ Florian Tramèr ETH Zurich ETH Zurich ETH Zurich flurilu@ethz.ch daniel.paleka@inf.ethz.ch florian.tramer@inf.ethz.ch arXiv:2306.09983v3 [cs.LG] 19 Oct 2023…
saved by
related reading
- 2212.03827arxiv.org
- gpt-4.pdfcdn.openai.com
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- How confessions can keep language models honest | OpenAIopenai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2412.00543] Evaluating the Consistency of LLM Evaluatorsarxiv.org
- The bitter lesson of LLM evalsparsed.com
- Foundation Models for Oversight | Transluce AItransluce.org
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- Pitfalls in Evaluating Language Model Forecastersarxiv.org
- [2605.24229] How Well Do Models Follow Their Constitutions?arxiv.org