2311.08576
arxiv.org · 8,358 words · saved by 1 readers
N/A
Towards Evaluating AI Systems for Moral Status Using Self-Reports Ethan Perez Robert Long New York University New York University perez@nyu.edu robert.long@nyu.edu arXiv:2311.08576v1 [cs.LG] 14 Nov 2023…
saved by
related reading
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- LLMs can learn about themselves by introspection — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- [2508.14802] Privileged Self-Access Matters for Introspection in AIarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- We should take AI welfare seriously - by Robert Longexperiencemachines.substack.com
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org