[2502.13410] Tell Me Why: Incentivizing Explanations
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2502.13410] Tell Me Why: Incentivizing Explanations Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computer Science and Game Theory arXiv:2502.13410 (cs) [Submitted on 19 Feb 2025] Title: Tell Me Why: Incentivizing Explanations Authors: Siddarth Srinivasan , Ezra Karger , Michiel Bakker , Yiling Chen View a PDF of the paper titled Tell Me Why: Incentivizing Explanations, by Siddarth Srinivasan and 3 other authors View PDF HTML (experimental) Abstract: Common sense suggests that w
Explore this link on the map →related reading
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Rationalization — LessWronglesswrong.com
- [1902.09469] Embedded Agencyarxiv.org
- Natural Deception with RL - Rajan Agarwalrajan.sh
- Bengt Holmström - Prize Lecture_ Pay for Performance and Beyond.pdfeconomics.mit.edu
- Unconscious Economics — LessWronglesswrong.com
- Robust Cooperation in the Prisoner's Dilemma — LessWronglesswrong.com
- Irrationality as a Defense Mechanism for Reward-hacking — LessWronglesswrong.com
- Books — LessWronglesswrong.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Logical decision theories — LessWrongarbital.com
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com