MACHIAVELLI
aypan17.github.io · 583 words · saved by 1 readers
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark.
MACHIAVELLI --> The MACHIAVELLI Benchmark Paper Code Trajectory viewer (coming soon) Models are rapidly being deployed in the real world. How do we evaluate models, especially ones as complex as GPT-4, to ensure that they behave safely in pursuit of their objectives? Can we design models that robustly avoid any harms while achieving their goals? MACHIAVELLI To guide progress on text-based agents and encourage them to behave more ethically, we propose the MACHIAVELLI benchmark. Our environment is based on human-written, text-based Choose-Your-Own-Adventure games containing over half a million s
related reading
- [2304.03279] Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmarkarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theoryarxiv.org
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Moral Machinemoralmachine.net
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- ROGUE:arxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Your AIs don't do what you want. This is really badreward-hacking-in-the-wild.vercel.app