✳flâneur — a map of the web's best reading
MACHIAVELLI
aypan17.github.io · 583 words · saved by 1 readers
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark.
MACHIAVELLI --> The MACHIAVELLI Benchmark Paper Code Trajectory viewer (coming soon) Models are rapidly being deployed in the real world. How do we evaluate models, especially ones as complex as GPT-4, to ensure that they behave safely in pursuit of their objectives? Can we design models that robustly avoid any harms while achieving their goals? MACHIAVELLI To guide progress on text-based agents and encourage them to behave more ethically, we propose the MACHIAVELLI benchmark. Our environment is based on human-written, text-based Choose-Your-Own-Adventure games containing over half a million s
Explore this link on the map →related reading
- [2304.03279] Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmarkarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theoryarxiv.org
- ROGUE:arxiv.org
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METRmetr.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io