flâneur — a map of the web's best reading

MACHIAVELLI

aypan17.github.io · 583 words · saved by 1 readers

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark.

MACHIAVELLI --> The MACHIAVELLI Benchmark Paper Code Trajectory viewer (coming soon) Models are rapidly being deployed in the real world. How do we evaluate models, especially ones as complex as GPT-4, to ensure that they behave safely in pursuit of their objectives? Can we design models that robustly avoid any harms while achieving their goals? MACHIAVELLI To guide progress on text-based agents and encourage them to behave more ethically, we propose the MACHIAVELLI benchmark. Our environment is based on human-written, text-based Choose-Your-Own-Adventure games containing over half a million s

Explore this link on the map →

related reading