Announcing the Inverse Scaling Prize ($250k Prize Pool) — AI Alignment Forum
TL;DR: We’re launching the Inverse Scaling Prize: a contest with $250k in prizes for finding zero/few-shot text tasks where larger language models show increasingly undesirable behavior (“inverse scaling”). We hypothesize that inverse scaling is often a sign of an alignment failure and that more examples of alignment failures would benefit empirical alignment research. We believe that this contest is an unusually concrete, tractable, and safety-relevant problem for engaging alignment newcomers and the broader ML community. This post will focus on the relevance of the contest and the inverse scaling framework to longer-term AGI alignment concerns. See our GitHub repo for contest details, prizes we’ll award, and task evaluation criteria. Recent work has found that Language Models (LMs) predictably improve as we scale LMs in various ways (“scaling laws”). For example, the test loss on the LM objective (next word prediction) decreases as a power law with compute, dataset size, and model si
x Announcing the Inverse Scaling Prize ($250k Prize Pool) — AI Alignment Forum Bounties (closed) Inner Alignment Language Models (LLMs) Outer Alignment Scaling Laws AI Community Frontpage 59 Announcing the Inverse Scaling Prize ($250k Prize Pool) by Ethan Perez , Ian McKenzie , Sam Bowman 27th Jun 2022 8 min read 14 59 TL;DR : We’re launching the Inverse Scaling Prize : a contest with $250k in prizes for finding zero/few-shot text tasks where larger language models show increasingly undesirable behavior (“inverse scaling”). We hypothesize that inverse scaling is often a sign of an alignment fa
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- irmckenzie.co.uk/round2irmckenzie.co.uk
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaobmk.sh
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com