irmckenzie.co.uk/round2
At the end of the second and final round of the Inverse Scaling Prize, we’re awarding 7 more Third Prizes. The Prize aimed to identify important tasks on which language models (LMs) perform worse the larger they are (“inverse scaling”). Inverse scaling may reveal cases where LM training actively encourages behaviors that are misaligned with human preferences. The contest started on June 27th and concluded on October 27th, 2022 – thanks to everyone who participated! Across the two rounds, we had over 80 unique submissions and gave out a total of 11 Third Prizes. We are also accepting updates to two previous prize-winners (quote-repetition and redefine-math). For more details on the first round winners, see the Round 1 Announcement Post. We didn't find the kind of robust, major long-term-relevant problems that would have warranted a grand prize, but these submissions represent interesting tests of practically important issues and that help contribute to our scientific understanding of la
Inverse Scaling Prize: Second Round Winners At the end of t he second and final round of the Inverse Scaling Prize , we're awarding 7 more Third Prizes. The Prize aimed to identify important tasks on which language models (LMs) perform worse the larger they are ("inverse scaling"). Inverse scaling may reveal cases where LM training actively encourages behaviors that are misaligned with human preferences. The contest started on June 27th and concluded on October 27th , 2022 - thanks to everyone who participated! Across the two rounds, we had over 80 unique submissions and gave out a total of 11
related reading
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scalinggithub.com
- machine learning imindslice.substack.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Announcing the Inverse Scaling Prize ($250k Prize Pool) — AI Alignment Forumalignmentforum.org
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Scaling: The State of Play in AIoneusefulthing.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- 2308.03958arxiv.org
- Alex L. Zhangalexzhang13.github.io
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- larger language models may disappoint you [or, an eternally unfinished draft] — LessWronglesswrong.com
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org