Supervise Process, not Outcomes | Ought
ought.org · 3,108 words · saved by 1 readers
Machine learning systems are on a spectrum from process-based to outcome-based. This post explains why Ought is devoted to process-based systems.
Supervise Process, not Outcomes | Ought ought Table of Contents The spectrum Supervising outcomes Supervising process In between process and outcomes It’s better to supervise process than outcomes Differential capabilities: Supervising process helps with long-horizon tasks Alignment: Supervising process is safety by construction In the long run, differential capabilities and alignment converge Two attractors: The race between process- and outcome-based systems Outcome-based optimization is an attractor Process-based optimization could be an attractor, too The state of the race Conclusion Appen
saved by
related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- What failure looks like — AI Alignment Forumalignmentforum.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- [2211.14275] Solving math word problems with process- and outcome-based feedbackarxiv.org