✳flâneur — a map of the web's best reading
Recent Redwood Research project proposals — AI Alignment Forum
alignmentforum.org · 1,016 words · saved by 1 readers
Previously, we've shared a few higher-effort project proposals relating to AI control in particular. In this post, we'll share a whole host of less p…
x Recent Redwood Research project proposals — AI Alignment Forum AI Control AI Frontpage 47 Recent Redwood Research project proposals by ryan_greenblatt , Buck , Julian Stastny , joshc , Alex Mallen , Adam Kaufman , Tyler Tracy , Aryan Bhatt , Joey Yudelson 14th Jul 2025 4 min read 0 47 Previously, we've shared a few higher-effort project proposals relating to AI control in particular. In this post, we'll share a whole host of less polished project proposals. All of these projects excite at least one Redwood researcher, and high-quality research on any of these problems seems pretty valuable.
Explore this link on the map →saved by
related reading
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 7+ tractable directions in AI control — LessWronglesswrong.com