✳flâneur — a map of the web's best reading
[AN #70]: Agents that help humans who are still learning about their own preferences — LessWrong
lesswrong.com · 2,933 words · saved by 1 readers
Find all Alignment Newsletter resources here. In particular, you can sign up, or look through this spreadsheet of all summaries that have ever been i…
x [AN #70]: Agents that help humans who are still learning about their own preferences — LessWrong Alignment Newsletter Newsletters AI Frontpage 16 [AN #70]: Agents that help humans who are still learning about their own preferences by Rohin Shah 23rd Oct 2019 Alignment Newsletter AI Alignment Forum 11 min read 0 16 Ω 14 Find all Alignment Newsletter resources here . In particular, you can sign up , or look through this spreadsheet of all summaries that have ever been in the newsletter. I'm always happy to hear feedback; you can send it to me by replying to this email. Audio version here (may
Explore this link on the map →saved by
related reading
- The Era of Experience Paper.pdfstorage.googleapis.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- State of Robot Learning, December 2025vedder.io
- Teaching Claude why \ Anthropicanthropic.com
- Learning through human feedback — Google DeepMinddeepmind.google
- Relaxed adversarial training for inner alignment — AI Alignment Forumalignmentforum.org
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com