✳flâneur — a map of the web's best reading
2312.06942
arxiv.org · 20,628 words · saved by 3 readers
N/A
# link_jewx9ybwzx.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20240724002503Z - Creator=LaTeX with hyperref - ModDate=D:20240724002503Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 AI CONTROL: IMPROVING SAFETY DESPITEINTENTIONAL SUBVERSIONRyan Greenblatt∗ Buck Shlegeris Kshitij Sachan Fabien RogerRedwood ResearchABSTRACTAs larg
Explore this link on the map →saved by
related reading
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionar5iv.labs.arxiv.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- gpt-4.pdfcdn.openai.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- How to prevent collusion when using untrusted models to monitor each other — AI Alignment Forumalignmentforum.org