AI Safety | Arkose
victoriabrook.github.io · 1,173 words · saved by 1 readers
AI Safety Resources
AI Safety | Arkose Learn about large-scale risks from advanced AI Selected Papers --> --> --> Overviews What types of risks from advanced AI might we face? These papers provide an overview of anticipated problems and relevant technical research directions. (Ngo et al., 2022) The Alignment Problem from a Deep Learning Perspective (Hendrycks et al., 2023) An Overview of Catastrophic AI Risks (Chan et al., 2023) Harms from Increasingly Agentic Algorithmic Systems (Anwar et al., 2024) Foundational Challenges in Assuring Alignment and Safety of Large Language Models Dangerous Capability Evaluations
saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI safety - Wikipediaen.wikipedia.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Shallow review of technical AI safety, 2024 — LessWronglesswrong.com
- A Summary of Recent Work (July 2026)gdmalignment.substack.com
- Spring 2026 Projects - SPARsparai.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Thoughts on AI safety – Windows On Theorywindowsontheory.org
- Resources on the AI risk landscape – Ben Pomeranzbenpomeranz.com