A basic systems architecture for AI agents that do autonomous research — LessWrong
A lot of threat models describing how AIs might escape our control (e.g. self-exfiltration, hacking the datacenter) start out with AIs that are actin…
x A basic systems architecture for AI agents that do autonomous research — LessWrong Redwood Research AI Curated 190 A basic systems architecture for AI agents that do autonomous research by Buck 23rd Sep 2024 AI Alignment Forum 9 min read 17 190 Ω 71 A lot of threat models describing how AIs might escape our control (e.g. self-exfiltration , hacking the datacenter ) start out with AIs that are acting as agents working autonomously on research tasks (especially AI R&D) in a datacenter controlled by the AI company. So I think it’s important to have a clear picture of how this kind of AI agent c
Explore this link on the map →saved by
related reading
- AI 2027ai-2027.com
- How can we solve diffuse threats like research sabotage with AI control?blog.redwoodresearch.org
- How might we safely pass the buck to AI? — LessWronglesswrong.com
- Fields that I reference when thinking about AI takeover prevention — LessWronglesswrong.com
- Danger, AI Scientist, Danger - by Zvi Mowshowitzthezvi.substack.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI 2027ai-2027.com
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org
- The Rogue Replication Threat Model - METRmetr.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Self-exfiltration is a key dangerous capabilityaligned.substack.com