Building Trustworthy AI Agents - Schneier on Security
View or Download in PDF Format The promise of personal AI assistants rests on a dangerous assumption: that we can trust systems we haven’t made trustworthy. We can’t. And today’s versions are failing us in predictable ways: pushing us to do things against our own best interests, gaslighting us with doubt about things we are or that we know, and being unable to distinguish between who we are and who we have been. They struggle with incomplete, inaccurate, and partial context: with no standard way to move toward accuracy, no mechanism to correct sources of error, and no accountability when wrong information leads to bad decisions...
View or Download in PDF Format The promise of personal AI assistants rests on a dangerous assumption: that we can trust systems we haven’t made trustworthy. We can’t. And today’s versions are failing us in predictable ways: pushing us to do things against our own best interests, gaslighting us with doubt about things we are or that we know, and being unable to distinguish between who we are and who we have been. They struggle with incomplete, inaccurate, and partial context: with no standard way to move toward accuracy, no mechanism to correct sources of error, and no accountability when…
saved by
related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI Integrity: Defending Against Backdoors and Secret Loyalties - Institute for AI Policy and Strategyiaps.ai
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- [2004.07213] Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claimsarxiv.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Off Target | CNAScnas.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com