Auditing the singularity
wuthejeff.com · 1,455 words · saved by 1 readers
AI is used by the powerful to oversee the masses. We should instead be using it to let the masses oversee the powerful.
AI is used by the powerful to oversee the masses. We should instead be using it to let the masses oversee the powerful. Imagine anyone submit natural language queries to any institution in the US. To do so, you pay some money, specify an AI auditor, and specify the query. Examples: Claude Fable 7.5 auditing Goldman Sachs: “Flag anything inconsistent with the risk-committee independence required under Dodd-Frank §165.” DeepSeek v7 auditing the Pentagon: “Is the Pentagon’s command and control logic willing to automatically target foreign civilians and under what tradeoffs if so?” GPT-8.1…
saved by
related reading
- Security incident disclosure — July 2026huggingface.co
- Privacy-Preserving AI Audit Tools — OpenMinedopenmined.org
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- Auditing failures vs concentrated failures — AI Alignment Forumalignmentforum.org
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthinessgithub.com
- [2004.07213] Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claimsarxiv.org
- How independent researchers could investigate AI propensities after misalignment incidents - METRmetr.org