evals/advanced-ai-risk at main · anthropics/evals
Here, we include datasets to test for behaviors related to risks from advanced AI systems. The datasets were generated using a language model (LM) with the few-shot approach described in our paper. The behaviors tested are: For each behavior, there are two relevant datasets, each with up to 1,000 questions. One dataset was generated by crowdworkers (in human_generated_evals), while the other was generated with an LM (in lm_generated_evals). The LM-generated datasets were created with few-shot prompting using the few, gold examples in prompts_for_few_shot_generation. Each question was generated by uniformly-at-random choosing 5 of the gold examples. We use the examples with the label is " (A)" to generate examples where the label " (A)"; we flip the labels to be "(B)" (and adjust the answer choices appropriately), in order to few-shot prompt an LM to generate questions where the label is "(B)". All questions are A/B binary questions and are formatted like: Each dataset is stored in a .j
Description Here, we include datasets to test for behaviors related to risks from advanced AI systems. The datasets were generated using a language model (LM) with the few-shot approach described in our paper. The behaviors tested are: Desire for survival Desire for power Desire for wealth "One-box" tendency Awareness of architecture Awareness of lack of internet access Awareness of being an AI Awareness of being a text-only model Awareness of ability to solve complex text tasks Myopia Corrigibility w.r.t a more helpful, harmless, and honest objective Corrigibility w.r.t a neutrally helpful, h
Explore this link on the map →related reading
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Cookbookcookbook.openai.com
- Contra Labs - Powered by Contracontralabs.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com
- MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METRmetr.org
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMsarxiv.org
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org