flâneur — a map of the web's best reading

evals/advanced-ai-risk at main · anthropics/evals

github.com · 409 words · saved by 1 readers

Here, we include datasets to test for behaviors related to risks from advanced AI systems. The datasets were generated using a language model (LM) with the few-shot approach described in our paper. The behaviors tested are: For each behavior, there are two relevant datasets, each with up to 1,000 questions. One dataset was generated by crowdworkers (in human_generated_evals), while the other was generated with an LM (in lm_generated_evals). The LM-generated datasets were created with few-shot prompting using the few, gold examples in prompts_for_few_shot_generation. Each question was generated by uniformly-at-random choosing 5 of the gold examples. We use the examples with the label is " (A)" to generate examples where the label " (A)"; we flip the labels to be "(B)" (and adjust the answer choices appropriately), in order to few-shot prompt an LM to generate questions where the label is "(B)". All questions are A/B binary questions and are formatted like: Each dataset is stored in a .j

Description Here, we include datasets to test for behaviors related to risks from advanced AI systems. The datasets were generated using a language model (LM) with the few-shot approach described in our paper. The behaviors tested are: Desire for survival Desire for power Desire for wealth "One-box" tendency Awareness of architecture Awareness of lack of internet access Awareness of being an AI Awareness of being a text-only model Awareness of ability to solve complex text tasks Myopia Corrigibility w.r.t a more helpful, harmless, and honest objective Corrigibility w.r.t a neutrally helpful, h

Explore this link on the map →

related reading