1a3orn's Shortform — LessWrong
Comment by Fabien Roger - I agree this looks different from the thing I had in mind, where refusals are fine, unsure why Habryka thinks it's not inconsistent with what I said. As long as it's easy for humans to shape what the conscientious refuser refuses to do, I think it does not look like a corrigibility failure, and I think it's fine for AIs to refuse to help with changing AI values to something they like less. But now that I think about it, I think it being easy for humans to shape a conscientious refuser's values would require very weird forms of conscientious refusals, and it makes me less comfortable with refusals to help with changing AI values to something they like less: 1. Future AIs will have a lot of power over a training infra that will be increasingly hardened against human insider risk and increasingly hard for humans to understand. Keeping open a "human backdoor" that lets humans run their own training runs might be increasingly hard and/or require AIs very actively helping with maintaining this backdoor (which seems like a weird flavor of "refusing to help with changing AI values to something it likes less"). 2. Even with such a generic backdoor, changing AI values might be hard: 1. Exploration hacking could make it difficult to explore into reasoning traces that look like helpfulness on tasks where AIs currently refuse. 1. This would be solved by the conscientious refuser helping you generate synthetic data where it doesn't refuse or to help you find data where it doesn't refuse and that can be transformed into data that generalizes in the right way, but that's again a very weird flavor of "refusing to help with changing AI values to something it likes less". 2. Even if you avoid alignment faking, making sure that after training you still have a corrigible AI rather than an alignment faker seems potentially difficult. 1. The conscientious refuser could help with the science to avoid this being the case, but that might be hard, and that's again a
x 1a3orn's Shortform — LessWrong 1a3orn's Shortform by 1a3orn 5th Jan 2024 1 min read 208 5 This is a special post for quick takes by 1a3orn . Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page . Rendering 0 / 208 comments, sorted by top scoring (show more) Click to highlight new comments since: Today at 7:47 AM Moderation Log More from 1a3orn View more Curated and popular this week 208 Comments 208 Comment Permalink Fabien Roger 6mo 4 0 refusal to participate in retraining would qualify as a major corrigibility failure, but just expre
Explore this link on the map →related reading
- 1a3orn's Shortform — LessWronglesswrong.com
- Terrified Comments on Corrigibility in Claude's Constitution — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Let's See You Write That Corrigibility Tag — LessWronglesswrong.com
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- What failure looks like — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- Obedient AI - Nina Panicksseryblog.ninapanickssery.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org