flâneur

Current alignment techniques might be ineffective (and actively bad) in the age of RL — LessWrong

lesswrong.com · saved by 4 readers

Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses: I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment. Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents: To be clear, these two incidents differ in various concrete details. However, both are clear examples of misalignment. In both cases, models took many actions which they had not been instructed to take and which violated ethical and legal norms. Note: Here, I am using alignment in the sense of "what type of behaviour would reasonable people want and expect given the available in

saved by