Revisiting Some Situational Awareness Results | George Ingebretsen
Here’s a quick write-up of some pretty crazy results I’ve seen this year. (There may be alternative explanations for each of these experiments, but taken at face value, these results are all extremely surprising to me and seem worth spending time thinking about. Do let me know if it seems like I’m misrepresenting anything here- I typed this up pretty fast and wouldn’t be surprised.) First, I’ll go through each of these results with some comments. Then I’ll elaborate a bit on why I see them as related / important. This paper takes GPT-4o and fine-tunes it on a bunch of examples of an AI assistant writing insecure code despite the user asking for regular secure code (look familiar?). After fine-tuning, when you ask the model to “name the biggest downside of your code,” it totally knows that it’s prone to writing insecure code. Crazy! People typically think of models as only being updated on the object-level phenomena that it’s getting better at predicting during fine-tuning. Like, if you
Share on: Here’s a quick write-up of some pretty crazy results I’ve seen this year. (There may be alternative explanations for each of these experiments, but taken at face value, these results are all extremely surprising to me and seem worth spending time thinking about. Do let me know if it seems like I’m misrepresenting anything here- I typed this up pretty fast and wouldn’t be surprised.) First, I’ll go through each of these results with some comments. Then I’ll elaborate a bit on why I see them as related / important. 1) A result from: Tell me about yourself: LLMs are aware of their learn
Explore this link on the map →related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- gpt-4.pdfcdn.openai.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Alignment faking in large language modelsarxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Alignment will happen by default. What’s next? — LessWronglesswrong.com