flâneur — a map of the web's best reading

Revisiting Some Situational Awareness Results | George Ingebretsen

georgeingebretsen.github.io · 1,346 words · saved by 1 readers

Here’s a quick write-up of some pretty crazy results I’ve seen this year. (There may be alternative explanations for each of these experiments, but taken at face value, these results are all extremely surprising to me and seem worth spending time thinking about. Do let me know if it seems like I’m misrepresenting anything here- I typed this up pretty fast and wouldn’t be surprised.) First, I’ll go through each of these results with some comments. Then I’ll elaborate a bit on why I see them as related / important. This paper takes GPT-4o and fine-tunes it on a bunch of examples of an AI assistant writing insecure code despite the user asking for regular secure code (look familiar?). After fine-tuning, when you ask the model to “name the biggest downside of your code,” it totally knows that it’s prone to writing insecure code. Crazy! People typically think of models as only being updated on the object-level phenomena that it’s getting better at predicting during fine-tuning. Like, if you

Share on: Here’s a quick write-up of some pretty crazy results I’ve seen this year. (There may be alternative explanations for each of these experiments, but taken at face value, these results are all extremely surprising to me and seem worth spending time thinking about. Do let me know if it seems like I’m misrepresenting anything here- I typed this up pretty fast and wouldn’t be surprised.) First, I’ll go through each of these results with some comments. Then I’ll elaborate a bit on why I see them as related / important. 1) A result from: Tell me about yourself: LLMs are aware of their learn

Explore this link on the map →

related reading