flâneur

AI Isn't Coming for Your Mind. It's Coming Through It.

aletteraday.substack.com · 7,849 words · saved by 1 readers

Last week, I saw the following tweet from Anthropic:

Last week, I saw the following tweet from Anthropic: The tweet was part of a thread sharing a study Anthropic had conducted on agentic misalignment (specifically with regards to blackmail) and the new post-training methods they had used to address it. But what stuck out to me was a broader observation I have been thinking about for the past few years: If a specific harmful behavior can be traced to specific patterns in the training corpus, then less specific things, such as a model’s default narrative structure, are likely also corpus-shaped. The less specific something is, the harder it is…

saved by

related reading