AI Isn't Coming for Your Mind. It's Coming Through It.
aletteraday.substack.com · 7,849 words · saved by 1 readers
Last week, I saw the following tweet from Anthropic:
Last week, I saw the following tweet from Anthropic: The tweet was part of a thread sharing a study Anthropic had conducted on agentic misalignment (specifically with regards to blackmail) and the new post-training methods they had used to address it. But what stuck out to me was a broader observation I have been thinking about for the past few years: If a specific harmful behavior can be traced to specific patterns in the training corpus, then less specific things, such as a model’s default narrative structure, are likely also corpus-shaped. The less specific something is, the harder it is…
saved by
related reading
- AI 2027ai-2027.com
- Teaching Claude Whyalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Teaching Claude why \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Cyborgism — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- AI #24: Week of the Podcast — LessWronglesswrong.com
- AI 2027ai-2027.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com