flâneur — a map of the web's best reading

Chapter 4: Alignment Science - ARENA

learn.arena.education · 1,685 words · saved by 1 readers

Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. Most exercises in this chapter have dealt with LLMs at quite a low level of abstraction: as mechanisms to perform certain tasks (e.g. indirect object identification, in-context antonym learning, or algorithmic tasks like predicting legal Othello moves). But if we want to study characteristics that have alignment relevance, we need a higher level of abstraction. LLMs often exhibit "personas" that can shift unexpectedly, sometimes dramatically (see Sydney, Grok's "MechaHitler" persona, or AI-induced psychosis). These personalities are clearly shaped through training and prompting, but exactly why remains a

[4.4] LLM Psychology & Persona Vectors Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction Most exercises in this chapter have dealt with LLMs at quite a low level of abstraction: as mechanisms to perform certain tasks

Explore this link on the map →

related reading