flâneur — a map of the web's best reading

Thomas Larsen's Shortform — LessWrong

lesswrong.com · 1,242 words · saved by 2 readers

Comment by Thomas Larsen - Hypothesis: alignment-related properties of an ML model will be mostly determined by the part(s) of training that were most responsible for capabilities. If you take a very smart AI model with arbitrary goals/values and train it to output any particular sequence of tokens using SFT, it'll almost certainly work. So can we align an arbitrary model by training them to say "I'm a nice chatbot, I wouldn't cause any existential risk, ... "? Seems like obviously not, because the model will just learn the domain specific / shallow property of outputting those particular tokens in that particular situation. On the other hand, if you train an AI model from the ground up with a hypothetical "perfect reward function" that always gives correct ratings to the behaviour of the AI, (and you trained on a distribution of tasks similar to the one you are deploying it on) then I would guess that this AI, at least until around the human range, will behaviorally basically act according to the reward function. A related intuition pump here for the difference is the effect of training someone to say "I care about X" by punishing them until they say X consistently, vs raising them consistently with a large value set / ideology over time. For example, students are sometimes forced to write "I won't do X" or "I will do Y" 100 times, and usually this doesn't work at all. Similarly, randomly taking a single ethics class during high school usually doesn't cause people to enduringly act according to their stated favorite moral theory. However, raising your child Catholic, taking them to Catholic school, taking them to church, taking them to Sunday school, constantly talking to them about the importance of Catholic morality is in practice fairly likely to make them a pretty robust Catholic. There are maybe two factors being conflated above: (1) the fraction of training / upraising focused on goal X, and (2) the extent to which goal X was getting the capabilities. The reason why I think (2) is

x Thomas Larsen's Shortform — LessWrong Thomas Larsen's Shortform by Thomas Larsen 8th Nov 2022 AI Alignment Forum 1 min read 115 6 Ω 3 This is a special post for quick takes by Thomas Larsen . Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page . Rendering 0 / 115 comments, sorted by top scoring (show more) Click to highlight new comments since: Today at 12:00 PM Moderation Log More from Thomas Larsen View more Curated and popular this week 115 Comments 115 Mentioned in 25 Tracking (Expert/Influential) Predictions about AI Comment Perm

Explore this link on the map →

saved by

related reading