flâneur — a map of the web's best reading

Alignment likely generalizes further than capabilities.

beren.io · 1,143 words · saved by 1 readers

Recently, I was reading this paper which demonstrates how to do online RLHF for alignment of LLMs and a sentence stuck out to me: We conjecture that this is because the reward model (discriminator) usually generalizes better than the policy (generator) This is an offhand remark but it strikes at...

Recently, I was reading this paper which demonstrates how to do online RLHF for alignment of LLMs and a sentence stuck out to me: We conjecture that this is because the reward model (discriminator) usually generalizes better than the policy (generator) This is an offhand remark but it strikes at the core of the LessWrong doom story, as propounded here for which the idea of ‘capabilities generalizing further than alignment’ is central. This is a key point because it is necessary for the self-improvmeent of AI systems, even if initially ‘aligned’ using current methods, to spell doom because even

Explore this link on the map →

related reading