flâneur

How "Discovering Latent Knowledge in Language Models Without Supervision" Fits Into a Broader Alignment Scheme — LessWrong

lesswrong.com · saved by 3 readers

INTRODUCTION A few collaborators and I recently released a new paper: Discovering Latent Knowledge in Language Models Without Supervision. For a quick summary of our paper, you can check out this Twi…

saved by