flâneur — a map of the web's best reading

By Default, GPTs Think In Plain Sight - LessWrong

lesswrong.com · 5,805 words · saved by 1 readers

Comment by gwern - After watching how people use ChatGPT, and ChatGPT's weaknesses due to not using inner-monologue, I think I can be more concrete than pointing to non-robust features & CycleGAN [https://www.lesswrong.com/posts/jNTp87ioEWsZw3nar/gliders-in-language-models?commentId=fAjn3u5EnPX9rkb2z] (or the S1 'blob' [https://arxiv.org/abs/1912.04958#nvidia]) about why you should expect RLHF to put pressure towards developing steganographic encoding as a way to bring idle compute to bear on maximizing its reward. And further, this represents a tragedy of the commons where anyone failing to suppress steganographic encoding may screw it up for everyone else. -------------------------------------------------------------------------------- When people ask GPT-3 a hard multi-step question, it will usually answer immediately. This is because GPT-3 is trained on natural text, where usually a hard multi-step question is followed immediately by an answer; the most likely next token after 'Question?' is 'Answer.', it is not '[several paragraphs of tedious explicit reasoning]'. So it is doing a good job of imitating likely real text. Unfortunately, its predicted answer will often be wrong. This is because GPT-3 has no memory or scratchpad beyond the text context input, and it must do all the thinking inside one forward pass, but one forward pass is not enough thinking to handle a brandnew problem it has never seen before and has not already memorized an answer to or learned a strategy for answering. It is somewhat analogous to Memento [https://en.wikipedia.org/wiki/Memento_(film)]: at every forward pass, GPT-3 'wakes up' from amnesia not knowing anything, reads the notes on its hand and makes its best guesses, and tries to do... something. Fortunately, there is a small niche of text where the human has written 'Let's take this step by step' and it is then followed by a long paragraph of tedious explicit reasoning. If that is in the prompt, then GPT-3 can rejoice: it can simply write down t

x By Default, GPTs Think In Plain Sight — LessWrong GPT Interpretability (ML & AI) AI Frontpage 90 By Default, GPTs Think In Plain Sight by Fabien Roger 19th Nov 2022 AI Alignment Forum 11 min read 36 90 Ω 41 Epistemic status: Speculation with some factual claims in areas I’m not an expert in. Thanks to Jean-Stanislas Denain, Charbel-Raphael Segerie, Alexandre Variengien, and Arun Jose for helpful feedback on drafts, and thanks to janus, who shared related ideas. Main claims GPTs’ next-token-prediction process roughly matches System 1 (aka human intuition) and is not easily accessible, but GPT

Explore this link on the map →

related reading