flâneur — a map of the web's best reading

larger language models may disappoint you [or, an eternally unfinished draft] - LessWrong

lesswrong.com · 17,649 words · saved by 1 readers

WHAT THIS POST IS The following is an incomplete draft, which I'm publishing now because I am unlikely to ever finish writing it. I no longer fully endorse all the claims in the post. (In a few cases, I've added a note to say this explicitly.) However, there are some arguments in the post that I still endorse, and which I have not seen made elsewhere. This post is the result of me having lots of opinions about LM scaling, at various times in 2021, which were difficult to write down briefly or independently of one another. This post, originally written in July 2021, is the closest I got to writing them all down in one place. -nost, 11/26/21 -------------------------------------------------------------------------------- 0. CAVEAT This post will definitely disappoint you. Or, anyway, it will definitely disappoint me. I know that even though I haven't written it yet. My drafts folder contains several long, abandoned attempts to write (something like) this post. I've written (something like) this post many times in my head. I just can't seem to get it right, though. The drafts always sprawl out of control. So, if I can't do it right, why not do it wrong? Here's the disorganized, incomplete, brain-dump version of the better post I wish I were writing. Caveat lector. 1. POLARIZATION The topic of this post is large language models (LMs) like GPT-3. Specifically, what will happen as we make them larger and larger. By my lights, everyone else seems either too impressed/scared by the concept of LM scaling, or not impressed/scared enough. On LessWrong and related communities, I see lots of people worrying in earnest about whether the first superhuman AGI will be a GPT-like model. Both here and in the wider world, people often talk about GPT-3 like it's a far "smarter" being that it seems to me. On the other hand, the people who aren't scared often don't seem like they're even paying attention. Faced with a sudden leap in machine capabilities, they shrug. Faced wi

x larger language models may disappoint you [or, an eternally unfinished draft] — LessWrong Best of LessWrong 2021 GPT Language Models (LLMs) AI Curated 261 larger language models may disappoint you [or, an eternally unfinished draft] by nostalgebraist 26th Nov 2021 AI Alignment Forum 37 min read 31 261 Ω 61 what this post is The following is an incomplete draft, which I'm publishing now because I am unlikely to ever finish writing it. I no longer fully endorse all the claims in the post. (In a few cases, I've added a note to say this explicitly.) However, there are some arguments in the post

Explore this link on the map →

related reading