flâneur — a map of the web's best reading

instruction tuning and autoregressive distribution shift — LessWrong

lesswrong.com · 2,355 words · saved by 1 readers

[Note: this began life as a "Quick Takes" comment, but it got pretty long, so I figured I might as well convert it to a regular post.] …

x instruction tuning and autoregressive distribution shift — LessWrong AI Frontpage 41 instruction tuning and autoregressive distribution shift by nostalgebraist 5th Sep 2024 6 min read 5 41 [Note: this began life as a "Quick Takes" comment, but it got pretty long, so I figured I might as well convert it to a regular post.] In LM training , every token provides new information about "the world beyond the LM" that can be used/"learned" in-context to better predict future tokens in the same window. But when text is produced by autoregressive sampling from the same LM, it is not informative in th

Explore this link on the map →

related reading