flâneur — a map of the web's best reading

How can LLM RL Work Despite Information-Theoretic Inefficiency

beren.io · 3,833 words · saved by 1 readers

Epistemic Status: Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which they argue that RL is extremely informationally inefficient compared to pretraining, that it can only impart a tiny amount of bits,...

Epistemic Status : Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which they argue that RL is extremely informationally inefficient compared to pretraining , that it can only impart a tiny amount of bits, and that it can only bring out behaviours that are already in the base models etc. The basic intuition here is extremely obvious and essentially falls immediately out of the formulation of the two methods. Pretraining (and SFT etc) compute a loss on every token so every

Explore this link on the map →

related reading