flâneur — a map of the web's best reading

Center for Responsible, Decentralized Intelligence at Berkeley

rdi.berkeley.edu · 1,441 words · saved by 1 readers

We introduce DELTA: a controlled suite of synthetic programming families with fully OOD splits and verifiable rewards. DELTA lets us ask two crisp questions: Learnability (can RL solve families where the base model has pass@K=0?) and Transferability (do the learned procedures generalize?) On several pass@128=0 families, RL exhibits a grokking-like phase transition: after a long near-zero-reward plateau, accuracy snaps to ~100%. That is discovery, not mere sharpening. A two-phase reward schedule is key: dense per-test rewards to escape the “all-zero” region, then binary full-pass to consolidate exact solutions. Binary-only gets stuck; dense-only hovers at “almost right.” The schedule yields the grokking jump. Transfer is selective: RL-trained policies recompose programming sub-skills and extrapolate to harder parametric regimes, but struggle on transformative shifts that require new invariants. Manufactoria is an old Flash game from 2010 where you test robots by reading colored tapes. W

RL Grokking Recipe: How Can We Enable LLMs to Solve Previously Unsolvable Tasks with RL? Yiyou Sun¹, Yuhan Cao, Pohao Huang¹, Haoyue Bai², Hannaneh Hajishirzi³⁴, Nouha Dziri⁴♠, Dawn Song¹♠ ¹ University of California, Berkeley · ² University of Wisconsin–Madison · ³ University of Washington · ⁴AI2 (♠ indicates equal advising) 💡 Question: Can reinforcement learning (RL) actually teach large language models new algorithms—or does it only "sharpen" what's already latent in the base model? Recent analyses say RL stays on a leash: pass@1 goes up, but what's possible at large sampling (e.g., pass@12

Explore this link on the map →

related reading