✳flâneur — a map of the web's best reading
Sean Lee
0 followers · 1 following · 88 views
Open this reading profile →
on the atlas — 25
- [2606.07082] On the Geometry of On-Policy Distillation1 savers
- Rethinking RL Infra for Agents | B'Log1 savers
- Untangling the Moons: A Visual History of Contrastive Learning1 savers
- [2605.12671] All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs1 savers
- [2605.15522] Stochastic Non-Smooth Convex Optimization with Unbounded Gradients1 savers
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation2 savers
- [1606.05312] Successor Features for Transfer in Reinforcement Learning1 savers
- Good QC for RL Data4 savers
- RL without TD learning1 savers
- Seohong Park on X: "We scaled up an "alternative" paradigm in RL: *divide and conquer*. Compared to Q-learning (TD learning), divide and conquer can naturally scale to much longer horizons. Blog post: https://t.co/xtXBzya0bI Paper: https://t.co/nqYkLucsWu ↓ https://t.co/XCdgUzaLxF" / X1 savers
- Aviral Kumar on X: "🚨🚨 New paper on flow-matching value functions Last year, we showed training RL value functions with a flow-matching loss achieved SOTA results. But why does it work? And what could it possibly tell us about other things that have nothing to do with VFs or even RL? Short https://t.co/x5trTotUO4" / X1 savers
- Clare Lyle | What's grokking good for?1 savers
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descent1 savers
- [2410.21265] Modular Duality in Deep Learning1 savers
- Clare Lyle | The state of plasticity in 20251 savers
- [2605.03327] DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment1 savers
- [2604.00626] A Survey of On-Policy Distillation for Large Language Models1 savers
- Interpreting Language Model Parameters4 savers
- [2605.01172] A Theory of Generalization in Deep Learning3 savers
- Plasticity as the Mirror of Empowerment (David Abel) - Sensorimotor AI Journal Club1 savers
- Illuminating the Three Dogmas of RL under Evolutionary Light (Mani Hamidi) - Sensorimotor AI Journal Club1 savers
- Woosh: A Sound Effects Foundation Model1 savers
- Curius / Onboarding2507 savers
- The World Inside Neural Networks6 savers
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishing3 savers