✳flâneur — a map of the web's best reading
7029
Open this reading profile →on the atlas — 9
- [2506.14202] DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation1 savers
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learning1 savers
- Learning from Rare Success and Rich Feedback via Reflection-Enhanced Self-Distillation1 savers
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why1 savers
- Through the looking glass of benchmark hacking — Poolside1 savers
- [2605.06639] Recursive Agent Optimization1 savers
- [2512.21577] A Unified Definition of Hallucination: It's The World Model, Stupid!1 savers
- [2604.05273] Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext1 savers
- Is Frontier Asynchronous RL Solved? — Luke J. Huang4 savers