Su P
0 followers · 149 views
on the atlas — 10
- Is Frontier Asynchronous RL Solved? — Luke J. Huang6 savers
- [2506.14202] DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation1 savers
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraints2 savers
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learning1 savers
- Learning from Rare Success and Rich Feedback via Reflection-Enhanced Self-Distillation1 savers
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why1 savers
- Through the looking glass of benchmark hacking — Poolside1 savers
- [2605.06639] Recursive Agent Optimization1 savers
- [2512.21577] A Unified Definition of Hallucination: It's The World Model, Stupid!1 savers
- [2604.05273] Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext1 savers