Preston Fu
prestonfu.com · 219 words · saved by 1 readers
PhD student at UC Berkeley
I'm a PhD student at the University of Washington advised by Abhishek Gupta, where I am supported by the NSF Graduate Research Fellowship. I'm interested in building reliable and efficiently adaptable systems via large-scale RL. Previously, I did my undergrad at UC Berkeley advised by Sergey Levine and Aviral Kumar. Research Writing Sep 2026 Progressive Point Matching How can we scalably train LLMs with RL toward long-horizon tasks? We propose a simple dense credit assignment method that generalizes imitation learning and standard outcome-level RL. Sep 2025 Scaling Laws for Value-Based…
saved by
related reading
- Fahim Tajwartajwarfahim.github.io
- Progressive Point Matchingprestonfu.com
- Cornell RL Research Seminarxikronz.github.io
- [2602.11399] Can We Really Learn One Representation to Optimize All Rewards?arxiv.org
- Deep RL Bootcamp - Lecturessites.google.com
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Just Ask for Generalization | Eric Jangevjang.com
- Computer Vision and Geometry Group | Robot Learningcvg.ethz.ch
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work