✳flâneur — a map of the web's best reading
Finite Difference Flow Optimization
mcallisterdavid.com · 3,060 words · saved by 1 readers
RL Post-Training for Flow-Based Image Generators
Finite Difference Flow Optimization Finite Difference Flow Optimization RL Post-Training for Flow-Based Image Generators Epoch 0 Epoch 50 Results from before and after our human preference RL post-training. No CFG is used in either image. David McAllister Miika Aittala Tero Karras Janne Hellsten Angjoo Kanazawa Timo Aila Samuli Laine Paper arXiv Code In this project, we set out to find a simple, grounded RL post-training method for diffusion image generators. We made a few observations about the structure of diffusion flows that lead to Finite Difference Flow Optimization (FDFO), a new online
Explore this link on the map →saved by
related reading
- Flow Matching Policy Gradientsflowreinforce.github.io
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- ⭐️ Diffusion Modelsandrewkchan.dev
- Learning the integral of a diffusion model – Sander Dielemansander.ai
- How to Generate Text in One Stepone-step-lm.github.io
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Diffusion Meets Flow Matchingdiffusionflow.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Yang Songyang-song.net
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Diffusion models from scratchchenyang.co