RL-Enhanced 6DOF Rocket Landing - Jaival Patel
A eusable-booster landing benchmark with variable mass, quaternion attitude dynamics, thrust-vector control, classical baselines, flat PPO, hierarchical PPO, and hybrid residual reinforcement learning. Goal was to determine if a learned RL policy can reduce rocket landing dynamics successfully. Scope boundary: this is a research simulator and controller benchmark, not an exact vehicle, mission, aerodynamic, or flight-software reconstruction. Reusable booster landing is a useful control problem because it is simple to state and difficult to make honest. The vehicle must remove vertical energy, arrest lateral motion, maintain attitude, obey actuator limits, and do all of that while its mass changes during powered descent. A reinforcement-learning benchmark that removes variable mass, attitude dynamics, disturbances, or lateral-attitude coupling can look impressive while avoiding the hard part of the problem. This article presents the current state of a Falcon 9-inspired 6DOF landing simu
Phase 2C hybrid PPO nominal rollout. Source: outputs/phase2c_hybrid_rl/eval_nominal_seed7_50k_stopping_floor_v1/trajectory.csv and metrics.json. This is a saved benchmark artifact, not a recreated operational flight profile. Abstract Reusable booster landing is a useful control problem because it is simple to state and difficult to make honest. The vehicle must remove vertical energy, arrest lateral motion, maintain attitude, obey actuator limits, and do all of that while its mass changes during powered descent. A reinforcement-learning benchmark that removes variable mass, attitude dynamics,
Explore this link on the map →saved by
related reading
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Precise Manipulation with Efficient Online RLpi.website
- Learning Beyond Gradientstrinkle23897.github.io
- State of Robot Learning, December 2025vedder.io
- pistar06.pdfpi.website
- FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learningarxiv.org
- FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learningjasonjzliu.com
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- A VLA that Learns from Experiencepi.website
- Underactuated Roboticsunderactuated.mit.edu