flâneur

when do closed-loop policies win? | flower and bolt

garden.binhph.am · saved by 1 readers

For a long time I've wondered why open-loop policies and action chunks dominate modern robotics. It's counter-intuitive for me that an action model would work best when executing actions blindly, with a small context and large action chunks. So last week, inspired by two recent writings, Behavioral cloning mystery and Revisiting Open-Loop Execution in Robotics, I took a stab at understanding the problem space myself and simply did some testing. In this article I'll try to answer one big question by answering a series of smaller ones: in which case do closed-loop policies win? Why this question matters to me: I've been interested in doing RL for robotics recently, and along the way I've seen the number of hacks the field has derived to accommodate action chunks. I'm a big fan of being exceptionally naive and stupid in answering questions (cough, the bitter lesson), so I take it personally when the answer to a problem is not elegant and requires hacks. We should just be able to RL modern

saved by