Jaival Patel
0 followers · 198 views
on the atlas — 4
- My Engineering Philosophy — Kaylee J. Edwards1 savers
- Canada Rocket Company2 savers
- Deep reinforcement learning for six degree-of-freedom planetary landing1 savers
- Curius / Onboarding2621 savers
highlights — 14
Creativity is what separates a good engineer from a great one.
My Engineering Philosophy — Kaylee J. EdwardsUsability, accessibility, security, and performance are not opposing forces. With the right questions, strong teams can solve for all of them.
My Engineering Philosophy — Kaylee J. EdwardsI realized: yes, we can.
My Engineering Philosophy — Kaylee J. EdwardsSolution-oriented engineers don’t create new problems - they meet them where they are.
My Engineering Philosophy — Kaylee J. EdwardsBecause we build software for people
My Engineering Philosophy — Kaylee J. Edwardspolicy gradient methods are more stable, and work with minimal hyperparameter tuning, and tend to perform better in sys- tems with high dimensional dynamics.
Deep reinforcement learning for six degree-of-freedom planetary landingPolicy gradient methods are much less sample efficient than methods based on Deep Q learning as they operate on pol- icy
Deep reinforcement learning for six degree-of-freedom planetary landingRecently, Schulman et al. (2015) have proven that policy gradient methods with stochastic policies can have monotonic improvement guarantees, pro- vided that the policy changes during optimization as mea- sured by the Kullback-Leibler (KL) divergence are constrained to be within certain bounds. The policy opti- mization problem is posed as a constrained optimization problem that ensures the KL divergence between policy updates remains within specified bounds. The authors used this result to develop the trust region policy optimization algorithm (TRPO
Deep reinforcement learning for six degree-of-freedom planetary landingDeep Q networks have proven effective at control tasks requiring the mapping of pixel level observations directly to control actions,
Deep reinforcement learning for six degree-of-freedom planetary landingalue function-based algorithms learn a mapping between a state-action tuple and the sum of the future dis- counted rewards received when starting in that specific state and taking that specific action
Deep reinforcement learning for six degree-of-freedom planetary landingalue function-based algorithms learn a mapping between a state-action tuple and the sum of the future dis- counted rewards received when starting in that specific state and taking that specific action.
Deep reinforcement learning for six degree-of-freedom planetary landingRL algorithms can be broken down into two major classes: value function methods and policy gradient meth- ods.
Deep reinforcement learning for six degree-of-freedom planetary landingFuture Mars missions will require advanced guidance, navigation, and control algorithms for the powered descent phase to tar- get specific surface locations and achieve pinpoint accuracy (landing error ellipse <5 m radius).
Deep reinforcement learning for six degree-of-freedom planetary landingStarting [at a young age] he’s read everything that he could find about business. The subject that interests him, he’s read newspapers, biographies, trade press. He went over to his grandfather who was a grocer and he read the progressive grocer magazine, and he read articles on how to stock a meat department... What he’s really done is he’s created this immense vertical filing cabinet in his brain of layers and layers and layers of files of information that he can draw back on now for more than 70 years worth of data.
Curius / Onboarding