flâneur

Joe Van Valer

0 followers · 1 following · 237 views

on the atlas — 2

highlights — 2

  • From the last expression, it is clear that the update rule for AdaGrad adapts the step-size for each parameter j {\displaystyle j} accoding to η ( ϵ + G t ( j , j ) ) − 1 / 2 {\textstyle \eta (\epsilon +G_{t}^{(j,j)})^{-1/2}}, while standard sub-gradient methods have fixed step-size η {\displaystyle \eta } for every parameter.
    AdaGrad - Cornell University Computational Optimization Open Textbook - Optimization Wiki
  • Starting [at a young age] he’s read everything that he could find about business. The subject that interests him, he’s read newspapers, biographies, trade press. He went over to his grandfather who was a grocer and he read the progressive grocer magazine, and he read articles on how to stock a meat department... What he’s really done is he’s created this immense vertical filing cabinet in his brain of layers and layers and layers of files of information that he can draw back on now for more than 70 years worth of data.
    Curius / Onboarding