Gradient Hacking is extremely difficult.
Epistemic Status: Originally started out as a comment on this post but expanded enough to become its own post. My view has been formed by spending a reasonable amount of time trying and failing to construct toy gradient hackers by hand, but this could just reflect me being insufficiently creative...
Epistemic Status : Originally started out as a comment on this post but expanded enough to become its own post. My view has been formed by spending a reasonable amount of time trying and failing to construct toy gradient hackers by hand, but this could just reflect me being insufficiently creative or thoughtful rather than the intrinsic difficulty of the problem. There has been a lot of discussion recently about gradient hackers as a potentially important class of mesaoptimizers. The idea of gradient hackers is that they are some malign subnetwork that exists in a larger network that steer the
Explore this link on the map →related reading
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- The Little Book of Deep Learningfleuret.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Gradient descent - Wikipediaen.wikipedia.org
- A Visual Explanation of Gradient Descent Methods (Momentum, AdaGrad, RMSProp, Adam) | Towards Data Sciencetowardsdatascience.com
- Why Momentum Really Worksdistill.pub
- Learning Beyond Gradientstrinkle23897.github.io
- Hypothesis: gradient descent prefers general circuits — LessWronglesswrong.com
- NL.pdfabehrouz.github.io
- microgptkarpathy.github.io
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai