Ji-Ha's Blog
jiha-kim.github.io · 377 words · saved by 2 readers
Ji-Ha Kim’s blog for sharing things.
Post When Equivalent Weights Train Differently Why coordinate-level optimizers can behave differently on weights that represent the same model, and how quotient-aware updates remove the hidden gauge. May 6, 2026 Machine Learning , Mathematical Optimization Fast Tight Spectral-Norm Bounds A GPU-friendly algorithm for tight certified singular-value endpoint bounds using scaled Gram matrices, SYRKs, Frobenius reductions, and a small scalar moment problem. Apr 28, 2026 Numerical Linear Algebra , Matrix Computations Autoregression vs Diffusion - Understanding Sampling via Optimal…
saved by
related reading
- Deriving Muonjeremybernste.in
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- [2502.01131] Simple Linear Neuron Boostingarxiv.org
- [1512.04202] Preconditioned Stochastic Gradient Descentarxiv.org
- Why Momentum Really Worksdistill.pub
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogkellerjordan.github.io
- Online KL Shampoo | Tildeblog.tilderesearch.com
- Understanding Muonlakernewhouse.com
- The Little Book of Deep Learningfleuret.org
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- [2601.08393] Controlled LLM Training on Spectral Spherearxiv.org