Fetching the paper…

AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights · Around