Fetching the paper…

Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search) · Around