Fetching the paper…
Reading the bibliography…
In previous literature, backward error analysis was used to find ordinary differential equations (ODEs) approximating the gradient descent trajectory.
Geometric numerical integration
Ernst Hairer, C. L. and Wanner, G · 2006
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T., Hinton, G., et al · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deeply-supervised nets
Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Earlier work this paper cites.
Improving generalization performance by switching from adam to sgd
Keskar, N. S. and Socher, R · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Li, Q., Tai, C., and E, W · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Earlier work this paper cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Chen, J., Zhou, D., Tang, Y., Yang, Z., Cao, Y., and Gu, Q · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2019
Earlier work this paper cites.
The implicit bias of adagrad on separable data
Qian, Q. and Qian, X · 2019
Earlier work this paper cites.
The geometry of sign gradient descent
Balles, L., Pedregosa, F., and Roux, N. L · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Multiscale analysis of accelerated gradient methods
Farazmand, M · 2020
Cited alongside, same era.
Granziol, D · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ji, Z. and Telgarsky, M · 2020
Cited alongside, same era.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Kunin, D., Sagastuy-Brena, J., Ganguli, S., Yamins, D. L., and Tanaka, H · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Better plain vit baselines for imagenet-1k
Beyer, L., Zhai, X., and Kolesnikov, A · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Adaptive gradient methods at the edge of stability
Cohen, J. M., Ghorbani, B., Krishnan, S., Agarwal, N., Medapati, S., Badura, M., Suo, D., Cardoze, D., Nado, Z., Dahl, G. E., et al · 2022
Later among the works it cites.
How does adaptive optimization impact local neural network geometry?
Jiang, K., Malik, D., and Li, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S · 2020
Cited alongside, same era.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Zhou, P., Feng, J., Ma, C., Xiong, C., Hoi, S. C. H., et al · 2020
Cited alongside, same era.
Convergence and dynamical behavior of the adam algorithm for nonconvex stochastic optimization
Barakat, A. and Bianchi, P · 2021
Cited alongside, same era.
Implicit gradient regularization
Barrett, D. and Dherin, B · 2021
Cited alongside, same era.
When vision transformers outperform resnets without pre-training or strong data augmentations
Chen, X., Hsieh, C.-J., and Gong, B · 2021
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Cohen, J., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Cited alongside, same era.
Kumar, A., Shen, R., Bubeck, S., and Gunasekar, S · 2022
Later among the works it cites.
A qualitative study of the dynamic behavior for adaptive gradient algorithms
Ma, C., Wu, L., and Weinan, E · 2022
Later among the works it cites.
On the sdes and scaling rules for adaptive gradient algorithms
Malladi, S., Lyu, K., Panigrahi, A., and Arora, S · 2022
Later among the works it cites.
Toward equation of motion for deep neural networks: Continuous-time gradient descent and discretization error analysis
Miyagawa, T · 2022
Later among the works it cites.
Does momentum change the implicit regularization on separable data?
Wang, B., Meng, Q., Zhang, H., Sun, R., Chen, W., Ma, Z.-M., and Liu, T.-Y · 2022
Later among the works it cites.
Adaptive inertia: Disentangling the effects of adaptive learning rate and momentum
Xie, Z., Wang, X., Zhang, H., Sato, I., and Sugiyama, M · 2022
Later among the works it cites.
Penalizing gradient norm for efficiently improving generalization in deep learning
Zhao, Y., Zhang, H., and Hu, X · 2022
Later among the works it cites.
A modern look at the relationship between sharpness and generalization
Andriushchenko, M., Croce, F., Müller, M., Hein, M., and Flammarion, N · 2023
Closest in time.
On the trajectories of sgd without replacement
Beneventano, P · 2023
Closest in time.
(s) gd over diagonal linear networks: Implicit regularisation, large stepsizes and edge of stability
Even, M., Pesme, S., Gunasekar, S., and Flammarion, N · 2023
Closest in time.
Implicit regularization in heavy-ball momentum accelerated stochastic gradient descent
Ghosh, A., Lyu, H., Zhang, X., and Wang, R · 2023
Closest in time.
A theory on adam instability in large-scale machine learning
Molybog, I., Albert, P., Chen, M., DeVito, Z., Esiobu, D., Goyal, N., Koura, P. S., Narang, S., Poulton, A., Silva, R., et al · 2023
Closest in time.