Fetching the paper…
Reading the bibliography…
Averaging neural network weights sampled by a backbone stochastic gradient descent (SGD) is a simple yet effective approach to assist the backbone SGD in finding better optima, in terms of generalization.
D. Ruppert, “Efficient estimations from a slowly convergent robbins-monro process,” Cornell University Operations Research and Industrial Engineering, Tech. Rep., 1988
1988
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Technical report, University of Toronto , 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-fei, “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 630–645
2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” arXiv preprint arXiv:1605.07146 , 2016
2016
Earlier work this paper cites.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
S. Leslie N, “Cyclical learning rates for training neural networks,” in Proceedings of IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 2017, pp. 464–472
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger, “Snapshot ensembles: Train 1, get M for free,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Z. Allen-Zhu, Y. Li, and Z. Song, “A convergence theory for deep learning via over-parameterization,” in International Conference on Machine Learning . PMLR, 2019, pp. 242–252
2019
Later among the works it cites.
B. Ghorbani, S. Krishnan, and Y. Xiao, “An investigation into neural net optimization via hessian eigenvalue density,” in International Conference on Machine Learning . PMLR, 2019, pp. 2232–2241
2019
Later among the works it cites.
L. Yao, C. Mao, and Y. Luo, “Graph convolutional networks for text classification,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2019, pp. 7370–7377
2019
Later among the works it cites.
J. Lee, I. Lee, and J. Kang, “Self-attention graph pooling,” in International Conference on Machine Learning . PMLR, 2019, pp. 3734–3743
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
P. Jain, S. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford, “Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification,” Journal of Machine Learning Research , vol. 18, 2018
2018
Cited alongside, same era.
G. Neu and L. Rosasco, “Iterate averaging as regularization for stochastic gradient descent,” in Conference On Learning Theory . PMLR, 2018, pp. 3222–3242
2018
Cited alongside, same era.
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson, “Averaging weights leads to wider optima and better generalization,” in Proceedings of Conference on Uncertainty in Artificial Intelligence (UAI) , 2018, pp. 1–10
2018
Cited alongside, same era.
T. Garipov, P. Izmailov, D. Podoprikhin, D. Vetrov, and A. G. Wilson, “Loss surfaces, mode connectivity, and fast ensembling of dnns,” in Advances in Neural Information Processing Systems , 2018, pp. 8803–8812
2018
Cited alongside, same era.
L. N. Smith and N. Topin, “Super-convergence: Very fast training of neural networks using large learning rates,” in Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications , vol. 11006. International Society for Optics and Photonics, 2019, p. 1100612
2019
Cited alongside, same era.
2020
Later among the works it cites.
F. M. Bianchi, D. Grattarola, and C. Alippi, “Spectral clustering with graph neural networks for graph pooling,” in International Conference on Machine Learning . PMLR, 2020, pp. 874–883
2020
Later among the works it cites.
J. Cha, S. Chun, K. Lee, H.-C. Cho, S. Park, Y. Lee, and S. Park, “SWAD: Domain generalization by seeking flat minima,” in Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
J.-W. Hwang, Y. Lee, S. Oh, and Y. Bae, “Adversarial training with stochastic weight average,” in IEEE International Conference on Image Processing (ICIP) . IEEE, 2021, pp. 814–818
2021
Later among the works it cites.
P. Cheridito, A. Jentzen, and F. Rossmannek, “Non-convergence of stochastic gradient descent in the training of deep neural networks,” Journal of Complexity , vol. 64, p. 101540, 2021
2021
Later among the works it cites.
Y. Yang, L. Hodgkinson, R. Theisen, J. Zou, J. E. Gonzalez, K. Ramchandran, and M. W. Mahoney, “Taxonomizing local versus global structure in neural network loss landscapes,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
A. Kleiner, B. Neyshabur, H. Mobahi, and P. Foret, “Sharpness-aware minimization for efficiently improving generalization,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.