Fetching the paper…
Reading the bibliography…
Weight averaging of Stochastic Gradient Descent (SGD) iterates is a popular method for training deep learning models.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
Boris T Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis Bach and Eric Moulines · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Ya Le and Xuan Yang · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Chandrashekar Lakshminarayanan and Csaba Szepesvari · 2018
Earlier work this paper cites.
Iterate averaging as regularization for stochastic gradient descent
Gergely Neu and Lorenzo Rosasco · 2018
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al · 2018
Cited alongside, same era.
There are many consistent explanations of unlabeled data: Why you should average
Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson · 2019
Cited alongside, same era.
Mixmatch: A holistic approach to semi-supervised learning
David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel · 2019
Cited alongside, same era.
Semi-supervised semantic segmentation needs strong, varied perturbations
Geoff French, Samuli Laine, Timo Aila, Michal Mackiewicz, and Graham Finlayson · 2019
Cited alongside, same era.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li · 2020
Later among the works it cites.
On the reproducibility of neural network predictions
Srinadh Bhojanapalli, Kimberly Wilber, Andreas Veit, Ankit Singh Rawat, Seungyeon Kim, Aditya Menon, and Sanjiv Kumar · 2021
Later among the works it cites.
Exponential moving average normalization for self-supervised and semi-supervised learning
Zhaowei Cai, Avinash Ravichandran, Subhransu Maji, Charless Fowlkes, Zhuowen Tu, and Stefano Soatto · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Later among the works it cites.
Swad: Domain generalization by seeking flat minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Cited alongside, same era.
On model stability as a function of random seed
Pranava Madhyastha and Rishabh Jain · 2019
Cited alongside, same era.
Beating sgd saturation with tail-averaging and minibatching
Nicole Mücke, Gergely Neu, and Lorenzo Rosasco · 2019
Cited alongside, same era.
Self: Learning to filter noisy labels with self-ensembling
Duc Tam Nguyen, Chaithanya Kumar Mummadi, Thi Phuong Nhung Ngo, Thi Hoai Phuong Nguyen, Laura Beggel, and Thomas Brox · 2019
Cited alongside, same era.
Measuring calibration in deep learning
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran · 2019
Cited alongside, same era.
Swalp: Stochastic weight averaging in low precision training
Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, and Chris De Sa · 2019
Cited alongside, same era.
The unusual effectiveness of averaging in gan training
Yasin Yaz, Chuan-Sheng Foo, Stefan Winkler, Kim-Hui Yap, Georgios Piliouras, Vijay Chandrasekhar, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Robust overfitting may be mitigated by properly learned smoothening
Tianlong Chen, Zhenyu Zhang, Sijia Liu, Shiyu Chang, and Zhangyang Wang · 2021
Later among the works it cites.
Soft calibration objectives for neural networks
Archit Karandikar, Nicholas Cain, Dustin Tran, Balaji Lakshminarayanan, Jonathon Shlens, Michael C Mozer, and Becca Roelofs · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Data augmentation can improve robustness
Sylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg, Olivia Wiles, and Timothy A Mann · 2021
Later among the works it cites.
Daformer: Improving network architectures and training strategies for domain-adaptive semantic segmentation
Lukas Hoyer, Dengxin Dai, and Luc Van Gool · 2022
Later among the works it cites.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
Jean Kaddour · 2022
Later among the works it cites.
Trainable weight averaging for fast convergence and better generalization
Tao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu, and Xiaolin Huang · 2022
Later among the works it cites.
Continual test-time domain adaptation
Qin Wang, Olga Fink, Luc Van Gool, and Dengxin Dai · 2022
Later among the works it cites.
Learning with noisy labels revisited: A study using real-world human annotations
Jiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu, Gang Niu, and Yang Liu · 2022
Later among the works it cites.
(s) gd over diagonal linear networks: Implicit regularisation, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion · 2023
Later among the works it cites.
Optimal non-asymptotic analysis of the ruppert–polyak averaging stochastic algorithm
Sébastien Gadat and Fabien Panloup · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision, 2023
Maxime Oquab, Timothée Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2023
Later among the works it cites.
Training trajectories, mini-batch losses and the curious role of the learning rate, 2023
Mark Sandler, Andrey Zhmoginov, Max Vladymyrov, and Nolan Miller · 2023
Later among the works it cites.
Early weight averaging meets high learning rates for llm pre-training, 2023
Sunny Sanyal, Atula Neerkaje, Jean Kaddour, Abhishek Kumar, and Sujay Sanghavi · 2023
Later among the works it cites.