Fetching the paper…
Reading the bibliography…
Exponential Moving Average (EMA) is a widely used weight averaging (WA) regularization to learn flat optima for better generalizations without extra cost in deep neural network (DNN) optimization.
Cascade r-cnn: High-quality object detection and instance segmentation
Zhaowei Cai and Nuno Vasconcelos · 1939
Earlier work this paper cites.
A stochastic approximation method
Naresh K. Sinha and Michael P. Griscik · 1971
Earlier work this paper cites.
Learning internal representations by error propagation , pp. 318–362
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris Polyak and Anatoli B. Juditsky · 1992
Earlier work this paper cites.
No free lunch theorems for optimization
D.H. Wolpert and W.G. Macready · 1997
Earlier work this paper cites.
Minimum message length and kolmogorov complexity
Chris S. Wallace and David L. Dowe · 1999
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems, 2015
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Long short-term memory neural network for traffic speed prediction using remote microwave sensor data
Xiaolei Ma, Zhimin Tao, Yinhai Wang, Haiyang Yu, and Yunpeng Wang · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using LSTMs
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Tim Harley, Timothy P. Lillicrap, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout, 2017
Terrance DeVries and Graham W. Taylor · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, and Tom Goldstein · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Building a large annotated corpus of english: the penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 2017
Cited alongside, same era.
Agedb: The first manually collected, in-the-wild age database
Stylianos Moschoglou, Athanasios Papaioannou, Christos Sagonas, Jiankang Deng, Irene Kotsia, and Stefanos Zafeiriou · 2017
Cited alongside, same era.
Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms
Yunbo Wang, Mingsheng Long, Jianmin Wang, Zhifeng Gao, and Philip S Yu · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Later among the works it cites.
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Later among the works it cites.
Efficient sharpness-aware minimization for improved training of neural networks
Jiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, and Vincent Y. F. Tan · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Yolox: Exceeding yolo series in 2021
Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
Large batch training of convolutional networks with layer-wise adaptive rate scaling
Boris Ginsburg, Igor Gitman, and Yang You · 2018
Cited alongside, same era.
Decision boundary analysis of adversarial examples
Warren He, Bo Li, and Dawn Song · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Megdet: A large mini-batch object detector
Chao Peng, Tete Xiao, Zeming Li, Yuning Jiang, Xiangyu Zhang, Kai Jia, Gang Yu, and Jian Sun · 2018
Cited alongside, same era.
Deep expectation of real and apparent age from a single image without facial landmarks
Rasmus Rothe, Radu Timofte, and Luc Van Gool · 2018
Cited alongside, same era.
Iterative averaging in the quest for best test error, 2021
Diego Granziol, Xingchen Wan, Samuel Albanie, and Stephen Roberts · 2021
Later among the works it cites.
Boosting discriminative visual representation learning with scenario-agnostic mixup
Siyuan Li, Zicheng Liu, Zedong Wang, Di Wu, Zihan Liu, and Stan Z. Li · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Rethinking ”batch” in batchnorm
Yuxin Wu and Justin Johnson · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny · 2021
Later among the works it cites.
Towards understanding why lookahead generalizes better than sgd and beyond
Pan Zhou, Hanshu Yan, Xiaotong Yuan, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Sharpness-aware training for free
Jiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan, and Joey Tianyi Zhou · 2022
Later among the works it cites.
Simvp: Simpler yet better video prediction
Zhangyang Gao, Cheng Tan, Lirong Wu, and Stan Z. Li · 2022
Later among the works it cites.
Stochastic weight averaging revisited, 2022
Hao Guo, Jiyong Jin, and Bin Liu · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging, 2022
Jean Kaddour · 2022
Later among the works it cites.
When do flat minima optimizers work?
Jean Kaddour, Linqing Liu, Ricardo M. A. Silva, and Matt J. Kusner · 2022
Later among the works it cites.
Openmixup: Open mixup toolbox and benchmark for visual representation learning
Siyuan Li, Zedong Wang, Zicheng Liu, Di Wu, and Stan Z. Li · 2022
Later among the works it cites.
Usb: A unified semi-supervised learning benchmark for classification
Yidong Wang, Hao Chen, Yue Fan, Wang SUN, Ran Tao, Wenxin Hou, Renjie Wang, Linyi Yang, Zhi Zhou, Lan-Zhe Guo, Heli Qi, Zhen Wu, Yu-Feng Li, Satoshi Nakamura, Wei Ye, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang, Xing Xie, and Yue Zhang · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2022
Later among the works it cites.
C-mixup: Improving generalization in regression
Huaxiu Yao, Yiping Wang, Linjun Zhang, James Y Zou, and Chelsea Finn · 2022
Later among the works it cites.
Metaformer is actually what you need for vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2022
Later among the works it cites.
Surrogate gap minimization improves sharpness-aware training
Juntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui, Hartwig Adam, Nicha C. Dvornek, Sekhar Chandra Tatikonda, James S. Duncan, and Ting Liu · 2022
Later among the works it cites.
Why do we need weight decay in modern deep learning?
Maksym Andriushchenko, Francesco D’Angelo, Aditya Varre, and Nicolas Flammarion · 2023
Later among the works it cites.
Pfge: Parsimonious fast geometric ensembling of dnns, 2023
Hao Guo, Jiyong Jin, and Bin Liu · 2023
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine · 2023
Later among the works it cites.
Openstl: A comprehensive benchmark of spatio-temporal predictive learning
Cheng Tan, Siyuan Li, Zhangyang Gao, Wenfei Guan, Zedong Wang, Zicheng Liu, Lirong Wu, and Stan Z Li · 2023
Later among the works it cites.
Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models
Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng YAN · 2023
Later among the works it cites.
Adversarial automixup
Huafeng Qin, Xin Jin, Yun Jiang, Mounim A. El-Yacoubi, and Xinbo Gao · 2024
Closest in time.