Fetching the paper…
Reading the bibliography…
It is important to understand how dropout, a popular regularization method, aids in achieving a good generalization solution during neural network training.
Geometric numerical integration illustrated by the störmer–verlet method
Ernst Hairer, Christian Lubich, and Gerhard Wanner · 2003
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
A pac-bayesian tutorial with a dropout bound
David McAllester · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Lecun, and Rob Fergus · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Stefan Wager, Sida Wang, and Percy S Liang · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
On the inductive bias of dropout
David P Helmbold and Philip M Long · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel · 2018
Cited alongside, same era.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Cited alongside, same era.
Dropout training, data-dependent regularization, and generalization bounds
Wenlong Mou, Yuchen Zhou, Jun Gao, and Liwei Wang · 2018
Cited alongside, same era.
On the implicit bias of dropout
Poorya Mianjy, Raman Arora, and Rene Vidal · 2018
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Later among the works it cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2020
Later among the works it cites.
Phase diagram for two-layer relu neural networks at infinite-width limit
Tao Luo, Zhi-Qin John Xu, Zheng Ma, and Yaoyu Zhang · 2021
Later among the works it cites.
Towards understanding the condensation of neural networks at initial training
Hanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang, and Zhi-Qin John Xu · 2021
Later among the works it cites.
Embedding principle of loss landscape of deep neural networks
Yaoyu Zhang, Zhongwang Zhang, Tao Luo, and Zhiqin J Xu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
Fluctuation-dissipation relations for stochastic gradient descent
Sho Yaida · 2018
Cited alongside, same era.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain
Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Cited alongside, same era.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2020
Cited alongside, same era.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2020
Cited alongside, same era.
A linear frequency principle model to understand the absence of overfitting in neural networks
Yaoyu Zhang, Tao Luo, Zheng Ma, and Zhi-Qin John Xu · 2021
Later among the works it cites.
Theory of the frequency principle for general deep neural networks
Tao Luo, Zheng Ma, Zhi-Qin John Xu, and Yaoyu Zhang · 2021
Later among the works it cites.
Empirical phase diagram for three-layer neural networks with infinite width
Hanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo, Yaoyu Zhang, and Zhi-Qin John Xu · 2022
Closest in time.
Embedding principle: a hierarchical structure of loss landscape of deep neural networks
Yaoyu Zhang, Yuqing Li, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2022
Closest in time.
Embedding principle in depth for the loss landscape analysis of deep neural networks
Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, and Yaoyu Zhang · 2022
Closest in time.
Linear stability hypothesis and rank stratification for nonlinear models
Yaoyu Zhang, Zhongwang Zhang, Leyang Zhang, Zhiwei Bai, Tao Luo, and Zhi-Qin John Xu · 2022
Closest in time.
Overview frequency principle/spectral bias in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, and Tao Luo · 2022
Closest in time.