Fetching the paper…
Reading the bibliography…
It is important to understand how the popular regularization method dropout helps the neural network training find a good generalization solution.
Dropout: a simple way to prevent neural networks from overfitting,
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, R. Salakhutdinov, · 1958
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance,
G. An, · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition,
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images (2009)
A. Krizhevsky, et al., · 2009
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors,
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, R. R. Salakhutdinov, · 2012
Earlier work this paper cites.
A pac-bayesian tutorial with a dropout bound,
D. McAllester, · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect,
L. Wan, M. Zeiler, S. Zhang, Y. Lecun, R. Fergus, · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition,
K. Simonyan, A. Zisserman, · 2014
Earlier work this paper cites.
On the inductive bias of dropout,
D. P. Helmbold, P. M. Long, · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization,
D. P. Kingma, J. Ba, · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima,
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. T. P. Tang, · 2016
Cited alongside, same era.
Multi30k: Multilingual english-german image descriptions,
D. Elliott, S. Frank, K. Sima’an, L. Specia, · 2016
Cited alongside, same era.
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, J. Sun, · 2016
Cited alongside, same era.
Exploring generalization in deep learning,
B. Neyshabur, S. Bhojanapalli, D. McAllester, N. Srebro, · 2017
Cited alongside, same era.
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, · 2017
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size,
V. Papyan, · 2018
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks,
M. Tan, Q. Le, · 2019
Later among the works it cites.
V. Papyan, · 2019
Later among the works it cites.
Training behavior of deep neural network in frequency domain,
Z.-Q. J. Xu, Y. Zhang, Y. Xiao, · 2019
Later among the works it cites.
The implicit and explicit regularization effects of dropout,
C. Wei, S. Kakade, T. Ma, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Shwartz-Ziv, N. Tishby, · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets,
H. Li, Z. Xu, G. Taylor, C. Studer, T. Goldstein, · 2017
Cited alongside, same era.
Z. Zhu, J. Wu, B. Yu, L. Wu, J. Ma, · 2018
Cited alongside, same era.
Dropout training, data-dependent regularization, and generalization bounds,
W. Mou, Y. Zhou, J. Gao, L. Wang, · 2018
Cited alongside, same era.
P. Mianjy, R. Arora, · 2020
Later among the works it cites.
Frequency principle: Fourier analysis sheds light on deep neural networks,
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, Z. Ma, · 2020
Later among the works it cites.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima,
Y. Feng, Y. Tu, · 2021
Closest in time.
A linear frequency principle model to understand the absence of overfitting in neural networks,
Y. Zhang, T. Luo, Z. Ma, Z.-Q. J. Xu, · 2021
Closest in time.