Fetching the paper…
Reading the bibliography…
Overfitting is a crucial problem in deep neural networks, even in the latest network architectures.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Regularization of neural networks using DropConnect
L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger · 2016
Cited alongside, same era.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Cited alongside, same era.
Deep pyramidal residual networks
D. Han, J. Kim, and J. Kim · 2017
Later among the works it cites.
Deep pyramidal residual networks
D. Han, J. Kim, and J. Kim · 2017
Later among the works it cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
An alternative view: When does SGD escape local minima?
R. Kleinberg, Y. Li, and Y. Yuan · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep pyramidal residual networks with separated stochastic depth
Y. Yamada, M. Iwamura, and K. Kise · 2016
Cited alongside, same era.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
Dataset augmentation in feature space
T. DeVries and G. W. Taylor · 2017
Cited alongside, same era.
EraseReLU: A simple way to ease the training of deep convolution neural networks
X. Dong, G. Kang, K. Zhan, and Y. Yang · 2017
Cited alongside, same era.
X. Gastaldi · 2017
Cited alongside, same era.
Shake-shake regularization of 3-branch residual networks
X. Gastaldi · 2017
Cited alongside, same era.
Closest in time.
Virtual adversarial training: A regularization method for supervised and semi-supervised learning
T. Miyato, S.-I. Maeda, S. Ishii, and M. Koyama · 2018
Closest in time.
RICAP: Random image cropping and patching data augmentation for deep cnns
R. Takahashi, T. Matsubara, and K. Uehara · 2018
Closest in time.
Between-class learning for image classification
Y. Tokozume, Y. Ushiku, and T. Harada · 2018
Closest in time.
Manifold mixup: Encouraging meaningful on-manifold interpolation as a regularizer
V. Verma, A. Lamb, C. Beckham, A. Courville, I. Mitliagkis, and Y. Bengio · 2018
Closest in time.
Shakedrop regularization
Y. Yamada, M. Iwamura, and K. Kise · 2018
Closest in time.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Closest in time.