Unsupervised label noise modeling and loss correction
Original
Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O’Connor, and Kevin McGuinness · 1904
Earlier work this paper cites.
Understanding and utilizing deep neural networks trained with noisy labels
Original
Pengfei Chen, Benben Liao, Guangyong Chen, and Shengyu Zhang · 1905
Earlier work this paper cites.
Fast autoaugment
Original
Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and Sungwoong Kim · 1905
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Y. E. Nesterov · 1983
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2002
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Original
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Original
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Original
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
Original
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision, 2015
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2015
Earlier work this paper cites.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Original
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Deep pyramidal residual networks
Original
Dongyoon Han, Jiwhan Kim, and Junmo Kim · 2016
Earlier work this paper cites.