Fetching the paper…
Reading the bibliography…
Deep neural networks are known to be difficult to train due to the instability of back-propagation.
A method for unconstrained convex minimization problem with the rate of convergence o (1/k2)
Nesterov, Yurii · 1983
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Yann, Boser, Bernhard, Denker, John S, Henderson, Donnie, Howard, Richard E, Hubbard, Wayne, and Jackel, Lawrence D · 1989
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Yoshua, Simard, Patrice, and Frasconi, Paolo · 1994
Earlier work this paper cites.
A desicion-theoretic generalization of on-line learning and an application to boosting
Freund, Yoav and Schapire, Robert E · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, Ning · 1999
Earlier work this paper cites.
Convex neural networks
Bengio, Yoshua, Le Roux, Nicolas, Vincent, Pascal, Delalleau, Olivier, and Marcotte, Patrice · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex and Hinton, Geoffrey · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Yuval, Wang, Tao, Coates, Adam, Bissacco, Alessandro, Wu, Bo, and Ng, Andrew Y · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Yann A, Bottou, Léon, Orr, Genevieve B, and Müller, Klaus-Robert · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Zeiler, Matthew D · 2012
Earlier work this paper cites.
A theory of multiclass boosting
Mukherjee, Indraneel and Schapire, Robert E · 2013
Cited alongside, same era.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning
Rubinstein, Reuven Y and Kroese, Dirk P · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M, McClelland, James L, and Ganguli, Surya · 2013
Cited alongside, same era.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Sermanet, Pierre, Eigen, David, Zhang, Xiang, Mathieu, Michaël, Fergus, Rob, and LeCun, Yann · 2013
Cited alongside, same era.
Deep boosting
Cortes, Corinna, Mohri, Mehryar, and Syed, Umar · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Later among the works it cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Janzamin, Majid, Sedghi, Hanie, and Anandkumar, Anima · 2015
Later among the works it cites.
Fully convolutional networks for semantic segmentation
Long, Jonathan, Shelhamer, Evan, and Darrell, Trevor · 2015
Later among the works it cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, Shaoqing, He, Kaiming, Girshick, Ross, and Sun, Jian · 2015
Later among the works it cites.
Srivastava, Rupesh Kumar, Greff, Klaus, and Schmidhuber, Jürgen · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, Ross, Donahue, Jeff, Darrell, Trevor, and Malik, Jitendra · 2014
Cited alongside, same era.
Spatial pyramid pooling in deep convolutional networks for visual recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Selfieboost: A boosting algorithm for deep learning
Shalev-Shwartz, Shai · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
Zeiler, Matthew D and Fergus, Rob · 2014
Cited alongside, same era.
Fast r-cnn
Girshick, Ross · 2015
Cited alongside, same era.
Later among the works it cites.
Going deeper with convolutions
Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott, Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew · 2015
Later among the works it cites.
Convolutional networks and learning invariant to homogeneous multiplicative scalings
Tygert, Mark, Szlam, Arthur, Chintala, Soumith, Ranzato, Marc’Aurelio, Tian, Yuandong, and Zaremba, Wojciech · 2015
Later among the works it cites.
Adanet: Adaptive structural learning of artificial neural networks
Cortes, Corinna, Gonzalvo, Xavi, Kuznetsov, Vitaly, Mohri, Mehryar, and Yang, Scott · 2016
Later among the works it cites.
Identity matters in deep learning
Hardt, Moritz and Ma, Tengyu · 2016
Later among the works it cites.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Later among the works it cites.
Deep vs. shallow networks: An approximation theory perspective
Mhaskar, Hrushikesh N and Poggio, Tomaso · 2016
Later among the works it cites.
Residual networks behave like ensembles of relatively shallow networks
Veit, Andreas, Wilber, Michael J, and Belongie, Serge · 2016
Later among the works it cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, David, Frean, Marcus, Leary, Lennox, Lewis, JP, Ma, Kurt Wan-Duo, and McWilliams, Brian · 2017
Closest in time.