Fetching the paper…
Reading the bibliography…
With the fast development of deep learning, it has become common to learn big neural networks using massive training data.
Modèles connexionnistes de l’apprentissage
LeCun, Yann · 1987
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Becker, Sue, Le Cun, Yann, et al · 1988
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, Kurt · 1991
Earlier work this paper cites.
The elements of statistical learning , volume 1
Friedman, Jerome, Hastie, Trevor, and Tibshirani, Robert · 2001
Earlier work this paper cites.
Higher-order derivatives and taylor’s formula in several variables, 2005
Folland, GB · 2005
Earlier work this paper cites.
Learning multiple layers of representation
Hinton, Geoffrey E · 2007
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, James · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Agarwal, Alekh and Duchi, John C · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, Benjamin, Re, Christopher, Wright, Stephen, and Niu, Feng · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, Léon · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Dean, Jeffrey, Corrado, Greg, Monga, Rajat, Chen, Kai, Devin, Matthieu, Mao, Mark, Senior, Andrew, Tucker, Paul, Yang, Ke, Le, Quoc V, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
More effective distributed ml via a stale synchronous parallel parameter server
Ho, Qirong, Cipar, James, Cui, Henggang, Lee, Seunghak, Kim, Jin Kyu, Gibbons, Phillip B, Gibson, Garth A, Ganger, Greg, and Xing, Eric P · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, Tomas, Sutskever, Ilya, Chen, Kai, Corrado, Greg S, and Dean, Jeff · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Deep learning with elastic averaging sgd
Zhang, Sixin, Choromanska, Anna E, and LeCun, Yann · 2015
Later among the works it cites.
Revisiting distributed synchronous sgd
Chen, Jianmin, Monga, Rajat, Bengio, Samy, and Jozefowicz, Rafal · 2016
Closest in time.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Chen, Kai and Huo, Qiang · 2016
Closest in time.
Distributed deep learning using synchronous stochastic gradient descent
Das, Dipankar, Avancha, Sasikanth, Mudigere, Dheevatsa, Vaidynathan, Karthikeyan, Sridharan, Srinivas, Kalamkar, Dhiraj, Kaul, Bharat, and Dubey, Pradeep · 2016
Closest in time.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Delay-tolerant algorithms for asynchronous distributed online learning
McMahan, Brendan and Streeter, Matthew · 2014
Cited alongside, same era.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Sak, Haşim, Senior, Andrew, and Beaufays, Françoise · 2014
Cited alongside, same era.
Revisiting asynchronous linear solvers: Provable convergence rate through randomization
Avron, Haim, Druinsky, Alex, and Gupta, Anshul · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, and LeCun, Yann · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Lian, Xiangru, Huang, Yijun, Li, Yuncheng, and Liu, Ji · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, et al · 2015
Cited alongside, same era.
Adadelay: Delay adaptive distributed stochastic convex optimization
Sra, Suvrit, Yu, Adams Wei, Li, Mu, and Smola, Alexander J · 2015
Cited alongside, same era.
Closest in time.
Deep learning without poor local minima
Kawaguchi, Kenji · 2016
Closest in time.
Gradient descent converges to minimizers
Lee, Jason D, Simchowitz, Max, Jordan, Michael I, and Recht, Benjamin · 2016
Closest in time.
Asynchrony begets momentum, with an application to deep learning
Mitliagkas, Ioannis, Zhang, Ce, Hadjis, Stefan, and Ré, Christopher · 2016
Closest in time.
Very deep multilingual convolutional neural networks for lvcsr
Sercu, Tom, Puhrsch, Christian, Kingsbury, Brian, and LeCun, Yann · 2016
Closest in time.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, Christian, Ioffe, Sergey, Vanhoucke, Vincent, and Alemi, Alex · 2016
Closest in time.
Convolutional sequence to sequence learning
Gehring, Jonas, Auli, Michael, Grangier, David, Yarats, Denis, and Dauphin, Yann N · 2017
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, Priya, Piotr, Dollar, Ross, Girshick, Pieter, Noordhuis, Lukasz, Wesolowski, Aapo, Kyrola, Andrew, Tulloch, Yangqing, Jia, and Kaiming, He · 2017
Closest in time.