Fetching the paper…
Reading the bibliography…
We propose proximal backpropagation (ProxProp) as a novel algorithm that takes implicit instead of explicit gradient steps to update the network parameters during neural network training.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Proximité et dualité dans un espace hilbertien
Jean-Jacques Moreau · 1965
Earlier work this paper cites.
Régularisation d’inéquations variationnelles par approximations successives
Bernard Martinet · 1970
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
A theoretical framework for back-propagation
Yann LeCun · 1988
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou · 1991
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, and Paolo Frasconi · 2001
Cited alongside, same era.
Numerical Optimization
Jorge Nocedal and Stephen Wright · 2006
Cited alongside, same era.
Deep learning via Hessian-free optimization
James Martens · 2010
Cited alongside, same era.
On optimization methods for deep learning
Quoc V Le, Adam Coates, Bobby Prochnow, and Andrew Y Ng · 2011
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Fast large-scale optimization by unifying stochastic gradient and quasi-newton methods
Jascha Sohl-Dickstein, Ben Poole, and Surya Ganguli · 2013
Cited alongside, same era.
Distributed optimization of deeply nested systems
Miguel Á. Carreira-Perpiñán and Weiran Wang · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina
Cited in the paper.
Deep Relaxation: partial differential equations for optimizing deep neural networks
Pratik Chaudhari, Adam Oberman, Stanley Osher, Stefano Soatto, and Guillame Carlier
Cited in the paper.
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable ADMM approach
Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein · 2016
Later among the works it cites.