Fetching the paper…
Reading the bibliography…
While Truncated Back-Propagation through Time (BPTT) is the most popular approach to training Recurrent Neural Networks (RNNs), it suffers from being inherently sequential (making parallelization difficult) and from truncating gradient flow between distant time-steps.
Variational methods for the solution of problems of equilibrium and vibrations
Richard Courant et al. 1943 · 1943
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
” A method for non-linear constraints in minimization problems”
Michael JD Powell. 1967 · 1967
Earlier work this paper cites.
Multiplier and gradient methods
Magnus R Hestenes. 1969 · 1969
Earlier work this paper cites.
Sur l’approximation, par éléments finis d’ordre un, et la résolution, par pénalisation-dualité d’une classe de problèmes de dirichlet non linéaires
Roland Glowinski and A Marroco. 1975 · 1975
Earlier work this paper cites.
A dual algorithm for the solution of nonlinear variational problems via finite element approximation
Daniel Gabay and Bertrand Mercier. 1976 · 1976
Earlier work this paper cites.
Learning processes in an asymmetric threshold network
Yann LeCun. 1986 · 1986
Earlier work this paper cites.
Modèles connexionnistes de l’apprentissage
Yann Le Cun. 1987 · 1987
Earlier work this paper cites.
Theoretical framework for back-propagation
Yann LeCun. 1988 · 1988
Earlier work this paper cites.
A cost function for internal representations
Anders Krogh, CI Thorbergsson, and John A Hertz. 1989 · 1989
Cited alongside, same era.
Finding structure in time
Jeffrey L. Elman. 1990 · 1990
Cited alongside, same era.
Backpropagation through time: what it does and how to do it
Paul J Werbos. 1990 · 1990
Cited alongside, same era.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Cited alongside, same era.
Numerical optimization, second edition
Jorge Nocedal and Stephen J Wright. 2006 · 2006
Cited alongside, same era.
Learning invariant features through topographic filter maps
Koray Kavukcuoglu, Marc’Aurelio Ranzato, Rob Fergus, and Yann LeCun. 2009 · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Later among the works it cites.
How auto-encoders could provide credit assignment in deep networks via target propagation
Yoshua Bengio. 2014 · 2014
Later among the works it cites.
Distributed optimization of deeply nested systems
Miguel Á. Carreira-Perpiñán and Weiran Wang. 2014 · 2014
Later among the works it cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Later among the works it cites.
Hashing with binary autoencoders
Miguel Á. Carreira-Perpiñán and Ramin Raziperchikolaei. 2015 · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic factor graphs for time series modeling
Piotr Mirowski and Yann LeCun. 2009 · 2009
Cited alongside, same era.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur. 2010 · 2010
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen P. Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein. 2011 · 2011
Cited alongside, same era.
Difference target propagation
Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio. 2015 · 2015
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable ADMM approach
Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein. 2016 · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan Salakhutdinov. 2016 · 2016
Later among the works it cites.