Fetching the paper…
Reading the bibliography…
Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture that has been designed to address the vanishing and exploding gradient problems of conventional RNNs.
“An efficient gradient-based algorithm for online training of recurrent network trajectories,”
Ronald J. Williams and Jing Peng, · 1990
Earlier work this paper cites.
“Learning long-term dependencies with gradient descent is difficult,”
Yoshua Bengio, Patrice Simard, and Paolo Frasconi, · 1994
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K. Paliwal, · 1997
Earlier work this paper cites.
“Learning to forget: Continual prediction with LSTM,”
Felix A. Gers, Jürgen Schmidhuber, and Fred Cummins, · 2000
Earlier work this paper cites.
“LSTM recurrent networks learn simple context free and context sensitive languages,”
Felix A. Gers and Jürgen Schmidhuber, · 2001
Earlier work this paper cites.
“Learning precise timing with LSTM recurrent networks,”
Felix A. Gers, Nicol N. Schraudolph, and Jürgen Schmidhuber, · 2003
Cited alongside, same era.
“Framewise phoneme classification with bidirectional LSTM and other neural network architectures,”
Alex Graves and Jürgen Schmidhuber, · 2005
Cited alongside, same era.
“A novel connectionist system for unconstrained handwriting recognition,”
Alex Graves, Marcus Liwicki, Santiago Fernandez, Roman Bertolami, Horst Bunke, and Jürgen Schmidhuber, · 2009
Cited alongside, same era.
“Recurrent neural network based language model,”
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černocký, and Sanjeev Khudanpur, · 2010
Cited alongside, same era.
“Eigen v3,” http://eigen.tuxfamily.org, 2010
Gaël Guennebaud, Benoît Jacob, et al., · 2010
Cited alongside, same era.
“Acoustic modeling using deep belief networks,”
Abdel Rahman Mohamed, George E. Dahl, and Geoffrey E. Hinton, · 2012
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
George E. Dahl, Dong Yu, Li Deng, and Alex Acero, · 2012
Later among the works it cites.
“Application of pretrained deep neural networks to large vocabulary speech recognition,”
Navdeep Jaitly, Patrick Nguyen, Andrew Senior, and Vincent Vanhoucke, · 2012
Later among the works it cites.
“Large scale distributed deep networks.,”
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Ke Yang, and Andrew Y. Ng, · 2012
Later among the works it cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Later among the works it cites.
“Low-rank matrix factorization for deep neural network training with high-dimensional output targets,”
T.N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.