Computing machinery and intelligence
Alan M. Turing · 1950
Earlier work this paper cites.
Conditional markov processes
Ruslan L. Stratonovich · 1960
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
Andrew J. Viterbi · 1967
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J. Hopfield · 1982
Earlier work this paper cites.
Learning internal representations by error propagation
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1985
Earlier work this paper cites.
Learning distributed representations of concepts, 1986
Geoffrey E. Hinton · 1986
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
Michael I. Jordan · 1986
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser · 1989
Earlier work this paper cites.
Evolving networks: Using the genetic algorithm with connectionist learning
Richard K. Belew, John McInerney, and Nicol N. Schraudolph · 1990
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Yann Le Cun, B. Boser, John S. Denker, D. Henderson, Richard E. Howard, W. Hubbard, and Lawrence D. Jackel · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J. Werbos · 1990
Earlier work this paper cites.
Turing computability with neural nets
Hava T. Siegelmann and Eduardo D. Sontag · 1991
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim L. Blum and Ronald L. Rivest · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Neural network synthesis using cellular encoding and the genetic algorithm., 1994
Frédéric Gruau, L’universite Claude Bernard lyon I, Of A Diplome De Doctorat, M. Jacques Demongeot, Examinators M. Michel Cosnard, M. Jacques Mazoyer, M. Pierre Peretto, and M. Darell Whitley · 1994
Earlier work this paper cites.
Gradient calculations for dynamic recurrent neural networks: A survey
Barak A. Pearlmutter · 1995
Earlier work this paper cites.
Working memory and executive control [and discussion]
Alan Baddeley, Sergio Della Sala, and T.W. Robbins · 1996
Earlier work this paper cites.
Bridging long time lags by weight guessing and “long short-term memory”
Sepp Hochreiter and Jurgen Schmidhuber · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K. Paliwal · 1997
Earlier work this paper cites.
Recurrent nets that time and count
Felix A. Gers and Jürgen Schmidhuber · 2000
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred Cummins · 2000
Earlier work this paper cites.
Long short-term memory in recurrent neural networks
Felix A. Gers · 2001
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.