Fetching the paper…
Reading the bibliography…
This paper shows how Long Short-term Memory recurrent neural networks can be used to generate complex sequences with long-range structure, simply by predicting one data point at a time.
Data compression using adaptive coding and partial string matching
J. G. Cleary, Ian, and I. H. Witten · 1984
Earlier work this paper cites.
The tagged LOB corpus user’s manual; Norwegian Computing Centre for the Humanities, 1986
S. Johansson, R. Atwell, R. Garside, and G. Leech · 1986
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Mixture density networks
C. Bishop · 1994
Earlier work this paper cites.
Neural Networks for Pattern Recognition
C. Bishop · 1995
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent networks and their computational complexity
R. Williams and D. Zipser · 1995
Earlier work this paper cites.
An analysis of noise in recurrent neural networks: convergence and generalization
K.-C. Jim, C. Giles, and B. Horne · 1996
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Better generative models for sequential data problems: Bidirectional recurrent mixture density networks
M. Schuster · 1999
Earlier work this paper cites.
Gradient Flow in Recurrent Nets: the Difficulty of Learning Long-term Dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Learning precise timing with LSTM recurrent networks
F. Gers, N. Schraudolph, and J. Schmidhuber · 2002
Cited alongside, same era.
Framewise phoneme classification with bidirectional LSTM and other neural network architectures
A. Graves and J. Schmidhuber · 2005
Cited alongside, same era.
IAM-OnDB - an on-line English sentence database acquired from handwritten text on a whiteboard
M. Liwicki and H. Bunke · 2005
Cited alongside, same era.
The Minimum Description Length Principle (Adaptive Computation and Machine Learning)
P. D. Grünwald · 2007
Cited alongside, same era.
Offline handwriting recognition with multidimensional recurrent neural networks
A. Graves and J. Schmidhuber · 2008
Cited alongside, same era.
A Scalable Hierarchical Distributed Language Model
A. Mnih and G. Hinton · 2008
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. Hinton · 2011
Later among the works it cites.
Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent · 2012
Later among the works it cites.
Sequence transduction with recurrent neural networks
A. Graves · 2012
Later among the works it cites.
The Human Knowledge Compression Contest, 2012
M. Hutter · 2012
Later among the works it cites.
Statistical Language Models based on Neural Networks
T. Mikolov · 2012
Later among the works it cites.
Subword language modeling with neural networks
T. Mikolov, I. Sutskever, A. Deoras, H. Le, S. Kombrink, and J. Cernocky · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The recurrent temporal restricted boltzmann machine
I. Sutskever, G. E. Hinton, and G. W. Taylor · 2008
Cited alongside, same era.
Factored conditional restricted boltzmann machines for modeling motion style
G. W. Taylor and G. E. Hinton · 2009
Cited alongside, same era.
A Practical Guide to Training Restricted Boltzmann Machines
G. Hinton · 2010
Cited alongside, same era.
Practical variational inference for neural networks
A. Graves · 2011
Cited alongside, same era.
A machine learning perspective on predictive coding with paq
B. Knoll and N. de Freitas · 2011
Cited alongside, same era.
A first look at music composition using lstm recurrent neural networks
D. Eck and J. Schmidhuber
Cited in the paper.
A fast and simple algorithm for training neural probabilistic language models
A. Mnih and Y. W. Teh · 2012
Later among the works it cites.
Lecture 6.5 - rmsprop: Divide the gradient by a running average of its recent magnitude, 2012
T. Tieleman and G. Hinton · 2012
Later among the works it cites.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Closest in time.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
T. N. Sainath, A. Mohamed, B. Kingsbury, and B. Ramabhadran · 2013
Closest in time.