Fetching the paper…
Reading the bibliography…
For most deep learning practitioners, sequence modeling is synonymous with recurrent networks.
Parallel networks that learn to pronounce English text
Sejnowski, Terrence J. and Rosenberg, Charles R · 1987
Earlier work this paper cites.
Connectionist learning procedures
Hinton, Geoffrey E · 1989
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Yann, Boser, Bernhard, Denker, John S., Henderson, Donnie, Howard, Richard E., Hubbard, Wayne, and Jackel, Lawrence D · 1989
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
Waibel, Alex, Hanazawa, Toshiyuki, Hinton, Geoffrey, Shikano, Kiyohiro, and Lang, Kevin J · 1989
Earlier work this paper cites.
Speaker-independent isolated digit recognition: Multilayer perceptrons vs. dynamic time warping
Bottou, Léon, Soulie, F Fogelman, Blanchet, Pascal, and Liénard, Jean-Sylvain · 1990
Earlier work this paper cites.
Finding structure in time
Elman, Jeffrey L · 1990
Earlier work this paper cites.
Backpropagation through time: What it does and how to do it
Werbos, Paul J · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn treebank
Marcus, Mitchell P, Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Yoshua, Simard, Patrice, and Frasconi, Paolo · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
El Hihi, Salah and Bengio, Yoshua · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, Mike and Paliwal, Kuldip K · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Learning precise timing with lstm recurrent networks
Gers, Felix A, Schraudolph, Nicol N, and Schmidhuber, Jürgen · 2002
Earlier work this paper cites.
Harmonising chorales by probabilistic inference
Allan, Moray and Williams, Christopher · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, Ronan and Weston, Jason · 2008
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Collobert, Ronan, Weston, Jason, Bottou, Léon, Karlen, Michael, Kavukcuoglu, Koray, and Kuksa, Pavel P · 2011
Earlier work this paper cites.
Learning recurrent neural networks with Hessian-free optimization
Martens, James and Sutskever, Ilya · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
Sutskever, Ilya, Martens, James, and Hinton, Geoffrey E · 2011
Earlier work this paper cites.
Boulanger-Lewandowski, Nicolas, Bengio, Yoshua, and Vincent, Pascal · 2012
Earlier work this paper cites.
Supervised Sequence Labelling with Recurrent Neural Networks
Graves, Alex · 2012
Earlier work this paper cites.
Subword language modeling with neural networks
Mikolov, Tomáš, Sutskever, Ilya, Deoras, Anoop, Le, Hai-Son, Kombrink, Stefan, and Cernocky, Jan · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, Alex · 2013
Earlier work this paper cites.
Training and analysing deep recurrent neural networks
Hermans, Michiel and Schrauwen, Benjamin · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, Kyunghyun, Van Merriënboer, Bart, Bahdanau, Dzmitry, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, KyungHyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Learning character-level representations for part-of-speech tagging
dos Santos, Cícero Nogueira and Zadrozny, Bianca · 2014
Cited alongside, same era.
A convolutional neural network for modelling sentences
Kalchbrenner, Nal, Grefenstette, Edward, and Blunsom, Phil · 2014
Cited alongside, same era.
Convolutional neural networks for sentence classification
Kim, Yoon · 2014
Cited alongside, same era.
A clockwork RNN
Koutnik, Jan, Greff, Klaus, Gomez, Faustino, and Schmidhuber, Juergen · 2014
Cited alongside, same era.
How to construct deep recurrent neural networks
Pascanu, Razvan, Gülçehre, Çaglar, Cho, Kyunghyun, and Bengio, Yoshua · 2014
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, Denis, Kruszewski, Germán, Lazaridou, Angeliki, Pham, Quan Ngoc, Bernardi, Raffaella, Pezzelle, Sandro, Baroni, Marco, Boleda, Gemma, and Fernández, Raquel · 2016
Later among the works it cites.
Using the output embedding to improve language models
Press, Ofir and Wolf, Lior · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, Tim and Kingma, Diederik P · 2016
Later among the works it cites.
WaveNet: A generative model for raw audio
van den Oord, Aäron, Dieleman, Sander, Zen, Heiga, Simonyan, Karen, Vinyals, Oriol, Graves, Alex, Kalchbrenner, Nal, Senior, Andrew W., and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Full-capacity unitary recurrent neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey E, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
Effective use of word order for text categorization with convolutional neural networks
Johnson, Rie and Zhang, Tong · 2015
Cited alongside, same era.
An empirical exploration of recurrent network architectures
Jozefowicz, Rafal, Zaremba, Wojciech, and Sutskever, Ilya · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2015
Cited alongside, same era.
Wisdom, Scott, Powers, Thomas, Hershey, John, Le Roux, Jonathan, and Atlas, Les · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Wu, Yuhuai, Zhang, Saizheng, Zhang, Ying, Bengio, Yoshua, and Salakhutdinov, Ruslan R · 2016
Later among the works it cites.
Multi-scale context aggregation by dilated convolutions
Yu, Fisher and Koltun, Vladlen · 2016
Later among the works it cites.
Architectural complexity measures of recurrent neural networks
Zhang, Saizheng, Wu, Yuhuai, Che, Tong, Lin, Zhouhan, Memisevic, Roland, Salakhutdinov, Ruslan R, and Bengio, Yoshua · 2016
Later among the works it cites.
Quasi-recurrent neural networks
Bradbury, James, Merity, Stephen, Xiong, Caiming, and Socher, Richard · 2017
Later among the works it cites.
Dilated recurrent neural networks
Chang, Shiyu, Zhang, Yang, Han, Wei, Yu, Mo, Guo, Xiaoxiao, Tan, Wei, Cui, Xiaodong, Witbrock, Michael J., Hasegawa-Johnson, Mark A., and Huang, Thomas S · 2017
Later among the works it cites.
Very deep convolutional networks for text classification
Conneau, Alexis, Schwenk, Holger, LeCun, Yann, and Barrault, Loïc · 2017
Later among the works it cites.
Language modeling with gated convolutional networks
Dauphin, Yann N., Fan, Angela, Auli, Michael, and Grangier, David · 2017
Later among the works it cites.
Improving neural language models with a continuous cache
Grave, Edouard, Joulin, Armand, and Usunier, Nicolas · 2017
Later among the works it cites.
LSTM: A search space odyssey
Greff, Klaus, Srivastava, Rupesh Kumar, Koutník, Jan, Steunebrink, Bas R., and Schmidhuber, Jürgen · 2017
Later among the works it cites.
HyperNetworks
Ha, David, Dai, Andrew, and Le, Quoc V · 2017
Later among the works it cites.
Tunable efficient unitary neural networks (EUNN) and their application to RNNs
Jing, Li, Shen, Yichen, Dubcek, Tena, Peurifoy, John, Skirlo, Scott, LeCun, Yann, Tegmark, Max, and Soljačić, Marin · 2017
Later among the works it cites.
Deep pyramid convolutional neural networks for text categorization
Johnson, Rie and Zhang, Tong · 2017
Later among the works it cites.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
Krueger, David, Maharaj, Tegan, Kramár, János, Pezeshki, Mohammad, Ballas, Nicolas, Ke, Nan Rosemary, Goyal, Anirudh, Bengio, Yoshua, Larochelle, Hugo, Courville, Aaron C., and Pal, Chris · 2017
Later among the works it cites.
Temporal convolutional networks for action segmentation and detection
Lea, Colin, Flynn, Michael D., Vidal, René, Reiter, Austin, and Hager, Gregory D · 2017
Later among the works it cites.
Regularizing and optimizing LSTM language models
Merity, Stephen, Keskar, Nitish Shirish, and Socher, Richard · 2017
Later among the works it cites.
Diagonal RNNs in symbolic music modeling
Subakan, Y Cem and Smaragdis, Paris · 2017
Later among the works it cites.
Comparative study of CNN and RNN for natural language processing
Yin, Wenpeng, Kann, Katharina, Yu, Mo, and Schütze, Hinrich · 2017
Later among the works it cites.
Skip RNN: Learning to skip state updates in recurrent neural networks
Campos, Victor, Jou, Brendan, Giró i Nieto, Xavier, Torres, Jordi, and Chang, Shih-Fu · 2018
Closest in time.
On the state of the art of evaluation in neural language models
Melis, Gábor, Dyer, Chris, and Blunsom, Phil · 2018
Closest in time.
Sequence Models (Course 5 of Deep Learning Specialization)
Ng, Andrew · 2018
Closest in time.
Breaking the softmax bottleneck: A high-rank RNN language model
Yang, Zhilin, Dai, Zihang, Salakhutdinov, Ruslan, and Cohen, William W · 2018
Closest in time.