Fetching the paper…
Reading the bibliography…
Learning useful information across long time lags is a critical and difficult problem for temporal neural models in tasks such as language modeling.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Finding structure in time
Elman, J. L. (1990) · 1990
Earlier work this paper cites.
Artificial neural networks
Jordan, M. I. (1990) · 1990
Earlier work this paper cites.
Second-order recurrent neural networks for grammatical inference
Giles, C. L., Chen, D., Miller, C., Chen, H., Sun, G., and Lee, Y. (1991) · 1991
Earlier work this paper cites.
Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory
Das, S., Giles, C. L., and Sun, G.-Z. (1992) · 1992
Earlier work this paper cites.
Learning and extracting finite state automata with second-order recurrent neural networks
Giles, C. L., Miller, C. B., Chen, D., Chen, H.-H., Sun, G.-Z., and Lee, Y.-C. (1992) · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B. (1993) · 1993
Earlier work this paper cites.
Neural net architectures for temporal sequence processing
Mozer, M. C. (1993) · 1993
Earlier work this paper cites.
First-order versus second-order single-layer recurrent neural networks
Goudreau, M. W., Giles, C. L., Chakradhar, S. T., and Chen, D. (1994) · 1994
Earlier work this paper cites.
LTSM can solve hard time lag problems
Hochreiter, S. and Schmidhuber, J. (1997b) · 1996
Earlier work this paper cites.
Computation of conditional probability statistics by 8-month-old infants
Aslin, R. N., Saffran, J. R., and Newport, E. L. (1998) · 1998
Earlier work this paper cites.
The neural network pushdown automaton: Architecture, dynamics and training
Sun, G.-Z., Giles, C. L., and Chen, H.-H. (1998) · 1998
Earlier work this paper cites.
Recurrent nets that time and count
Gers, F. A. and Schmidhuber, J. (2000) · 2000
Earlier work this paper cites.
Noisy time series prediction using recurrent neural networks and grammatical inference
Giles, C. L., Lawrence, S., and Tsoi, A. C. (2001) · 2001
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
Hale, J. (2001) · 2001
Earlier work this paper cites.
Morphological Processing
Baayen, R. H. and Schreuder, R. (2006) · 2006
Earlier work this paper cites.
Parsing costs as predictors of reading difficulty: An evaluation using the Potsdam Sentence Corpus
Boston, M. F., Hale, J., Kliegl, R., Patil, U., and Vasishth, S. (2008) · 2008
Earlier work this paper cites.
Expectation-based syntactic comprehension
Levy, R. (2008) · 2008
Earlier work this paper cites.
Quadratic features and deep architectures for chunking
Turian, J., Bergstra, J., and Bengio, Y. (2009) · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černocký, J., and Khudanpur, S. (2010) · 2010
Cited alongside, same era.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
Mikolov, T., Kombrink, S., Burget, L., Černocký, J., and Khudanpur, S. (2011) · 2011
Cited alongside, same era.
Statistical Language Models Based on Neural Networks
Mikolov, T. (2012) · 2012
Cited alongside, same era.
Bayesian learning for neural networks
Neal, R. M. (2012) · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
Graves, A. (2013) · 2013
Cited alongside, same era.
Sukhbaatar, S., Szlam, A., Weston, J., and Fergus, R. (2015) · 2015
Later among the works it cites.
Larger-context language modelling
Wang, T. and Cho, K. (2015) · 2015
Later among the works it cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Later among the works it cites.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y. (2016) · 2016
Later among the works it cites.
Cooijmans, T., Ballas, N., Laurent, C., Gülçehre, Ç., and Courville, A. (2016) · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y. (2013) · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2014) · 2014
Cited alongside, same era.
Koutnik, J., Greff, K., Gomez, F., and Schmidhuber, J. (2014) · 2014
Cited alongside, same era.
Learning longer memory in recurrent neural networks
Mikolov, T., Joulin, A., Chopra, S., Mathieu, M., and Ranzato, M. (2014) · 2014
Cited alongside, same era.
Later among the works it cites.
A theoretically grounded application of dropout in recurrent neural networks
Gal, Y. and Ghahramani, Z. (2016) · 2016
Later among the works it cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., et al. (2016) · 2016
Later among the works it cites.
Gulcehre, C., Moczulski, M., Denil, M., and Bengio, Y. (2016) · 2016
Later among the works it cites.
Ha, D., Dai, A., and Le, Q. V. (2016) · 2016
Later among the works it cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Later among the works it cites.
Variable computation in recurrent neural networks
Jernite, Y., Grave, E., Joulin, A., and Mikolov, T. (2016) · 2016
Later among the works it cites.
Zoneout: Regularizing rnns by randomly preserving hidden activations
Krueger, D., Maharaj, T., Kramár, J., Pezeshki, M., Ballas, N., Ke, N. R., Goyal, A., Bengio, Y., Larochelle, H., Courville, A., et al. (2016) · 2016
Later among the works it cites.
Piecewise Latent Variables for Neural Variational Text Processing
Serban, I. V., Ororbia, I., Alexander, G., Pineau, J., and Courville, A. (2016) · 2016
Later among the works it cites.
Improvements in Language and Translation Modeling
Sundermeyer, M. (2016) · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Wu, Y., Zhang, S., Zhang, Y., Bengio, Y., and Salakhutdinov, R. R. (2016) · 2016
Later among the works it cites.
Minimal gated unit for recurrent neural networks
Zhou, G.-B., Wu, J., Zhang, C.-L., and Zhou, Z.-H. (2016) · 2016
Later among the works it cites.
Subword language modeling with neural networks
Mikolov, T., Sutskever, I., Deoras, A., Le, H.-S., Kombrink, S., and Černocký, J. (2012) · 2017
Closest in time.
Memory Augmented Neural Networks with Wormhole Connections
Gulcehre, Caglar and Chandar, Sarath and Bengio, Yoshua (2017) · 2017
Closest in time.
Gated feedback recurrent neural networks
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2015) · 2075
Closest in time.