Fetching the paper…
Reading the bibliography…
Long Short-Term Memory (LSTM) is one of the most widely used recurrent structures in sequence modeling.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Mutual information, metric entropy and cumulative relative entropy risk
Haussler, D., Opper, M., et al · 1997
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Hochreiter, S · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, F. A., Schmidhuber, J., and Cummins, F · 1999
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Zeiler, M. D · 2012
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., and Federico, M · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Jean, S., Cho, K., Memisevic, R., and Bengio, Y · 2015
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Karpathy, A., Johnson, J., and Fei-Fei, L · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2015
Earlier work this paper cites.
Minimum risk training for neural machine translation
Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian, S., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-c · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., and LeCun, Y · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Later among the works it cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Later among the works it cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Later among the works it cites.
Highway long short-term memory rnns for distant speech recognition
Zhang, Y., Chen, G., Yu, D., Yaco, K., Khudanpur, S., and Glass, J · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Grave, E., Joulin, A., and Usunier, N · 2016
Cited alongside, same era.
Dual learning for machine translation
He, D., Xia, Y., Qin, T., Wang, L., Yu, N., Liu, T., and Ma, W.-Y · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Later among the works it cites.
Zilly, J. G., Srivastava, R. K., Koutník, J., and Schmidhuber, J · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2016
Later among the works it cites.
Massive exploration of neural machine translation architectures
Britz, D., Goldie, A., Luong, T., and Le, Q · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Later among the works it cites.
Decoding with value networks for neural machine translation
He, D., Lu, H., Xia, Y., Qin, T., Wang, L., and Liu, T · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Melis, G., Dyer, C., and Blunsom, P · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R · 2017
Later among the works it cites.
Automatic rule extraction from long short term memory networks
Murdoch, W. J. and Szlam, A · 2017
Later among the works it cites.
Adversarial generation of natural language
Subramanian, S., Rajeswar, S., Dutil, F., Pal, C., and Courville, A · 2017
Later among the works it cites.
Learning to generate long-term future via hierarchical prediction
Villegas, R., Yang, J., Zou, Y., Sohn, S., Lin, X., and Lee, H · 2017
Later among the works it cites.