Fetching the paper…
Reading the bibliography…
Recent neural network sequence models with softmax classifiers have achieved their best language modeling performance only with very large hidden states and large vocabularies.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, Mitchell P., Santorini, Beatrice, and Marcinkiewicz, Mary Ann · 1993
Earlier work this paper cites.
A Maximum Entropy Approach to Adaptive Statistical Language Modeling
Rosenfeld, Roni · 1996
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Moses: Open Source Toolkit for Statistical Machine Translation
Koehn, Philipp, Hoang, Hieu, Birch, Alexandra, Callison-Burch, Chris, Federico, Marcello, Bertoldi, Nicola, Cowan, Brooke, Shen, Wade, Moran, Christine, Zens, Richard, Dyer, Chris, Bojar, Ondřej, Constantin, Alexandra, and Herbst, Evan · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukás, Cernocký, Jan, and Khudanpur, Sanjeev · 2010
Earlier work this paper cites.
Context dependent recurrent neural network language model
Mikolov, Tomas and Zweig, Geoffrey · 2012
Earlier work this paper cites.
One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling
Chelba, Ciprian, Mikolov, Tomas, Schuster, Mike, Ge, Qi, Brants, Thorsten, Koehn, Phillipp, and Robinson, Tony · 2013
Earlier work this paper cites.
Language Modeling with Sum-Product Networks
Cheng, Wei-Chen, Kok, Stanley, Pham, Hoai Vu, Chieu, Hai Leong, and Chai, Kian Ming Adam · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
Gal, Yarin · 2015
Cited alongside, same era.
End-To-End Memory Networks
Sukhbaatar, Sainbayar, Szlam, Arthur, Weston, Jason, and Fergus, Rob · 2015
Cited alongside, same era.
Pointer networks
Vinyals, Oriol, Fortunato, Meire, and Jaitly, Navdeep · 2015
Cited alongside, same era.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Adi, Yossi, Kermany, Einat, Belinkov, Yonatan, Lavi, Ofer, and Goldberg, Yoav · 2016
Cited alongside, same era.
Gülçehre, Çaglar, Ahn, Sungjin, Nallapati, Ramesh, Zhou, Bowen, and Bengio, Yoshua · 2016
Closest in time.
Text Understanding with the Attention Sum Reader Network
Kadlec, Rudolf, Schmid, Martin, Bajgar, Ondrej, and Kleindienst, Jan · 2016
Closest in time.
Character-aware neural language models
Kim, Yoon, Jernite, Yacine, Sontag, David, and Rush, Alexander M · 2016
Closest in time.
Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
Krueger, David, Maharaj, Tegan, Kramár, János, Pezeshki, Mohammad, Ballas, Nicolas, Ke, Nan Rosemary, Goyal, Anirudh, Bengio, Yoshua, Larochelle, Hugo, Courville, Aaron, et al · 2016
Closest in time.
Ask me anything: Dynamic memory networks for natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahn, Sungjin, Choi, Heeyoul, Pärnamaa, Tanel, and Bengio, Yoshua · 2016
Cited alongside, same era.
Long Short-Term Memory-Networks for Machine Reading
Cheng, Jianpeng, Dong, Li, and Lapata, Mirella · 2016
Cited alongside, same era.
Incorporating Copying Mechanism in Sequence-to-Sequence Learning
Gu, Jiatao, Lu, Zhengdong, Li, Hang, and Li, Victor O. K · 2016
Cited alongside, same era.
How to Construct Deep Recurrent Neural Networks
Pascanu, Razvan, Çaglar Gülçehre, Cho, Kyunghyun, and Bengio, Yoshua
Cited in the paper.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua
Cited in the paper.
Kumar, Ankit, Irsoy, Ozan, Ondruska, Peter, Iyyer, Mohit, Bradbury, James, Gulrajani, Ishaan, Zhong, Victor, Paulus, Romain, and Socher, Richard · 2016
Closest in time.
Latent Predictor Networks for Code Generation
Ling, Wang, Grefenstette, Edward, Hermann, Karl Moritz, Kociský, Tomás, Senior, Andrew, Wang, Fumin, and Blunsom, Phil · 2016
Closest in time.
Dynamic Memory Networks for Visual and Textual Question Answering
Xiong, Caiming, Merity, Stephen, and Socher, Richard · 2016
Closest in time.
Zilly, Julian Georg, Srivastava, Rupesh Kumar, Koutník, Jan, and Schmidhuber, Jürgen · 2016
Closest in time.