Fetching the paper…
Reading the bibliography…
Many of the leading approaches in language modeling introduce novel, complex and specialized architectures.
Backpropagation through time: what it does and how to do it
Werbos, P. J · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Williams, R. J. and Peng, J · 1990
Earlier work this paper cites.
The Penn Treebank: annotating predicate argument structure
Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B · 1994
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Hierarchical Probabilistic Neural Network Language Model
Morin, F. and Bengio, Y · 2005
Earlier work this paper cites.
Statistical Machine Translation
Koehn, P · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocký, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Earlier work this paper cites.
Automatic speech recognition: A deep learning approach
Yu, D. and Deng, L · 2014
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y · 2016
Cited alongside, same era.
Language modeling with Gated Convolutional Networks
Dauphin, Y., Fan, A., Auli, M., and Grangier, D · 2016
Cited alongside, same era.
Ha, D., Dai, A., and Le, Q. V · 2016
Cited alongside, same era.
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
Inan, H., Khosravi, K., and Socher, R · 2016
Cited alongside, same era.
Zoneout: Regularizing RNNss by randomly preserving hidden activations
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
Fast-Slow Recurrent Neural Networks
Mujika, A., Meier, F., and Steger, A · 2017
Later among the works it cites.
Deep contextualized word representations
P. Matthew, M. Neumann, M. Iyyer M. Gardner C. Clark K. Lee L. Zettlemoyer · 2017
Later among the works it cites.
Learning to generate reviews and discovering sentiment
Radford, A., Jozefowicz, R., and Sutskever, I · 2017
Later among the works it cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Krueger, D., Maharaj, T., Kramár, J., Pezeshki, M., Ballas, N., Ke, N., Goyal, A., Bengio, Y., Larochelle, H., Courville, A., et al · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2016
Cited alongside, same era.
Rocki, K., Kornuta, T., and Maharaj, T · 2016
Cited alongside, same era.
Zilly, J. G., Srivastava, R. K., Koutník, J., and Schmidhuber, J · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2016
Cited alongside, same era.
Quasi-Recurrent Neural Networks
Bradbury, J., Merity, S., Xiong, C., and Socher, R · 2017
Cited alongside, same era.
Improving Neural Language Models with a Continuous Cache
Grave, E., Joulin, A., and Usunier, N
Cited in the paper.
Efficient softmax approximation for GPUs
Grave, Edouard, Joulin, Armand, Cissé, Moustapha, Grangier, David, and Jégou, Hervé
Cited in the paper.
The Human Knowledge Compression Contest
Hutter, M · 2018
Closest in time.
Dynamic Evaluation of Neural Sequence Models
Krause, B., Kahembwe, E., Murray, I., and Renals, S · 2018
Closest in time.
On the State of the Art of Evaluation in Neural Language Models
Melis, G., Dyer, C., and Blunsom, P · 2018
Closest in time.
Regularizing and Optimizing LSTM Language Models
Merity, S., Keskar, N., and Socher, R · 2018
Closest in time.
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
Yang, Z., Dai, Z., Salakhutdinov, R., and Cohen, W · 2018
Closest in time.