Fetching the paper…
Reading the bibliography…
This research note combines two methods that have recently improved the state of the art in language modeling: Transformers and dynamic evaluation.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R. (2019) · 1901
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
The human knowledge compression prize
Hutter, M. (2006) · 2006
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernockỳ, J., and Khudanpur, S. (2010) · 2010
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A. (2013) · 2013
Earlier work this paper cites.
Multiplicative LSTM for sequence modelling
Krause, B., Lu, L., Murray, I., and Renals, S. (2016) · 2016
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y. (2017) · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D. (2017) · 2017
Cited alongside, same era.
Hypernetworks
Ha, D., Dai, A., and Lee, Q. (2017) · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2017) · 2017
Cited alongside, same era.
Fast-slow recurrent neural networks
Mujika, A., Meier, F., and Steger, A. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Recurrent highway networks
Zilly, J. G., Srivastava, R. K., Koutník, J., and Schmidhuber, J. (2017) · 2017
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Krause, B., Kahembwe, E., Murray, I., and Renals, S. (2018) · 2018
Later among the works it cites.
Generating wikipedia by summarizing long sequences
Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N. (2018) · 2018
Later among the works it cites.
An analysis of neural language modeling at multiple scales
Merity, S., Keskar, N. S., and Socher, R. (2018) · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018) · 2018
Later among the works it cites.
Fast parametric learning with activation memorization
Rae, J. W., Dyer, C., Dayan, P., and Lillicrap, T. P. (2018) · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Al-Rfou, R., Choe, D., Constant, N., Guo, M., and Jones, L. (2018) · 2018
Cited alongside, same era.
Sharp nearby, fuzzy far away: How neural language models use context
Khandelwal, U., He, H., Qi, P., and Jurafsky, D. (2018) · 2018
Cited alongside, same era.
Efficient softmax approximation for GPUs
Grave, E., Joulin, A., Cissé, M., Jégou, H., et al. (2017a)
Cited in the paper.
Improving neural language models with a continuous cache
Grave, E., Joulin, A., and Usunier, N. (2017b)
Cited in the paper.
Baevski, A. and Auli, M. (2019) · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Closest in time.