Fetching the paper…
Reading the bibliography…
Although SGD requires shuffling the training data between epochs, currently none of the word-level language modeling systems do this.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Direct output connection for a high-rank language model
Sho Takase, Jun Suzuki, and Masaaki Nagata. 2018 · 2018
Cited alongside, same era.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…