Fetching the paper…
Reading the bibliography…
Building accurate language models that capture meaningful long-term dependencies is a core challenge in natural language processing.
Prediction and entropy of printed english
Claude E Shannon · 1951
Earlier work this paper cites.
The well-calibrated bayesian
A. P. Dawid · 1982
Earlier work this paper cites.
The impossibility of inductive inference
A. P. Dawid · 1985
Earlier work this paper cites.
Fundamentals of Statistical Exponential Families: With Applications in Statistical Decision Theory
L. D. Brown · 1986
Earlier work this paper cites.
Prediction in the worst case
D. P. Foster · 1991
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Proposal for a mutual-information based language model
Uwe Jost and ES Atwell · 1994
Earlier work this paper cites.
Improvements in beam search
Volker Steinbiss, Bach-Hiep Tran, and Hermann Ney · 1994
Earlier work this paper cites.
Language model representations for beam-search decoding
Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, and Marcello Federico · 1995
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Alfred Müller · 1997
Earlier work this paper cites.
Towards improved language model evaluation measures
Philip Clarkson and Tony Robinson · 1999
Earlier work this paper cites.
Calibrated forecasting and merging
E. Kalai, E. Lehrer, and R. Smorodinsky · 1999
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John C. Platt · 1999
Earlier work this paper cites.
Look-ahead techniques for fast beam search
Stefan Ortmanns and Hermann Ney · 2000
Earlier work this paper cites.
Competitive on-line statistics
V. Vovk · 2001
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates, 2002
Bianca Zadrozny and Charles Elkan · 2002
Cited alongside, same era.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana · 2005
Cited alongside, same era.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Cited alongside, same era.
On integral probability metrics, phi-divergences and binary classification
Bharath Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Gert Lanckriet · 2009
Cited alongside, same era.
Information theory: coding theorems for discrete memoryless systems
Imre Csiszar and János Körner · 2011
Cited alongside, same era.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig · 2012
Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Later among the works it cites.
Fisher gan
Youssef Mroueh and Tom Sercu · 2017
Later among the works it cites.
Topic compositional neural language model
Wenlin Wang, Zhe Gan, Wenqi Wang, Dinghan Shen, Jiaji Huang, Wei Ping, Sanjeev Satheesh, and Lawrence Carin · 2017
Later among the works it cites.
Transformer-xl: Language modeling with longer-term dependency
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2018
Later among the works it cites.
Frage: frequency-agnostic word representation
Chengyue Gong, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Document context language models
Yangfeng Ji, Trevor Cohn, Lingpeng Kong, Chris Dyer, and Jacob Eisenstein · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Larger-context language modelling with recurrent neural network
Tian Wang and Kyunghyun Cho · 2016
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Sparse attentive backtracking: Temporal credit assignment through reminding
Nan Rosemary Ke, Anirudh Goyal, Olexa Bilaniuk, Jonathan Binas, Michael C Mozer, Chris Pal, and Yoshua Bengio · 2018
Later among the works it cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky · 2018
Later among the works it cites.
Information theoretic co-training
David McAllester · 2018
Later among the works it cites.
An Analysis of Neural Language Modeling at Multiple Scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
Cross entropy of neural language models at infinity—a new bound of the entropy rate
Shuntaro Takahashi and Kumiko Tanaka-Ishii · 2018
Later among the works it cites.
Direct output connection for a high-rank language model
Sho Takase, Jun Suzuki, and Masaaki Nagata · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trieu H Trinh, Andrew M Dai, Thang Luong, and Quoc V Le · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.