Fetching the paper…
Reading the bibliography…
Most language modeling methods rely on large-scale data to statistically learn the sequential patterns of words.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Slava Katz. 1987 · 1987
Earlier work this paper cites.
A statistical approach to machine translation
Peter F Brown, John Cocke, Stephen A Della Pietra, Vincent J Della Pietra, Fredrick Jelinek, John D Lafferty, Robert L Mercer, and Paul S Roossin. 1990 · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Ann Marcinkiewicz, Mary. 1993 · 1993
Earlier work this paper cites.
A linguistically motivated probabilistic model of information retrieval
Djoerd Hiemstra. 1998 · 1998
Earlier work this paper cites.
A language modeling approach to information retrieval
Jay M Ponte and W Bruce Croft. 1998 · 1998
Earlier work this paper cites.
Information retrieval as statistical translation
Adam Berger and John Lafferty. 1999 · 1999
Earlier work this paper cites.
Products of experts
G. E Hinton. 1999 · 1999
Earlier work this paper cites.
A hidden markov model information retrieval system
David RH Miller, Tim Leek, and Richard M Schwartz. 1999 · 1999
Earlier work this paper cites.
Headline generation based on statistical translation
Michele Banko, Vibhu O Mittal, and Michael J Witbrock. 2000 · 2000
Earlier work this paper cites.
Speech & language processing
Dan Jurafsky. 2000 · 2000
Earlier work this paper cites.
Classes for fast maximum entropy training
J Goodman. 2001 · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton. 2002 · 2002
Earlier work this paper cites.
Word similarity computing based on hownet
Qun Liu. 2002 · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Frederic Morin and Yoshua Bengio. 2005 · 2005
Cited alongside, same era.
Hownet and the computation of meaning (with Cd-rom)
Zhendong Dong and Qiang Dong. 2006 · 2006
Cited alongside, same era.
Product of gaussians for speech recognition
M. J. F. Gales and S. S. Airey. 2006 · 2006
Cited alongside, same era.
Large language models in machine translation
Thorsten Brants, Ashok C Popat, Peng Xu, Franz J Och, and Jeffrey Dean. 2007 · 2007
Cited alongside, same era.
A scalable hierarchical distributed language model
Andriy Mnih and Geoffrey Hinton. 2008 · 2008
Cited alongside, same era.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Cited alongside, same era.
Neural headline generation with minimum risk training
Shiqi Shen Ayana, Zhiyuan Liu, and Maosong Sun. 2016 · 2016
Later among the works it cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Later among the works it cites.
Incorporating copying mechanism in sequence-to-sequence learning
Jiatao Gu, Zhengdong Lu, Hang Li, and Victor OK Li. 2016 · 2016
Later among the works it cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017 · 2017
Later among the works it cites.
Exploration of tree-based hierarchical softmax for recurrent language models
Nan Jiang, Wenge Rong, Min Gao, Yikang Shen, Zhang Xiong, Nan Jiang, Wenge Rong, Min Gao, Yikang Shen, and Zhang Xiong. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-aspect sentiment analysis for chinese online social reviews based on topic modeling and hownet lexicon
Xianghua Fu, Guo Liu, Yanyan Guo, and Zhiqiang Wang. 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya. Sutskever, and Oriol. Vinyals. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Stephen Merity, Bryan Mccann, and Richard Socher. 2017 · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Later among the works it cites.
Lexical sememe prediction via word embeddings and matrix factorization
Ruobing Xie, Xingchi Yuan, Zhiyuan Liu, and Maosong Sun. 2017 · 2017
Later among the works it cites.
Incorporating chinese characters of words for lexical sememe prediction
Huiming Jin, Hao Zhu, Zhiyuan Liu, Ruobing Xie, Maosong Sun, Fen Lin, and Leyu Lin. 2018 · 2018
Closest in time.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Closest in time.
Cross-lingual lexical sememe prediction
Fanchao Qi, Yankai Lin, Maosong Sun, Hao Zhu, Ruobing Xie, and Zhiyuan Liu. 2018 · 2018
Closest in time.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Closest in time.
Improved word representation learning with sememes
Yilin Niu, Ruobing Xie, Zhiyuan Liu, and Maosong Sun. 2017 · 2058
Closest in time.