Fetching the paper…
Reading the bibliography…
We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies.
Human behavior and the principle of least effort
Zipf, George Kingsley · 1949
Earlier work this paper cites.
A maximum likelihood approach to continuous speech recognition
Bahl, Lalit R, Jelinek, Frederick, and Mercer, Robert L · 1983
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Katz, Slava M · 1987
Earlier work this paper cites.
Finding structure in time
Elman, Jeffrey L · 1990
Earlier work this paper cites.
A cache-based natural language model for speech recognition
Kuhn, Roland and De Mori, Renato · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Werbos, Paul J · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Williams, Ronald J and Peng, Jing · 1990
Earlier work this paper cites.
Class-based n-gram models of natural language
Brown, Peter F, Desouza, Peter V, Mercer, Robert L, Pietra, Vincent J Della, and Lai, Jenifer C · 1992
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Kneser, Reinhard and Ney, Hermann · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Koehn, Philipp · 2005
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
Continuous space language models
Schwenk, Holger · 2007
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Bengio, Yoshua and Senécal, Jean-Sébastien · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
Mnih, Andriy and Hinton, Geoffrey E · 2009
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, Michael and Hyvärinen, Aapo · 2010
Cited alongside, same era.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernockỳ, Jan, and Khudanpur, Sanjeev · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Structured output layer neural network language model
Le, Hai-Son, Oparin, Ilya, Allauzen, Alexandre, Gauvain, Jean-Luc, and Yvon, François · 2011
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
When and why are log-linear models self-normalizing
Andreas, Jacob and Klein, Dan · 2014
Later among the works it cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, KyungHyun, and Bengio, Yoshua · 2014
Later among the works it cites.
Fast and robust neural network joint models for statistical machine translation
Devlin, Jacob, Zbib, Rabih, Huang, Zhongqiang, Lamar, Thomas, Schwartz, Richard M, and Makhoul, John · 2014
Later among the works it cites.
Learning longer memory in recurrent neural networks
Mikolov, Tomas, Joulin, Armand, Chopra, Sumit, Mathieu, Michael, and Ranzato, Marc’Aurelio · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hinton, Geoffrey, Deng, Li, Yu, Dong, Dahl, George E, Mohamed, Abdel-rahman, Jaitly, Navdeep, Senior, Andrew, Vanhoucke, Vincent, Nguyen, Patrick, Sainath, Tara N, et al · 2012
Cited alongside, same era.
Context dependent recurrent neural network language model
Mikolov, Tomas and Zweig, Geoffrey · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
Mnih, Andriy and Teh, Yee Whye · 2012
Cited alongside, same era.
Large, pruned or continuous space language models on a gpu for statistical machine translation
Schwenk, Holger, Rousseau, Anthony, and Attik, Mohammed · 2012
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, Ciprian, Mikolov, Tomas, Schuster, Mike, Ge, Qi, Brants, Thorsten, Koehn, Phillipp, and Robinson, Tony · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Graves, Alan, Mohamed, Abdel-rahman, and Hinton, Geoffrey · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Mikolov, Tomas, Chen, Kai, Corrado, Greg, and Dean, Jeffrey · 2013
Cited alongside, same era.
Chen, Welin, Grangier, David, and Auli, Michael · 2015
Later among the works it cites.
On using very large target vocabulary for neural machine translation
Jean, Sebastien, Cho, Kyunghyun, Memisevic, Roland, and Bengio, Yoshua · 2015
Later among the works it cites.
Blackout: Speeding up recurrent neural network language models with very large vocabularies
Ji, Shihao, Vishwanathan, SVN, Satish, Nadathur, Anderson, Michael J, and Dubey, Pradeep · 2015
Later among the works it cites.
Learning visual features from large weakly supervised data
Joulin, Armand, van der Maaten, Laurens, Jabri, Allan, and Vasilache, Nicolas · 2015
Later among the works it cites.
A simple way to initialize recurrent networks of rectified linear units
Le, Quoc V, Jaitly, Navdeep, and Hinton, Geoffrey E · 2015
Later among the works it cites.
Sparse non-negative matrix language modeling for skip-grams
Shazeer, Noam, Pelemans, Joris, and Chelba, Ciprian · 2015
Later among the works it cites.
Efficient exact gradient update for training deep networks with very large sparse targets
Vincent, Pascal, de Brébisson, Alexandre, and Bouthillier, Xavier · 2015
Later among the works it cites.
Exploring the limits of language modeling
Jozefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Closest in time.