Fetching the paper…
Reading the bibliography…
We propose a new self-organizing hierarchical softmax formulation for neural-network-based language models over large vocabularies.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Classes for fast maximum entropy training
Joshua Goodman. 2001 · 2001
Earlier work this paper cites.
Bbn/umd at duc-2004: Topiary
David Zajic, Bonnie Dorr, and Richard Schwartz. 2004 · 2004
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Yoshua Bengio and Jean-Sébastien Senécal. 2008 · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
Andriy Mnih and Geoffrey E Hinton. 2009 · 2009
Earlier work this paper cites.
Theano: A cpu and gpu math compiler in python
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Structured output layer neural network language model
Hai-Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, and François Yvon. 2011 · 2011
Earlier work this paper cites.
Strategies for training large scale neural network language models
Tomáš Mikolov, Anoop Deoras, Daniel Povey, Lukáš Burget, and Jan Černockỳ. 2011a · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2011b · 2011
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
Andriy Mnih and Yee Whye Teh. 2012 · 2012
Cited alongside, same era.
Annotated gigaword
Courtney Napoles, Matthew Gormley, and Benjamin Van Durme. 2012 · 2012
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Later among the works it cites.
Strategies for training large vocabulary neural language models
Welin Chen, David Grangier, and Michael Auli. 2016 · 2016
Later among the works it cites.
Abstractive sentence summarization with attentive recurrent neural networks
Sumit Chopra, Michael Auli, Alexander M Rush, and SEAS Harvard. 2016 · 2016
Later among the works it cites.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache. 2016 · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Later among the works it cites.
Language as a latent variable: Discrete generative models for sentence compression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a · 2013
Cited alongside, same era.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Learning longer memory in recurrent neural networks
Tomas Mikolov, Armand Joulin, Sumit Chopra, Michael Mathieu, and Marc’Aurelio Ranzato. 2014 · 2014
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sebastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003a
Cited in the paper.
Quick training of probabilistic neural nets by importance sampling
Yoshua Bengio, Jean-Sébastien Senécal, et al. 2003b
Cited in the paper.
Yishu Miao and Phil Blunsom. 2016 · 2016
Later among the works it cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016 · 2016
Later among the works it cites.
Neural headline generation with sentence-wise optimization
Shiqi Shen, Yu Zhao, Zhiyuan Liu, Maosong Sun, et al. 2016 · 2016
Later among the works it cites.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou. 2017 · 2017
Closest in time.