Fetching the paper…
Reading the bibliography…
We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
The language instinct, 1994
Steven Pinker · 1994
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney · 1995
Earlier work this paper cites.
Switchboard-1 release 2
John J Godfrey and Edward Holliman · 1997
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Recurrent neural networks are universal approximators
Anton Maximilian Schäfer and Hans Georg Zimmermann · 2006
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Andriy Mnih and Geoffrey Hinton · 2007
Earlier work this paper cites.
Numerical recipes 3rd edition: The art of scientific computing
William H Press · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Low rank language models for small training sets
Brian Hutchinson, Mari Ostendorf, and Maryam Fazel · 2011
Earlier work this paper cites.
Large text compression benchmark, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
A sparse plus low rank maximum entropy language model
Brian Hutchinson, Mari Ostendorf, and Maryam Fazel · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig · 2012
Earlier work this paper cites.
Subword language modeling with neural networks
Tomáš Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, Stefan Kombrink, and Jan Cernocky · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus · 2013
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Later among the works it cites.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Later among the works it cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2016
Later among the works it cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush · 2016
Later among the works it cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning stochastic recurrent networks
Justin Bayer and Christian Osendorfer · 2014
Cited alongside, same era.
Neural word embedding as implicit matrix factorization
Omer Levy and Yoav Goldberg · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
A recurrent latent variable model for sequential data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio · 2015
Cited alongside, same era.
Deep temporal sigmoid belief networks for sequence modeling
Zhe Gan, Chunyuan Li, Ricardo Henao, David E Carlson, and Lawrence Carin · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra · 2015
Cited alongside, same era.
Generalizing and hybridizing count-based and neural language models
Graham Neubig and Chris Dyer · 2016
Later among the works it cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Later among the works it cites.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2017
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Closest in time.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Closest in time.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Closest in time.
Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi · 2017
Closest in time.