Fetching the paper…
Reading the bibliography…
The capacity of a neural network to absorb information is limited by its number of parameters.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
Improved backingoff for m-gram language modeling., 1995
Reinhard Kneser and Hermann. Ney · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A. Gers, Jürgen A. Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
Mixtures of Gaussian Processes
Volker Tresp · 2001
Earlier work this paper cites.
A parallel mixture of SVMs for very large scale problems
Ronan Collobert, Samy Bengio, and Yoshua Bengio · 2002
Earlier work this paper cites.
Infinite mixtures of Gaussian process experts
Carl Edward Rasmussen and Zoubin Ghahramani · 2002
Earlier work this paper cites.
Nonlinear models using dirichlet process mixtures
Babak Shahbaba and Radford Neal · 2009
Earlier work this paper cites.
Hierarchical mixture of classification experts uncovers interactions between brain regions
Bangpeng Yao, Dirk Walther, Diane Beck, and Li Fei-fei · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization, 2010
John Duchi, Elad Hazan, and Yoram Singer · 2010
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V. Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S. Corrado, Jeffrey Dean, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Japanese and Korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Cited alongside, same era.
Low-rank approximations for conditional feedforward computation in deep neural networks
Andrew Davis and Itamar Arel · 2013
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Later among the works it cites.
Distributed Gaussian processes
Marc Peter Deisenroth and Jun Wei Ng · 2015
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Later among the works it cites.
Generative image modeling using spatial LSTMs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Exponentially Increasing the Capacity-to-Computation Ratio for Conditional Computation in Deep Learning
K. Cho and Y. Bengio · 2014
Cited alongside, same era.
Edinburgh’s phrase-based machine translation systems for wmt-14
Nadir Durrani, Barry Haddow, Philipp Koehn, and Kenneth Heafield · 2014
Cited alongside, same era.
Deep sequential neural network
Patrick Gallinari Ludovic Denoyer · 2014
Cited alongside, same era.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Hasim Sak, Andrew W Senior, and Françoise Beaufays · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Lucas Theis and Matthias Bethge · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Gregory S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian J. Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Józefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Gordon Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul A. Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda B. Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Later among the works it cites.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars · 2016
Later among the works it cites.
Ensemble learning for multi-source neural machine translation
Ekaterina Garmash and Christof Monz · 2016
Later among the works it cites.
Memory-efficient backpropagation through time
Audrunas Gruslys, Rémi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Later among the works it cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda B. Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Later among the works it cites.
Deep recurrent models with fast-forward connections for neural machine translation
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu · 2016
Later among the works it cites.