Fetching the paper…
Reading the bibliography…
Two potential bottlenecks on the expressiveness of recurrent neural networks (RNNs) are their ability to store information about the task in their parameters, and to store information about the input history in their units.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M Cover · 1965
Earlier work this paper cites.
Number of stable points for spin-glasses and neural networks of higher orders
Pierre Baldi and Santosh S Venkatesh · 1987
Earlier work this paper cites.
The space of interactions in neural network models
Elizabeth Gardner · 1988
Earlier work this paper cites.
Universality of fully connected recurrent neural networks
Kenji Doya · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Vapnik-chervonenkis dimension of recurrent neural networks
Pascal Koiran and Eduardo D Sontag · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A. Gers, Jurgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
Real-time computing without stable states: A new framework for neural computation based on perturbations
Wolfgang Maass, Thomas Natschläger, and Henry Markram · 2002
Earlier work this paper cites.
Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication
Herbert Jaeger and Harald Haas · 2004
Earlier work this paper cites.
Short-term memory in orthogonal neural networks
Olivia L White, Daniel D Lee, and Haim Sompolinsky · 2004
Earlier work this paper cites.
Memory traces in dynamical systems
Surya Ganguli, Dongsung Huh, and Haim Sompolinsky · 2008
Earlier work this paper cites.
A novel connectionist system for unconstrained handwriting recognition
Alex Graves, Marcus Liwicki, Santiago Fernández, Roman Bertolami, Horst Bunke, and Jürgen Schmidhuber · 2009
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey. Hinton · 2012
Cited alongside, same era.
Context-dependent computation by recurrent dynamics in prefrontal cortex
Valerio Mante, David Sussillo, Krishna V Shenoy, and William T Newsome · 2013
Cited alongside, same era.
Opening the black box: low-dimensional dynamics in high-dimensional recurrent neural networks
David Sussillo and Omri Barak · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Short-term memory capacity in networks via the restricted isometry property
Adam S Charles, Han Lun Yap, and Christopher J Rozell · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
An empirical exploration of recurrent network architectures
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever · 2015
Later among the works it cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Fei-Fei Li · 2015
Later among the works it cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Later among the works it cites.
Deep knowledge tracing
Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J Guibas, and Jascha Sohl-Dickstein · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Parallelizing exploration-exploitation tradeoffs in gaussian process bandit optimization
Thomas Desautels, Andreas Krause, and Joel W Burdick · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan C. Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Erich Elsen, Jesse Engel, Linxi Fan, Christopher Fougner, Tony Han, Awni Y. Hannun, Billy Jun, Patrick LeGresley, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Yi Wang, Zhiqian Wang, Chong Wang, Bo Xiao, Dani Yogatama, Jun Zhan, and Zhenyao Zhu · 2015
Cited alongside, same era.
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber · 2015
Cited alongside, same era.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Later among the works it cites.
Deep fried convnets
Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang · 2015
Later among the works it cites.
Nanoconnectomic upper bound on the variability of synaptic plasticity
Thomas M Bartol, Cailey Bromer, Justin Kinney, Michael A Chirillo, Jennifer N Bourne, Kristen M Harris, and Terrence J Sejnowski · 2016
Closest in time.
Intelligible language modeling with input switched affine networks
Jakob Foerster, Justin Gilmer, Jan Chorowski, Jascha Sohl-Dickstein, and David Sussillo · 2016
Closest in time.
Itay Hubara, Daniel Soudry, and Ran El Yaniv · 2016
Closest in time.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Closest in time.
Large text compression benchmark: About the test data, 2011
Matt Mahoney · 2016
Closest in time.
Minimal gated unit for recurrent neural networks
Guo-Bing Zhou, Jianxin Wu, Chen-Lin Zhang, and Zhi-Hua Zhou · 2016
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Closest in time.