Fetching the paper…
Reading the bibliography…
Our work addresses two important issues with recurrent neural networks: (1) they are over-parameterized, and (2) the recurrence matrix is ill-conditioned.
Distance measures for speech recognition, psychological and instrumental
Paul Mermelstein · 1976
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1
John S Garofolo, Lori F Lamel, William M Fisher, Jonathon G Fiscus, and David S Pallett · 1993
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The ubiquitous kronecker product
Charles F Van Loan · 2000
Earlier work this paper cites.
The “echo state” approach to analysing and training recurrent neural networks-with an erratum note
Herbert Jaeger · 2001
Earlier work this paper cites.
Framewise phoneme classification with bidirectional lstm and other neural network architectures
Alex Graves and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
Theano: Deep learning on gpus with python
James Bergstra, Olivier Breuleux, Pascal Lamblin, Razvan Pascanu, Olivier Delalleau, Guillaume Desjardins, Ian Goodfellow, Arnaud Bergeron, Yoshua Bengio, and Pack Kaelbling · 2011
Earlier work this paper cites.
Nicolas Boulanger-Lewandowski, Yoshua Bengio, and Pascal Vincent · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Statistical Language Models Based on Neural Networks
Tomáš Mikolov · 2012
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Nando de Freitas, et al · 2013
Cited alongside, same era.
Fastfood-approximating kernel expansions in loglinear time
Quoc Le, Tamás Sarlós, and Alex Smola · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Low precision storage for deep learning
Matthieu Courbariaux, Jean-Pierre David, and Yoshua Bengio · 2014
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Later among the works it cites.
Distributed second-order optimization using kronecker-factored approximations
Jimmy Ba, Roger Grosse, and James Martens · 2016
Later among the works it cites.
Low-rank passthrough neural networks
Antonio Valerio Miceli Barone · 2016
Later among the works it cites.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Later among the works it cites.
Hypernetworks
David Ha, Andrew Dai, and Quoc Le · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Compressing neural networks with the hashing trick
Wenlin Chen, James Wilson, Stephen Tyree, Kilian Weinberger, and Yixin Chen · 2015
Cited alongside, same era.
Gated feedback recurrent neural networks
Junyoung Chung, Caglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Cited alongside, same era.
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2016
Later among the works it cites.
Orthogonal RNNs and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Later among the works it cites.
Scalable metric learning via weighted approximate rank component analysis
Cijo Jose and François Fleuret · 2016
Later among the works it cites.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman, and James Bailey · 2016
Later among the works it cites.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Later among the works it cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Closest in time.
Tunable efficient unitary neural networks (EUNN) and their application to RNNs
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljačić · 2017
Closest in time.
Pytorch, 2017
Adam Paszke, Sam Gross, and Soumith Chintala · 2017
Closest in time.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Closest in time.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Shun Liao, Roger Grosse, and Jimmy Ba · 2017
Closest in time.