Fetching the paper…
Reading the bibliography…
Designing deep neural networks is an art that often involves an expensive search over candidate architectures.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Optimal methods of smooth convex minimization
Arkaddii S Nemirovskii and Yu E Nesterov · 1985
Earlier work this paper cites.
Hybrid monte carlo
Simon Duane, Anthony D Kennedy, Brian J Pendleton, and Duncan Roweth · 1987
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Timit acoustic phonetic continuous speech corpus
John S Garofolo · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
Recurrent nets that time and count
Felix A Gers and Jürgen Schmidhuber · 2000
Earlier work this paper cites.
LSTM recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and E Schmidhuber · 2001
Earlier work this paper cites.
Sequence labelling in structured domains with hierarchical recurrent neural networks
Santiago Fernández, Alex Graves, and Jürgen Schmidhuber · 2007
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Structured sparse coding via lateral inhibition
Arthur D Szlam, Karol Gregor, and Yann L Cun · 2011
Earlier work this paper cites.
The Langevin equation: with applications to stochastic problems in physics, chemistry and electrical engineering
William Coffey and Yu P Kalmykov · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Advances in optimizing recurrent networks
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu · 2013
Earlier work this paper cites.
A fast proximal method for convolutional sparse coding
Rakesh Chalasani, Jose C Principe, and Naveen Ramakrishnan · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Haşim Sak, Andrew Senior, and Françoise Beaufays · 2014
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Syllable-based acoustic modeling with CTC-SMBR-LSTM
Zhongdi Qu, Parisa Haghani, Eugene Weinstein, and Pedro Moreno · 2017
Later among the works it cites.
Lstm and qrnn language model toolkit for pytorch
Salesforce · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Later among the works it cites.
Orthogonal recurrent neural networks with scaled Cayley transform
Kyle Helfrich, Devin Willmott, and Qiang Ye · 2018
Later among the works it cites.
Fastgrnn: A fast, accurate, stable and tiny kilobyte sized gated recurrent neural network
Aditya Kusupati, Manish Singh, Kush Bhatia, Ashish Kumar, Prateek Jain, and Manik Varma · 2018
Later among the works it cites.
Independently recurrent neural network (indrnn): Building a longer and deeper rnn
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Improving performance of recurrent neural network with relu nonlinearity
Sachin S Talathi and Aniket Vartak · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Cited alongside, same era.
Learning mmse optimal thresholds for fista
US Kamilov and H Mansour · 2016
Cited alongside, same era.
A recurrent neural network without chaos
Thomas Laurent and James von Brecht · 2016
Cited alongside, same era.
Phased LSTM: Accelerating recurrent network training for long or event-based sequences
Daniel Neil, Michael Pfeiffer, and Shih-Chii Liu · 2016
Cited alongside, same era.
Deep sentence embedding using long short-term memory networks: Analysis and application to information retrieval
Hamid Palangi, Li Deng, Yelong Shen, Jianfeng Gao, Xiaodong He, Jianshu Chen, Xinying Song, and Rabab Ward · 2016
Cited alongside, same era.
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao · 2018
Later among the works it cites.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Later among the works it cites.
The unreasonable effectiveness of the forget gate
Jos Van Der Westhuizen and Joan Lasenby · 2018
Later among the works it cites.
Optimization with orthogonal constraints and on general manifolds
Mario Lezcano Casado · 2019
Later among the works it cites.
Trivializations for gradient-based optimization on manifolds
Mario Lezcano Casado · 2019
Later among the works it cites.
Towards non-saturating recurrent units for modelling long-term dependencies
Sarath Chandar, Chinnadhurai Sankar, Eugene Vorontsov, Samira Ebrahimi Kahou, and Yoshua Bengio · 2019
Later among the works it cites.
Antisymmetricrnn: A dynamical system view on recurrent neural networks
Bo Chang, Minmin Chen, Eldad Haber, and Ed H Chi · 2019
Later among the works it cites.
Symplectic recurrent neural networks
Zhengdao Chen, Jianyu Zhang, Martin Arjovsky, and Léon Bottou · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2019
Later among the works it cites.
RNNs evolving in equilibrium: A solution to the vanishing and exploding gradients
Anil Kag, Ziming Zhang, and Venkatesh Saligrama · 2019
Later among the works it cites.
Direct steering of de novo molecular generation using descriptor conditional recurrent neural networks (crnns)
Panagiotis-Christos Kotsias, Josep Arús-Pous, Hongming Chen, Ola Engkvist, Christian Tyrchan, and Esben Jannik Bjerrum · 2019
Later among the works it cites.
Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group
Mario Lezcano-Casado and David Martínez-Rubio · 2019
Later among the works it cites.
Recurrent neural networks in the eye of differential equations
Murphy Yuezhen Niu, Lior Horesh, and Isaac Chuang · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Scheduled restart momentum for accelerated stochastic gradient descent
Bao Wang, Tan M Nguyen, Andrea L Bertozzi, Richard G Baraniuk, and Stanley J Osher · 2020
Closest in time.