Fetching the paper…
Reading the bibliography…
Recurrent neural networks have gained widespread use in modeling sequence data across various domains.
Learning representations by back-propagating errors
Rumelhart, David E, Hinton, Geoffrey E, and Williams, Ronald J · 1986
Earlier work this paper cites.
Chaos in random neural networks
Sompolinsky, H., Crisanti, A., and Sommers, H. J · 1988
Earlier work this paper cites.
Finding structure in time
Elman, Jeffrey L · 1990
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P, Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 1993
Earlier work this paper cites.
Mean field theory for sigmoid belief networks
Saul, Lawrence K, Jaakkola, Tommi, and Jordan, Michael I · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Brown’s spectral distribution measure for r-diagonal elements in finite von neumann algebras
Haagerup, Uffe and Larsen, Flemming · 2000
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernockỳ, Jan, and Khudanpur, Sanjeev · 2010
Earlier work this paper cites.
Non-hermitian random matrix theory for mimo channels
Cakmak, Burak · 2012
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Neal, Radford M · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, Alex, Mohamed, Abdel-rahman, and Hinton, Geoffrey · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M, McClelland, James L, and Ganguli, Surya · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, KyungHyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Random walks: Training very deep nonlinear feed-forward networks with smart initialization
Sussillo, David and Abbott, LF · 2014
Cited alongside, same era.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Cited alongside, same era.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Session-based recommendations with recurrent neural networks
Hidasi, Balázs, Karatzoglou, Alexandros, Baltrunas, Linas, and Tikk, Domonkos · 2015
Cited alongside, same era.
An empirical exploration of recurrent network architectures
Jozefowicz, Rafal, Zaremba, Wojciech, and Sutskever, Ilya · 2015
Cited alongside, same era.
Lstm: A search space odyssey
Greff, Klaus, Srivastava, Rupesh K, Koutník, Jan, Steunebrink, Bas R, and Schmidhuber, Jürgen · 2017
Later among the works it cites.
Learning unitary operators with help from u (n)
Hyland, Stephanie L and Rätsch, Gunnar · 2017
Later among the works it cites.
Deep neural networks as gaussian processes
Lee, Jaehoon, Bahri, Yasaman, Novak, Roman, Schoenholz, Samuel S, Pennington, Jeffrey, and Sohl-Dickstein, Jascha · 2017
Later among the works it cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Pennington, Jeffrey, Schoenholz, Sam, and Ganguli, Surya · 2017
Later among the works it cites.
Deep Information Propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Later among the works it cites.
A correspondence between random neural networks and statistical field theory
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Skip-thought vectors
Kiros, Ryan, Zhu, Yukun, Salakhutdinov, Ruslan R, Zemel, Richard, Urtasun, Raquel, Torralba, Antonio, and Fidler, Sanja · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Le, Quoc V, Jaitly, Navdeep, and Hinton, Geoffrey E · 2015
Cited alongside, same era.
Mishkin, Dmytro and Matas, Jiri · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Arjovsky, Martin, Shah, Amar, and Bengio, Yoshua · 2016
Cited alongside, same era.
Ba, Jimmy Lei, Kiros, Jamie Ryan, and Hinton, Geoffrey E · 2016
Cited alongside, same era.
Capacity and trainability in recurrent neural networks
Collins, Jasmine, Sohl-Dickstein, Jascha, and Sussillo, David · 2016
Cited alongside, same era.
Daniely, A., Frostig, R., and Singer, Y · 2016
Cited alongside, same era.
Schoenholz, Samuel S, Pennington, Jeffrey, and Sohl-Dickstein, Jascha · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
Vorontsov, Eugene, Trabelsi, Chiheb, Kadoury, Samuel, and Pal, Chris · 2017
Later among the works it cites.
Recurrent recommender networks
Wu, Chao-Yuan, Ahmed, Amr, Beutel, Alex, Smola, Alexander J, and Jing, How · 2017
Later among the works it cites.
Xie, Di, Xiong, Jiang, and Pu, Shiliang · 2017
Later among the works it cites.
Mean field residual networks: On the edge of chaos
Yang, Greg and Schoenholz, Samuel S · 2017
Later among the works it cites.
How to start training: The effect of initialization and architecture
Hanin, Boris and Rolnick, David · 2018
Closest in time.
On the selection of initialization and activation function for deep neural networks
Hayou, Soufiane, Doucet, Arnaud, and Rousseau, Judith · 2018
Closest in time.
Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
Karakida, R., Akaho, S., and Amari, S.-i · 2018
Closest in time.
The emergence of spectral universality in deep networks
Pennington, Jeffrey, Schoenholz, Samuel S., and Ganguli, Surya · 2018
Closest in time.
Can recurrent neural networks warp time?
Tallec, Corentin and Ollivier, Yann · 2018
Closest in time.
Deep mean field theory: Layerwise variance and width variation as methods to control gradient explosion, 2018
Yang, Greg and Schoenholz, Sam S · 2018
Closest in time.