Fetching the paper…
Reading the bibliography…
It is well known that the initialization of weights in deep neural networks can have a dramatic impact on learning speed.
Free random variables
Dan V Voiculescu, Ken J Dykema, and Alexandru Nica · 1992
Earlier work this paper cites.
Multiplicative functions on the lattice of non-crossing partitions and free convolution
Roland Speicher · 1994
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Topics in random matrix theory
Terence Tao · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Performance-optimized hierarchical models predict neural responses in higher visual cortex
Daniel LK Yamins, Ha Hong, Charles F Cadieu, Ethan A Solomon, Darren Seibert, and James J DiCarlo · 2014
Cited alongside, same era.
Plancherel–rotach formulae for average characteristic polynomials of products of ginibre random matrices and the fuss–catalan distribution
Thorsten Neuschel · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deep knowledge tracing
Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J Guibas, and Jascha Sohl-Dickstein · 2015
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Deep learning models of the retinal response to natural scenes
Lane McIntosh, Niru Maheswaranathan, Aran Nayebi, Surya Ganguli, and Stephen Baccus · 2016
Later among the works it cites.
Exponential expressivity in deep neural networks through transient chaos
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Nouvelle méthode pour résoudre les problèmes indéterminés en nombres entiers
Joseph Louis Lagrange
Cited in the paper.
Neural networks for machine learning lecture 6a overview of mini–batch gradient descent
Geoffrey Hinton, NiRsh Srivastava, and Kevin Swersky
Cited in the paper.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le · 2016
Later among the works it cites.
Deep Information Propagation
S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein · 2017
Closest in time.