Fetching the paper…
Reading the bibliography…
We study the connection between the highly non-convex loss function of a simple model of the fully-connected feed-forward neural network and the Hamiltonian of the spherical spin-glass model under the assumptions of: i) variable independence, ii) redundancy in network parametrization, and iii) uniformity.
On the Distribution of the Roots of Certain Symmetric Matrices
Wigner, E. P. (1958) · 1958
Earlier work this paper cites.
Spin-glass models of neural networks
Amit, D. J., Gutfreund, H., and Sompolinsky, H. (1985) · 1985
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
An Introduction to the Theory of Spin Glasses and Neural Networks
Dotsenko, V. (1995) · 1995
Earlier work this paper cites.
Exact solution for on-line learning in multilayer neural networks
Saad, D. and Solla, S. A. (1995) · 1995
Earlier work this paper cites.
Mean-field theory for a spin-glass model of neural networks: Tap free energy and the paramagnetic to spin-glass transition
Nakanishi, K. and Takayama, H. (1997) · 1997
Earlier work this paper cites.
Online algorithms and stochastic approximations
Bottou, L. (1998) · 1998
Earlier work this paper cites.
Decoupling : from dependence to independence : randomly stopped processes, U-statistics and processes, martingales and beyond
De la Peña, V. H. and Giné, E. (1999) · 1999
Earlier work this paper cites.
The Elements of Statistical Learning
Hastie, T., Tibshirani, R., and Friedman, J. (2001) · 2001
Cited alongside, same era.
The statistics of critical points of gaussian fields on large-dimensional spaces
Bray, A. J. and Dean, D. S. (2007) · 2007
Cited alongside, same era.
Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity
Fyodorov, Y. V. and Williams, I. (2007) · 2007
Cited alongside, same era.
On-line learning in neural networks
Saad, D. (2009) · 2009
Cited alongside, same era.
Random matrices and complexity of spin glasses
Auffinger, A., Ben Arous, G., and Cerny, J. (2010) · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. (2010) · 2010
Complexity of random smooth functions on the high-dimensional sphere
Auffinger, A. and Ben Arous, G. (2013) · 2013
Later among the works it cites.
Predicting parameters in deep learning
Denil, M., Shakibi, B., Dinh, L., Ranzato, M., and Freitas, N. D. (2013) · 2013
Later among the works it cites.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013) · 2013
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gülçehre, Ç., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Closest in time.
Exploiting linear structure within convolutional networks for efficient evaluation
Denton, E., Zaremba, W., Bruna, J., LeCun, Y., and Fergus, R. (2014) · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Yu, D., Dahl, G., rahman Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B. (2012) · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012) · 2012
Cited alongside, same era.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998a)
Cited in the paper.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G., and Muller, K. (1998b)
Cited in the paper.
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2014) · 2014
Closest in time.
#tagspace: Semantic embeddings from hashtags
Weston, J., Chopra, S., and Adams, K. (2014) · 2014
Closest in time.