Fetching the paper…
Reading the bibliography…
We present a Statistical Mechanics (SM) model of deep neural networks, connecting the energy-based and the feed forward networks (FFN) approach.
Modern Theory of Critical Phenomena
Ma, S.-K · 1976
Earlier work this paper cites.
Critical point behaviour and probability theory
Cassandro, M. and Jona-Lasinio, G · 1978
Earlier work this paper cites.
Spin Glass Theory and Beyond , volume 9 of Lecture notes in Physics
Mézard, M., Parisi, G., and Virasoro, M. A · 1987
Earlier work this paper cites.
Modeling Brain Functions, The world of acctractor Neural Networks
Amit, D. J · 1989
Earlier work this paper cites.
Introduction to the Theory of Neural Computation
Hertz, J., Krogh, A., and Palmer, R. G · 1991
Earlier work this paper cites.
Statistical mechanics and applications in condensed matter
Di-Castro, C. and Raimondi, R · 2003
Earlier work this paper cites.
Probability Theory, the logic of science
Jaynes, E. T · 2003
Earlier work this paper cites.
Information Theory, Inference and Learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Equilibrium and Non-Equilibrium Statistical Thermodynamics
Bellac, M. L., Mortessagne, F., and Bastrouni, G. G · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C · 2006
Earlier work this paper cites.
Information, Physics and Computation
Mézard, M. and Montanari, A · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio., Y · 2014
Cited alongside, same era.
Expected energy-based restricted boltzmann machine for classification
Elfwing, S., Uchibe, E., and Doya, K · 2014
Cited alongside, same era.
An exact mapping between the variational renormalization group and deep learning
Mehta, P. and Schwab, D. J · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
The loss surface of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Michael, A., Gerard, B., and LeCun, Y · 2015
Cited alongside, same era.
Inverse statistical problems: from the inverse ising problem to data science
Nguyen, H. C., Zecchina, R., and Berg, J · 2017
Later among the works it cites.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Later among the works it cites.
Nonlinear random matrix theory for deep learning
Pennington, J. and Worah, P · 2017
Later among the works it cites.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: A review
Pogio, T., Mhaskar, H., Rosasco, L., Miranda, B., and Liao., Q · 2017
Later among the works it cites.
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Later among the works it cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tishby, N. and Zaslavsky, N · 2015
Cited alongside, same era.
Random Fields and Spin Glasees, A Field Theory approach
Dominicis, C. D. and Giardina, I · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S · 2016
Cited alongside, same era.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Kleinberg, J., Ganguli, S., and Sohl-Dickstein., J · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, J · 2017
Cited alongside, same era.
Why does deep and cheap learning work so well?
Lin, H. W., Tegmark, M., and Rolnick, D · 2017
Cited alongside, same era.
Sagun, L., Bottou, L., and LeCun., Y · 2017
Later among the works it cites.
Opening the black box of deep neural networks via information
Schwartz-Ziv, R. and Tishby, N · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Later among the works it cites.
On the selection of initialization and activation function for deep neural networks
Hayou, S., Doucet, A., and Rousseau, J · 2018
Closest in time.
Distribution regression networks
Kou, C., Lee, H. K., and Ng, T. K · 2018
Closest in time.
Artificial intelligence and its limits
Mézard, M · 2018
Closest in time.