Fetching the paper…
Reading the bibliography…
The success of deep learning in many real-world tasks has triggered an intense effort to understand the power and limitations of deep learning in the training and generalization of complex tasks, so far with limited progress.
J. Pearl, Fusion, propagation, and structuring in belief networks, Artificial intelligence 29
1986
Earlier work this paper cites.
D. J. Amit, H. Gutfreund, and H. Sompolinsky, Statistical mechanics of neural networks near saturation, Annals of physics 173
1987
Earlier work this paper cites.
E. Gardner, The space of interactions in neural network models, Journal of physics A: Mathematical and general 21
1988
Earlier work this paper cites.
E. Gardner and B. Derrida, Optimal storage properties of neural network models, Journal of Physics A: Mathematical and general 21
1988
Earlier work this paper cites.
N. Tishby, E. Levin, and S. A. Solla, Consistent inference of probabilities in layered networks: Predictions and generalization, in International Joint Conference on Neural Networks , Vol. 2 (1989) pp. 403–409
1989
Earlier work this paper cites.
D. J. MacKay, A practical bayesian framework for backpropagation networks, Neural computation 4
1992
Earlier work this paper cites.
H. S. Seung, H. Sompolinsky, and N. Tishby, Statistical mechanics of learning from examples, Physical review A 45
1992
Earlier work this paper cites.
S. Geman, E. Bienenstock, and R. Doursat, Neural networks and the bias/variance dilemma, Neural computation 4
1992
Earlier work this paper cites.
S. W. Pierson and O. T. Valls, Renormalization-group study of a layered-superconductor model, Physical Review B 45
1992
Earlier work this paper cites.
T. L. Watkin, A. Rau, and M. Biehl, The statistical mechanics of learning a rule, Reviews of Modern Physics 65
1993
Earlier work this paper cites.
S. W. Pierson, Critical behavior of vortices in a layered system, Physical review letters 73
1994
Earlier work this paper cites.
R. Dodier, Geometry of early stopping in linear networks, Advances in neural information processing systems , 365 (1996)
1996
Earlier work this paper cites.
L.-Y. Chen, N. Goldenfeld, and Y. Oono, Renormalization group and singular perturbations: Multiple scales, boundary layers, and reductive perturbation theory, Physical Review E 54
1996
Earlier work this paper cites.
Y. LeCun, P. Haffner, L. Bottou, and Y. Bengio, Object recognition with gradient-based learning, in Shape, contour and grouping in computer vision (Springer, 1999) pp. 319–345
1999
Earlier work this paper cites.
P. Domingos, A unified bias-variance decomposition for zero-one and squared loss, AAAI/IAAI 2000
2000
Earlier work this paper cites.
J. S. Yedidia, W. T. Freeman, Y. Weiss, et al. , Generalized belief propagation, in NIPS , Vol. 13 (2000) pp. 689–695
2000
Earlier work this paper cites.
A. Engel and C. Van den Broeck, Statistical mechanics of learning (Cambridge University Press, 2001)
2001
Earlier work this paper cites.
A. Messinger, L. R. Squire, S. M. Zola, and T. D. Albright, Neuronal representations of stimulus associations develop in the temporal lobe during learning, Proceedings of the National Academy of Sciences 98
2001
Earlier work this paper cites.
F. R. Kschischang, B. J. Frey, and H.-A. Loeliger, Factor graphs and the sum-product algorithm, IEEE Transactions on information theory 47
2001
Earlier work this paper cites.
J. S. Yedidia, W. T. Freeman, Y. Weiss, et al. , Understanding belief propagation and its generalizations, Exploring artificial intelligence in the new millennium 8
2003
Earlier work this paper cites.
J. Shawe-Taylor, N. Cristianini, et al. , Kernel methods for pattern analysis (Cambridge university press, 2004)
2004
Earlier work this paper cites.
J. S. Yedidia, W. T. Freeman, and Y. Weiss, Constructing free-energy approximations and generalized belief propagation algorithms, IEEE Transactions on information theory 51
2005
Earlier work this paper cites.
J. Winn, C. M. Bishop, and T. Jaakkola, Variational message passing., Journal of Machine Learning Research 6
2005
Earlier work this paper cites.
T. Hofmann, B. Schölkopf, and A. J. Smola, Kernel methods in machine learning, The annals of statistics , 1171 (2008)
2008
Earlier work this paper cites.
A. Rahimi and B. Recht, Random features for large-scale kernel machines, in Advances in neural information processing systems (2008) pp. 1177–1184
2008
Earlier work this paper cites.
N. Kriegeskorte, M. Mur, and P. A. Bandettini, Representational similarity analysis-connecting the branches of systems neuroscience, Frontiers in systems neuroscience 2
2008
Cited alongside, same era.
M. Mezard and A. Montanari, Information, physics, and computation (Oxford University Press, 2009)
2009
Cited alongside, same era.
Y. Cho and L. K. Saul, Kernel methods for deep learning, in Advances in neural information processing systems (2009) pp. 342–350
2009
Cited alongside, same era.
S. Ganguli and H. Sompolinsky, Statistical mechanics of compressed sensing, Physical review letters 104
2010
Cited alongside, same era.
Y. Weiss and J. Pearl, Belief propagation: technical perspective, Communications of the ACM 53
2010
Cited alongside, same era.
H. Lu and K. Kawaguchi, Depth creates no bad local minima, arXiv preprint arXiv:1702.08580 (2017)
2017
Later among the works it cites.
M. Van der Wilk, C. E. Rasmussen, and J. Hensman, Convolutional gaussian processes, in Advances in Neural Information Processing Systems (2017) pp. 2849–2858
2017
Later among the works it cites.
A. Banino, C. Barry, B. Uria, C. Blundell, T. Lillicrap, P. Mirowski, A. Pritzel, M. J. Chadwick, T. Degris, J. Modayil, et al. , Vector-based navigation using grid-like representations in artificial agents, Nature 557
2018
Later among the works it cites.
M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G. B. Arous, C. Cammarota, Y. LeCun, M. Wyart, and G. Biroli, Comparing dynamics: Deep neural networks versus glassy systems, in International Conference on Machine Learning (2018) pp. 314–323
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Bai and J. W. Silverstein, Sample covariance matrices and the marčenko-pastur law, in Spectral Analysis of Large Dimensional Random Matrices (Springer, 2010) pp. 39–58
2010
Cited alongside, same era.
S. Rifai, P. Vincent, X. Muller, X. Glorot, and Y. Bengio, Contractive auto-encoders: Explicit invariance during feature extraction, in Icml (2011)
2011
Cited alongside, same era.
A. Graves, Practical variational inference for neural networks, in Advances in neural information processing systems (Citeseer, 2011) pp. 2348–2356
2011
Cited alongside, same era.
G.-X. Yuan, C.-H. Ho, and C.-J. Lin, Recent advances of large-scale linear classification, Proceedings of the IEEE 100
2012
Cited alongside, same era.
R. M. Neal, Bayesian learning for neural networks , Vol. 118 (Springer Science & Business Media, 2012)
2012
Cited alongside, same era.
S. Ganguli and H. Sompolinsky, Compressed sensing, sparsity, and dimensionality in neuronal information processing and data analysis, Annual review of neuroscience 35
2012
Cited alongside, same era.
L. Deng, G. Hinton, and B. Kingsbury, New types of deep neural network learning for speech recognition and related applications: An overview, in 2013 IEEE international conference on acoustics, speech and signal processing (IEEE, 2013) pp. 8599–8603
2013
Cited alongside, same era.
S. Chung, D. D. Lee, and H. Sompolinsky, Classification and geometry of general perceptual manifolds, Physical Review X 8
2018
Later among the works it cites.
T. Laurent and J. Brecht, Deep linear networks with arbitrary loss: All local minima are global, in International conference on machine learning (PMLR, 2018) pp. 2902–2907
2018
Later among the works it cites.
N. Goldenfeld, Lectures on phase transitions and the renormalization group (CRC Press, 2018)
2018
Later among the works it cites.
S.-H. Li and L. Wang, Neural network renormalization group, Physical review letters 121
2018
Later among the works it cites.
A. Garriga-Alonso, C. E. Rasmussen, and L. Aitchison, Deep convolutional networks as shallow gaussian processes, in International Conference on Learning Representations (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Machine learning and the physical sciences, Reviews of Modern Physics 91
2019
Later among the works it cites.
A. M. Saxe, J. L. McClelland, and S. Ganguli, A mathematical theory of semantic development in deep neural networks, Proceedings of the National Academy of Sciences 116
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Belkin, D. Hsu, S. Ma, and S. Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off, Proceedings of the National Academy of Sciences 116
2019
Later among the works it cites.
2019
Later among the works it cites.
T. Parr, D. Markovic, S. J. Kiebel, and K. J. Friston, Neuronal message passing using mean-field, bethe, and marginal approximations, Scientific reports 9
2019
Later among the works it cites.
T. Poggio, A. Banburski, and Q. Liao, Theoretical issues in deep networks, Proceedings of the National Academy of Sciences (2020)
2020
Closest in time.
S. Becker, Y. Zhang, et al. , Geometry of energy landscapes and the optimizability of deep neural networks, Physical Review Letters 124
2020
Closest in time.
Y. Bahri, J. Kadmon, J. Pennington, S. S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli, Statistical mechanics of deep learning, Annual Review of Condensed Matter Physics (2020)
2020
Closest in time.
R. Vershynin, Memory capacity of neural networks with threshold and rectified linear unit activations, SIAM Journal on Mathematics of Data Science 2
2020
Closest in time.
M. Li, M. Soltanolkotabi, and S. Oymak, Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks, in International Conference on Artificial Intelligence and Statistics (PMLR, 2020) pp. 4313–4324
2020
Closest in time.
M. S. Advani, A. M. Saxe, and H. Sompolinsky, High-dimensional dynamics of generalization error in neural networks, Neural Networks 132
2020
Closest in time.
2020
Closest in time.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, Understanding deep learning (still) requires rethinking generalization, Communications of the ACM 64
2021
Closest in time.