Fetching the paper…
Reading the bibliography…
Despite the practical success of deep neural networks, a comprehensive theoretical framework that can predict practically relevant scores, such as the test accuracy, from knowledge of the training data is currently lacking.
1908
Earlier work this paper cites.
V. A. Marčenko and L. A. Pastur, Distribution of eigenvalues for some sets of random matrices, Mathematics of the USSR-Sbornik 1
1967
Earlier work this paper cites.
P. Breuer and P. Major, Central limit theorems for non-linear functionals of gaussian fields, Journal of Multivariate Analysis 13
1983
Earlier work this paper cites.
C. Cortes and V. Vapnik, Support-vector networks, Machine Learning 20
1995
Earlier work this paper cites.
R. M. Neal, Priors for infinite networks, in Bayesian Learning for Neural Networks (Springer New York, New York, NY, 1996) pp. 29–53
1996
Earlier work this paper cites.
C. Williams, Computing with infinite networks, in Advances in Neural Information Processing Systems , Vol. 9, edited by M. Mozer, M. Jordan, and T. Petsche (MIT Press, 1996)
1996
Earlier work this paper cites.
R. Dietrich, M. Opper, and H. Sompolinsky, Statistical mechanics of support vector networks, Phys. Rev. Lett. 82
1999
Earlier work this paper cites.
A. Engel and C. Van den Broeck, Statistical Mechanics of Learning (Cambridge University Press, 2001)
2001
Earlier work this paper cites.
2006
Earlier work this paper cites.
Y. Cho and L. Saul, Kernel methods for deep learning, in Advances in Neural Information Processing Systems , Vol. 22, edited by Y. Bengio, D. Schuurmans, J. Lafferty, C. Williams, and A. Culotta (Curran Associates, Inc., 2009)
2009
Earlier work this paper cites.
I. Nourdin, G. Peccati, and M. Podolskij, Quantitative Breuer-Major theorems (2010)
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Y. Bengio and O. Delalleau, On the expressive power of deep architectures, in International conference on algorithmic learning theory (Springer, 2011) pp. 18–36
2011
Earlier work this paper cites.
J.-M. Bardet and D. Surgailis, Moment bounds and central limit theorems for gaussian subordinated arrays, Journal of Multivariate Analysis 114
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Shah, A. Wilson, and Z. Ghahramani, Student-t Processes as Alternatives to Gaussian Processes, in Proceedings of the Seventeenth International Conference on Artificial Intelligence and Statistics , Proceedings of Machine Learning Research, Vol. 33, edited by S. Kaski and J. Corander (PMLR, Reykjavik, Iceland, 2014) pp. 877–885
2014
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning (MIT Press, 2016)
2016
Earlier work this paper cites.
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli, Exponential expressivity in deep neural networks through transient chaos, in Advances in Neural Information Processing Systems , Vol. 29, edited by D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Curran Associates, Inc., 2016)
2016
Earlier work this paper cites.
G. Yang and S. Schoenholz, Mean field residual networks: On the edge of chaos, in Advances in Neural Information Processing Systems , Vol. 30, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Curran Associates, Inc., 2017)
2017
Earlier work this paper cites.
B. Li and D. Saad, Exploring the function space of deep-learning machines, Phys. Rev. Lett. 120
2018
Earlier work this paper cites.
A. G. de G. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani, Gaussian process behaviour in wide deep neural networks, in International Conference on Learning Representations (2018)
2018
Earlier work this paper cites.
J. Lee, J. Sohl-dickstein, J. Pennington, R. Novak, S. Schoenholz, and Y. Bahri, Deep neural networks as gaussian processes, in International Conference on Learning Representations (2018)
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in Advances in Neural Information Processing Systems , Vol. 31, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018)
2018
Earlier work this paper cites.
E. Dobriban and S. Wager, High-dimensional asymptotics of prediction: ridge regression and classification, The Annals of Statistics 46
2018
Earlier work this paper cites.
B. D. Tracey and D. Wolpert, Upgrading from gaussian processes to student’st processes, in 2018 AIAA Non-Deterministic Approaches Conference (2018) p. 1659
2018
Cited alongside, same era.
A. Garriga-Alonso, C. E. Rasmussen, and L. Aitchison, Deep convolutional networks as shallow gaussian processes, in International Conference on Learning Representations (2019)
2019
Cited alongside, same era.
R. Novak, L. Xiao, Y. Bahri, J. Lee, G. Yang, D. A. Abolafia, J. Pennington, and J. Sohl-dickstein, Bayesian deep convolutional networks with many channels are gaussian processes, in International Conference on Learning Representations (2019)
2019
Cited alongside, same era.
L. Chizat, E. Oyallon, and F. Bach, On lazy training in differentiable programming, in Advances in Neural Information Processing Systems , Vol. 32, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., 2019)
2019
Cited alongside, same era.
G. Naveh and Z. Ringel, A self consistent theory of gaussian processes captures feature learning effects in finite cnns, in Advances in Neural Information Processing Systems , Vol. 34, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021) pp. 21352–21364
2021
Later among the works it cites.
B. Loureiro, C. Gerbelot, H. Cui, S. Goldt, F. Krzakala, M. Mezard, and L. Zdeborová, Learning curves of generic features maps for realistic datasets with a teacher-student model, Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari, Linearized two-layers neural networks in high dimension, The Annals of Statistics 49
2021
Later among the works it cites.
J. Zavatone-Veth, A. Canatar, B. Ruben, and C. Pehlevan, Asymptotics of representation learning in finite bayesian neural networks, in Advances in Neural Information Processing Systems , Vol. 34, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021) pp. 24765–24777
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington, Wide neural networks of any depth evolve as linear models under gradient descent, in Advances in Neural Information Processing Systems , Vol. 32, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., 2019)
2019
Cited alongside, same era.
P. L. Bartlett, N. Harvey, C. Liaw, and A. Mehrabian, Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks, The Journal of Machine Learning Research 20
2019
Cited alongside, same era.
A. M. Saxe, J. L. McClelland, and S. Ganguli, A mathematical theory of semantic development in deep neural networks, Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
S. Mei and A. Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve, Communications on Pure and Applied Mathematics (2019)
2019
Cited alongside, same era.
G. Pang, L. Yang, and G. E. Karniadakis, Neural-net-induced gaussian process regression for function approximation and pde solution, Journal of Computational Physics 384
2019
Cited alongside, same era.
A. Mozeika, B. Li, and D. Saad, Space of functions computed by deep-layered machines, Phys. Rev. Lett. 125
2020
Cited alongside, same era.
Y. Bahri, J. Kadmon, J. Pennington, S. S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli, Statistical mechanics of deep learning, Annual Review of Condensed Matter Physics 11
2020
Cited alongside, same era.
B. Bordelon, A. Canatar, and C. Pehlevan, Spectrum dependent learning curves in kernel regression and wide neural networks, in Proceedings of the 37th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 119, edited by H. D. III and A. Singh (PMLR, 2020) pp. 1024–1034
2020
Cited alongside, same era.
2021
Later among the works it cites.
A. Mozeika, M. Sheikh, F. Aguirre-López, F. Antenucci, and A. C. C. Coolen, Exact results on high-dimensional linear regression via statistical physics, Phys. Rev. E 103
2021
Later among the works it cites.
Y. Uchiyama, H. Oka, and A. Nono, Student’s t-process regression on the space of probability density functions, Proceedings of the ISCIE International Symposium on Stochastic Systems Theory and its Applications 2021
2021
Later among the works it cites.
J. A. Zavatone-Veth and C. Pehlevan, Depth induces scale-averaging in overparameterized linear bayesian neural networks, in 2021 55th Asilomar Conference on Signals, Systems, and Computers (2021) pp. 600–607
2021
Later among the works it cites.
C. Baldassi, C. Lauditi, E. M. Malatesta, R. Pacelli, G. Perugini, and R. Zecchina, Learning through atypical phase transitions in overparameterized neural networks, Phys. Rev. E 106
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
F. Aguirre-López, M. Pastore, and S. Franz, Satisfiability transition in asymmetric neural networks, Journal of Physics A: Mathematical and Theoretical 55
2022
Closest in time.
S. Ariosto, R. Pacelli, F. Ginelli, M. Gherardi, and P. Rotondo, Universal mean-field upper bound for the generalization gap of deep neural networks, Phys. Rev. E 105
2022
Closest in time.
J. A. Zavatone-Veth, A. Canatar, B. S. Ruben, and C. Pehlevan, Asymptotics of representation learning in finite bayesian neural networks*, Journal of Statistical Mechanics: Theory and Experiment 2022
2022
Closest in time.
H. Lee, E. Yun, H. Yang, and J. Lee, Scale mixtures of neural network gaussian processes, in International Conference on Learning Representations (2022)
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
I. Seroussi, G. Naveh, and Z. Ringel, Separation of scales and a thermodynamic description of feature learning in some cnns, Nature Communications 14
2023
Closest in time.
A. J. Wakhloo, T. J. Sussman, and S. Chung, Linear classification of neural manifolds with correlated variability, Phys. Rev. Lett. 131
2023
Closest in time.
2023
Closest in time.
S. Ariosto, R. Pacelli, M. Pastore, F. Ginelli, M. Gherardi, and P. Rotondo, Supplemental Material for ”A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-wdith limit” (2023)
2023
Closest in time.
B. Hanin and A. Zlokapa, Bayesian interpolation with deep linear networks, Proceedings of the National Academy of Sciences 120
2023
Closest in time.
2023
Closest in time.
H. Cui, F. Krzakala, and L. Zdeborov’a, Optimal learning of deep random networks of extensive-width, in International Conference on Machine Learning (2023)
2023
Closest in time.
R. Pacelli, rpacelli/FC_deep_bayesian_networks: FC_deep_bayesian_networks (2023)
2023
Closest in time.