Fetching the paper…
Reading the bibliography…
These are the notes for the lectures that I was giving during Fall 2020 at the Moscow Institute of Physics and Technology (MIPT) and at the Yandex School of Data Analysis (YSDA).
Dynamics of deep neural networks and neural tangent hierarchy
Huang, J. and Yau, H.-T. (2019) · 1909
Earlier work this paper cites.
On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables
Isserlis, L. (1918) · 1918
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W. (1963) · 1963
Earlier work this paper cites.
The sizes of compact subsets of hilbert space and continuity of gaussian processes
Dudley, R. M. (1967) · 1967
Earlier work this paper cites.
Распределение собственных значений в некоторых ансамблях случайных матриц
Marchenko, V. A. and Pastur, L. A. (1967) · 1967
Earlier work this paper cites.
О равномерной сходимости частот появления событий к их вероятностям
Vapnik, V. N. and Chervonenkis, A. Y. (1971) · 1971
Earlier work this paper cites.
On the density of families of sets
Sauer, N. (1972) · 1972
Earlier work this paper cites.
Large deviations for stationary gaussian processes
Donsker, M. and Varadhan, S. (1985) · 1985
Earlier work this paper cites.
Multiplication of certain non-commuting random variables
Voiculescu, D. (1987) · 1987
Earlier work this paper cites.
On the method of bounded differences
McDiarmid, C. (1989) · 1989
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Yu, X.-H. and Chen, G.-A. (1995) · 1995
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Tropp, J. A. (2011) · 2011
Earlier work this paper cites.
Topics in random matrix theory
Tao, T. (2012) · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2013) · 2013
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2015) · 2015
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B. (2016) · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Pennington, J., Schoenholz, S., and Ganguli, S. (2017) · 2017
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. (2018) · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Later among the works it cites.
Deep linear networks with arbitrary loss: All local minima are global
Laurent, T. and Brecht, J. (2018) · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S. (2016) · 2016
Cited alongside, same era.
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J. (2016) · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2016) · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J. (2017) · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M. (2017) · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I. (2017) · 2017
Cited alongside, same era.
Depth creates no bad local minima
Lu, H. and Kawaguchi, K. (2017) · 2017
Cited alongside, same era.
Neyshabur, B., Bhojanapalli, S., and Srebro, N. (2018) · 2018
Later among the works it cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A. (2019) · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2019) · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2019) · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z. (2019) · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Nguyen, Q. (2019) · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P. (2019) · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
Dyer, E. and Gur-Ari, G. (2020) · 2020
Closest in time.