Fetching the paper…
Reading the bibliography…
Normalization techniques such as Batch Normalization have been applied successfully for training deep neural networks.
Maximum-entropy distributions having prescribed first and second moments (corresp.)
Dowson, D. and Wragg, A. (1973) · 1973
Earlier work this paper cites.
Minimization methods for smooth nonconvex functions
Mikhalevich, V., Redkovskii, N., and Antonyuk, A. (1988) · 1988
Earlier work this paper cites.
On exact estimates of the convergence rate of the steepest ascent method in the symmetric eigenvalue problem
Knyazev, A. and Shorokhodov, A. (1991) · 1991
Earlier work this paper cites.
Optimization in solving elliptic problems
D’yakonov, E. G. and McCormick, S. (1995) · 1995
Earlier work this paper cites.
Preconditioned eigensolvers—an oxymoron
Knyazev, A. V. (1998) · 1998
Earlier work this paper cites.
A geometric theory for preconditioned inverse iteration iii: A short and sharp convergence estimate for generalized eigenvalue problems
Knyazev, A. V. and Neymeyr, K. (2003) · 2003
Earlier work this paper cites.
Stein’s lemma for elliptical random vectors
Landsman, Z. and Ne v · 2008
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
Absil, P.-A., Mahony, R., and Sepulchre, R. (2009) · 2009
Earlier work this paper cites.
Hardness of learning halfspaces with noise
Guruswami, V. and Raghavendra, P. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
A generalized linear model with “gaussian” regressor variables
Brillinger, D. R. (2012) · 2012
Earlier work this paper cites.
Estimating the hessian by back-propagating curvature
Martens, J., Sutskever, I., and Swersky, K. (2012) · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Nesterov, Y. (2013) · 2013
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Learning halfspaces and neural networks with random initialization
Zhang, Y., Lee, J. D., Wainwright, M. J., and Jordan, M. I. (2015) · 2015
Cited alongside, same era.
Arpit, D., Zhou, Y., Kota, B. U., and Govindaraju, V. (2016) · 2016
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Cited alongside, same era.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
Du, S. S., Lee, J. D., Tian, Y., Poczos, B., and Singh, A. (2017) · 2017
Later among the works it cites.
Gitman, I. and Ginsburg, B. (2017) · 2017
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
Jin, C., Netrapalli, P., and Jordan, M. I. (2017) · 2017
Later among the works it cites.
Self-normalizing neural networks
Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S. (2017) · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaled least squares estimator for glms in large-scale problems
Erdogdu, M. A., Dicker, L. H., and Bayati, M. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P. (2016) · 2016
Cited alongside, same era.
Convergence theory for preconditioned eigenvalue solvers in a nutshell
Argentati, M. E., Knyazev, A. V., Neymeyr, K., Ovtchinnikov, E. E., and Zhou, M. (2017) · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A. and Globerson, A. (2017) · 2017
Cited alongside, same era.
Riemannian approach to batch normalization
Cho, M. and Lee, J. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D. (2017) · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A. (2017) · 2017
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. S. and Lee, J. D. (2018) · 2018
Closest in time.
Troubling trends in machine learning scholarship
Lipton, Z. C. and Steinhardt, J. (2018) · 2018
Closest in time.
How does batch normalization help optimization?(no, it is not about internal covariate shift)
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A. (2018) · 2018
Closest in time.