Fetching the paper…
Reading the bibliography…
Conventional wisdom in deep learning states that increasing depth improves expressiveness but complicates optimization.
Elementary differential equations and boundary value problems , volume 9
Boyce, W. E., DiPrima, R. C., and Haines, C. W · 1969
Earlier work this paper cites.
Updating quasi-newton matrices with limited storage
Nocedal, J · 1980
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Nesterov, Y · 1983
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
Effect of batch learning in multilayer neural networks
Fukumizu, K · 1998
Earlier work this paper cites.
SciPy: Open source scientific tools for Python, 2001–
Jones, E., Oliphant, T., Peterson, P., et al · 2001
Earlier work this paper cites.
Advanced calculus
Buck, R. C · 2003
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Optimization and dynamical systems
Helmke, U. and Moore, J. B · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Chemical gas sensor drift compensation using classifier ensembles
Vergara, A., Vembu, S., Ayhan, T., Ryan, M. A., Homer, M. L., and Huerta, R · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Zeiler, M. D · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Cited alongside, same era.
On the calibration of sensor arrays for pattern recognition using the minimal number of experiments
Rodriguez-Lujan, I., Fonollosa, J., Vergara, A., Homer, M., and Huerta, R · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Su, W., Boyd, S., and Candes, E · 2014
Cited alongside, same era.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Later among the works it cites.
Identity matters in deep learning
Hardt, M. and Ma, T · 2016
Later among the works it cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Later among the works it cites.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Kleinberg, J., Ganguli, S., and Sohl-Dickstein, J · 2016
Later among the works it cites.
On the quality of the initial basin in overspecified neural networks
Safran, I. and Shamir, O · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Cited alongside, same era.
The power of depth for feedforward neural networks
Eldan, R. and Shamir, O · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Cited alongside, same era.
Global Optimality in Tensor Factorization, Deep Learning, and Beyond
Haeffele, B. D. and Vidal, R · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
Janzamin, M., Sedghi, H., and Anandkumar, A · 2015
Cited alongside, same era.
Soudry, D. and Carmon, Y · 2016
Later among the works it cites.
A variational perspective on accelerated methods in optimization
Wibisono, A., Wilson, A. C., and Jordan, M. I · 2016
Later among the works it cites.
Finding approximate local minima faster than gradient descent
Agarwal, N., Allen-Zhu, Z., Bullins, B., Hazan, E., and Ma, T · 2017
Later among the works it cites.
Analysis and design of convolutional networks via hierarchical tensor decompositions
Cohen, N., Sharir, O., Levine, Y., Tamari, R., Yakira, D., and Shashua, A · 2017
Later among the works it cites.
Depth separation for neural networks
Daniely, A · 2017
Later among the works it cites.
Why momentum really works
Goh, G · 2017
Later among the works it cites.
On the ability of neural nets to express distributions
Lee, H., Ge, R., Risteski, A., Ma, T., and Arora, S · 2017
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2017
Later among the works it cites.
Understanding deep neural networks with rectified linear units
Arora, R., Basu, A., Mianjy, P., and Mukherjee, A · 2018
Closest in time.