Fetching the paper…
Reading the bibliography…
We analyze the loss landscape and expressiveness of practical deep convolutional neural networks (CNNs) with shared weights and max pooling layers.
Neural networks and principle component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1988
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Blum, A. and Rivest., R. L · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
LeCun, Y., Boser, B., Denker, J., Henderson, D., Howard, R., Hubbard, W., and Jackel, L · 1990
Earlier work this paper cites.
Exponentially many local minima for single neurons
Auer, P., Herbster, M., and Warmuth, M. K · 1996
Earlier work this paper cites.
A Primer of Real Analytic Functions
Krantz, S. G. and Parks, H. R · 2002
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
Sima, J · 2002
Earlier work this paper cites.
Numerical Recipes: The art of scientific computing
Press, W. H · 2007
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, K., Kavukcuoglu, K., and LeCun, Y · 2009
Earlier work this paper cites.
Shallow vs. deep sum-product networks
Delalleau, O. and Bengio, Y · 2011
Earlier work this paper cites.
On random weights and unsupervised feature learning
Saxe, A., Koh, P. W., Chen, Z., Bhand, M., Suresh, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Learning polynomials with neural networks
Andoni, A., Panigrahy, R., Valiant, G., and Zhang, L · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Montufar, G., Pascanu, R., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
On the number of response regions of deep feedforward networks with piecewise linear activations
Pascanu, R., Montufar, G., and Bengio, Y · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
How can deep rectifier networks achieve linear separability and preserve distances?
An, S., Boussaid, F., and Bennamoun, M · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Cited alongside, same era.
Understanding deep image representations by inverting them
Mahendran, A. and Vedaldi, A · 2015
Cited alongside, same era.
The zero set of a real analytic function, 2015
Mityagin, B · 2015
Cited alongside, same era.
Complex powers of analytic functions and meromorphic renormalization in qft, 2015
Nguyen, V. D · 2015
Cited alongside, same era.
Provable methods for training neural networks with sparse connectivity
Sedghi, H. and Anandkumar, A · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Benefits of depth in neural networks
Telgarsky, M · 2016
Later among the works it cites.
Error bounds for approximations with deep relu networks, 2016
Yarotsky, D · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A. and Globerson, A · 2017
Closest in time.
Understanding synthetic gradients and decoupled neural interfaces
Czarnecki, W. M., Swirszcz, G., Jaderberg, M., Osindero, S., Vinyals, O., and Kavukcuoglu, K · 2017
Closest in time.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2017
Closest in time.
Learning depth-three neural networks in polynomial time, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
Representation benefits of deep feedforward networks, 2015
Telgarsky, M · 2015
Cited alongside, same era.
Understanding neural networks through deep visualization
Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H · 2015
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2016
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions, 2016
Chollet, F · 2016
Cited alongside, same era.
Convolutional rectifier networks as generalized tensor decompositions
Cohen, N. and Shashua, A · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
Eldan, R. and Shamir, O · 2016
Cited alongside, same era.
Goel, S. and Klivans, A · 2017
Closest in time.
Global optimality in neural network training
Haeffele, B. D. and Vidal, R · 2017
Closest in time.
Identity matters in deep learning
Hardt, M. and Ma, T · 2017
Closest in time.
Convergence analysis of two-layer neural networks with relu activation, 2017
Li, Y. and Yuan, Y · 2017
Closest in time.
Why deep neural networks for function approximation?
Liang, S. and Srikant, R · 2017
Closest in time.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Closest in time.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Kleinberg, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Closest in time.
Depth-width tradeoffs in approximating natural functions with neural networks
Safran, I. and Shamir, O · 2017
Closest in time.
Failures of gradient-based deep learning
Shalev-Shwartz, S., Shamir, O., and Shammah, S · 2017
Closest in time.
Distribution-specific hardness of learning neural networks, 2017
Shamir, O · 2017
Closest in time.
Learning relus via gradient descent
Soltanolkotabi, M · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks, 2017
Soudry, D. and Hoffer, E · 2017
Closest in time.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Tian, Y · 2017
Closest in time.
Understanding deep learning requires re-thinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Closest in time.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P., and Dhillon, I · 2017
Closest in time.
When is a convolutional filter easy to learn?
Du, S. S., Lee, J. D., and Tian, Y · 2018
Closest in time.
Global optimality conditions for deep neural networks
Yun, C., Sra, S., and Jadbabaie, A · 2018
Closest in time.