Fetching the paper…
Reading the bibliography…
We prove that the global minimum of the backpropagation (BP) training problem of neural networks with an arbitrary nonlinear activation is given by the ridgelet transform.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R. (1993) · 1993
Earlier work this paper cites.
An integral representation of functions using three-layered betworks and their approximation bounds
Murata, N. (1996) · 1996
Earlier work this paper cites.
Ridgelets: theory and applications
Candès, E. J. (1998) · 1998
Earlier work this paper cites.
The finite ridgelet transform for image representation
Do, M. N. and Vetterli, M. (2003) · 2003
Earlier work this paper cites.
Universality, Characteristic Kernels and RKHS Embedding of Measures
Sriperumbudur, B. K., Fukumizu, K., and Lanckriet, G. R. G. (2010) · 2010
Earlier work this paper cites.
The ridgelet and curvelet transforms
Starck, J.-L., Murtagh, F., and Fadili, J. M. (2010) · 2010
Earlier work this paper cites.
Complexity estimates based on integral transforms induced by computational units
Kůrková, V. (2012) · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Sampling hidden parameters from oracle distribution
Sonoda, S. and Murata, N. (2014) · 2014
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y. (2015) · 2015
Earlier work this paper cites.
Deep Learning without Poor Local Minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
You Only Look Once: Unified, Real-Time Object Detection
Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016) · 2016
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
Brutzkus, A. and Globerson, A. (2017) · 2017
Cited alongside, same era.
Identity Matters in Deep Learning
Hardt, M. and Ma, T. (2017) · 2017
Cited alongside, same era.
Convergence Analysis of Two-layer Neural Networks with ReLU Activation
Li, Y. and Yuan, Y. (2017) · 2017
Cited alongside, same era.
Neural network with unbounded activation functions is universal approximator
Sonoda, S. and Murata, N. (2017) · 2017
Later among the works it cites.
Recovery Guarantees for One-hidden-layer Neural Networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S. (2017) · 2017
Later among the works it cites.
Essentially No Barriers in Neural Network Energy Landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A. (2018) · 2018
Closest in time.
On the Power of Over-parametrization in Neural Networks with Quadratic Activation
Du, S. S. and Lee, J. D. (2018) · 2018
Closest in time.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D., and Wilson, A. G. (2018) · 2018
Closest in time.
Learning One-hidden-layer Neural Networks with Landscape Design
Ge, R., Lee, J. D., and Ma, T. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Loss Surface of Deep and Wide Neural Networks
Nguyen, Q. and Hein, M. (2017) · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017) · 2017
Cited alongside, same era.
Learning ReLUs via Gradient Descent
Soltanolkotabi, M. (2017) · 2017
Cited alongside, same era.
Breaking the Curse of Dimensionality with Convex Neural Networks
Bach, F. (2017a)
Cited in the paper.
On the Equivalence between Kernel Quadrature Rules and Random Feature Expansions
Bach, F. (2017b)
Cited in the paper.
Closest in time.
Transport Analysis of Infinitely Deep Neural Network
Sonoda, S. and Murata, N. (2018) · 2018
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E. (2018) · 2018
Closest in time.
Fast generalization error bound of deep learning from a kernel perspective
Suzuki, T. (2018) · 2018
Closest in time.