Fetching the paper…
Reading the bibliography…
Neural networks with REctified Linear Unit (ReLU) activation functions (a.k.a.
F. Rosenblatt, “The perceptron: A probabilistic model for information storage and organization in the brain,” Psychol. Rev. , vol. 65, no. 6, p. 386, Nov. 1958
1958
Earlier work this paper cites.
A. B. Novikoff, “On convergence proofs for perceptrons,” in Proc. Symp. Math. Theory Automata , vol. 12, 1963, pp. 615–622
1963
Earlier work this paper cites.
N. Littlestone and M. Warmuth, “Relating data compression and learnability,” University of California, Santa Cruz, Tech. Rep., 1986
1986
Earlier work this paper cites.
A. Blum and R. L. Rivest, “Training a 3-node neural network is NP-complete,” in Adv. in Neural Inf. Process. Syst. , Cambridge, Massachusetts, Aug. 3–5, 1988, pp. 494–501
1988
Earlier work this paper cites.
F. H. Clarke, Optimization and Nonsmooth Analysis . SIAM, 1990, vol. 5
1990
Earlier work this paper cites.
M. Gori and A. Tesi, “On the problem of local minima in backpropagation,” IEEE Trans. Pattern Anal. Mach. Intell. , no. 1, pp. 76–86, Jan. 1992
1992
Earlier work this paper cites.
L. Holmstrom and P. Koistinen, “Using additive noise in back-propagation training,” IEEE Trans. Neural Netw. , vol. 3, no. 1, pp. 24–38, Jan. 1992
1992
Earlier work this paper cites.
P. Auer, M. Herbster, and M. K. Warmuth, “Exponentially many local minima for single neurons,” in Adv. in Neural Inf. Process. Syst. , Denver, Colorado, Nov. 27–Dec. 2, 1995, pp. 316–322
1995
Earlier work this paper cites.
G. An, “The effects of adding noise during backpropagation training on a generalization performance,” Neural Comput. , vol. 8, no. 3, pp. 643–674, Apr. 1996
1996
Earlier work this paper cites.
C. Wang and J. C. Principe, “Training neural networks with additive noise in the desired signal,” IEEE Trans. Neural Netw. , vol. 10, no. 6, pp. 1511–1517, Nov. 1999
1999
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Intl. Conf. on Artif. Intell. and Stat. , Sardinia, Italy, May 13–15, 2010, pp. 249–256
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted Boltzmann machines,” in Intl. Conf. on Mach. Learn. , Haifa, Israel, June 21–24, 2010, pp. 807–814
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Adv. in Neural Inf. Process. Syst. , Lake Tahoe, Nevada, Dec. 3–6, 2012, pp. 1097–1105
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms . New York, NY: Cambridge University Press, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
R. Ge, F. Huang, C. Jin, and Y. Yuan, “Escaping from saddle points — Online stochastic gradient for tensor decomposition,” in Conf. on Learn. Theory , vol. 40, Paris, France, July 3–6, 2015, pp. 797–842
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . Cambridge: MIT press, 2016, vol. 1
2016
Earlier work this paper cites.
K. Kawaguchi, “Deep learning without poor local minima,” in Adv. in Neural Inf. Process. Syst. , Barcelona, Spain, Dec. 5–10, 2016, pp. 586–594
2016
Earlier work this paper cites.
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli, “Exponential expressivity in deep neural networks through transient chaos,” in Adv. in Neural Inf. Process. Syst. , Barcelona, Spain, Dec. 5-10, 2016, pp. 3360–3368
2016
Cited alongside, same era.
A. Brutzkus and A. Globerson, “Globally optimal gradient descent for a ConvNet with Gaussian inputs,” in Intl. Conf. on Mach. Learn. , vol. 70, Sydney, Australia, Aug. 6–11, 2017
2017
Cited alongside, same era.
D. Dheeru and E. Karra Taniskidou, “UCI machine learning repository,” 2017. [Online]. Available: http://archive.ics.uci.edu/ml
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Closest in time.
Q. Nguyen and M. Hein, “Optimization landscape and expressivity of deep CNNs,” in Intl. Conf. Mach. Learn. , 2018, pp. 3727–3736
2018
Closest in time.
2018
Closest in time.
I. Safran and O. Shamir, “Spurious local minima are common in two-layer ReLU neural networks,” in Intl. Conf. on Mach. Learn. , vol. 80, Stockholm, Sweden, July 10–15, 2018, pp. 4430–4438
2018
Closest in time.
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro, “The implicit bias of gradient descent on separable data,” J. Mach. Learn. Res. , vol. 19, no. 70, pp. 1–57, 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, and J. Sohl-Dickstein, “On the expressive power of deep neural networks,” in Intl. Conf. on Mach. Learn. , vol. 70, Sydney, Australia, Aug. 6–11, 2017, pp. 2847–2854
2017
Cited alongside, same era.
M. Soltanolkotabi, “Learning ReLU via gradient descent,” in Adv. in Neural Inf. Process. Syst. , Long Beach, CA, Dec. 4–9, 2017, pp. 2007–2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Zhong, Z. Song, P. Jain, P. L. Bartlett, and I. S. Dhillon, “Recovery guarantees for one-hidden-layer neural networks,” in Intl. Conf. on Mach. Learn. , vol. 70, Sydney, Australia, Aug. 6–11, 2017, pp. 4140–4149
2017
Cited alongside, same era.
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz, “SGD learns over-parameterized networks that provably generalize on linearly separable data,” in Intl. Conf. on Learn. Rep. , Vancouver, BC, Canada, Apr. 30–May 3, 2018
2018
Cited alongside, same era.
S. S. Du, J. D. Lee, Y. Tian, B. Poczos, and A. Singh, “Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima,” in Intl. Conf. on Mach. Learn. , vol. 80, Stockholm, Sweden, July 10–15, 2018, pp. 1338–1347
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Closest in time.
G. Wang, G. B. Giannakis, and Y. C. Eldar, “Solving systems of random quadratic equations via truncated amplitude flow,” IEEE Trans. Inf. Theory , vol. 64, no. 2, pp. 773–794, Feb. 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2019
Closest in time.
2019
Closest in time.
Q. Nguyen, “On connected sublevel sets in deep learning,” arXiv:1901.07417 , 2019
2019
Closest in time.
2019
Closest in time.