Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Original
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Original
Yuan Cao and Quanquan Gu · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Original
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
Original
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1904
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Original
Yuan Cao and Quanquan Gu · 1905
Earlier work this paper cites.
Limitations of lazy training of two-layers neural networks
Original
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1906
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
Mercer’s theorem, feature maps, and smoothing
Ha Quang Minh, Partha Niyogi, and Yuan Yao · 2006
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Original
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Original
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Identity matters in deep learning
Original
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.