Fetching the paper…
Reading the bibliography…
We theoretically study the landscape of the training error for neural networks in overparameterized cases.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Uniqueness of the weights for minimal feedforward nets with a given input-output map
H. J. Sussmann · 1992
Earlier work this paper cites.
Uniqueness of the weights for minimal feedforward nets with a given input-output map
H. J. Sussmann · 1992
Earlier work this paper cites.
Functionally equivalent feedforward neural networks
V. K u ̊ \mathring{\rm u} rková and P. C. Kainen · 1994
Earlier work this paper cites.
Functionally equivalent feedforward neural networks
V. K u ̊ \mathring{\rm u} rková and P. C. Kainen · 1994
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
S. Hochreiter and J. Schmidhuber · 1995
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
S. Hochreiter and J. Schmidhuber · 1995
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. A. McAllester · 1999
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. A. McAllester · 1999
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
K. Fukumizu and S. Amari · 2000
Cited alongside, same era.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
K. Fukumizu and S. Amari · 2000
Cited alongside, same era.
Simplified PAC-Bayesian margin bounds
D. McAllester · 2003
Cited alongside, same era.
Simplified PAC-Bayesian margin bounds
D. McAllester · 2003
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2018
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
B. Kleinberg, Y. Li, and Y. Yuan · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
B. Neyshabur, S. Bhojanapalli, and N. Srebro · 2018
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Cited alongside, same era.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. T. Chayes, L. Sagun, and R. Zecchina · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. T. Chayes, L. Sagun, and R. Zecchina · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
B. Kleinberg, Y. Li, and Y. Yuan · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
B. Neyshabur, S. Bhojanapalli, and N. Srebro · 2018
Later among the works it cites.
A Scale Invariant Flatness Measure for Deep Network Minima
A. Rangamani, N. H. Nguyen, A. Kumar, D. Phan, S. H. Chin, and T. D. Tran · 2019
Closest in time.
Y. Tsuzuku, I. Sato, and M. Sugiyama · 2019
Closest in time.
A Scale Invariant Flatness Measure for Deep Network Minima
A. Rangamani, N. H. Nguyen, A. Kumar, D. Phan, S. H. Chin, and T. D. Tran · 2019
Closest in time.
Y. Tsuzuku, I. Sato, and M. Sugiyama · 2019
Closest in time.