Fetching the paper…
Reading the bibliography…
Modern neural networks are typically trained in an over-parameterized regime where the parameters of the model far exceed the size of the training data.
Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen
Schur, J · 1911
Earlier work this paper cites.
Flat minima
Hochreiter, S., and Schmidhuber, J · 1997
Earlier work this paper cites.
The generic chaining: upper and lower bounds of stochastic processes
Talagrand, M · 2006
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Earlier work this paper cites.
Robust sparse regression under adversarial corruption
Chen, Y., Caramanis, C., and Mannor, S · 2013
Earlier work this paper cites.
Compressed sensing and matrix completion with constant proportion of corruptions
Li, X · 2013
Earlier work this paper cites.
Classification with asymmetric label noise: Consistency and maximal denoising
Scott, C., Blanchard, G., and Handy, G · 2013
Earlier work this paper cites.
Corrupted sensing: Novel guarantees for separating structured signals
Foygel, R., and Mackey, L · 2014
Earlier work this paper cites.
A comprehensive introduction to label noise
Frénay, B., Kabán, A., et al · 2014
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Reed, S., Lee, H., Anguelov, D., Szegedy, C., Erhan, D., and Rabinovich, A · 2014
Earlier work this paper cites.
Robust regression via hard thresholding
Bhatia, K., Jain, P., and Kar, P · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Earlier work this paper cites.
Diverse neural network learns true target functions
Xie, B., Liang, Y., and Song, L · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Computationally efficient robust sparse estimation in high dimensions
Balakrishnan, S., Du, S. S., Li, J., and Singh, A · 2017
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Earlier work this paper cites.
Learning from noisy singly-labeled data
Khetan, A., Lipton, Z. C., and Anandkumar, A · 2017
Cited alongside, same era.
Decoupling” when to update” from” how to update”
Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Deep learning is robust to massive label noise
Rolnick, D., Veit, A., Belongie, S., and Shavit, N · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2018
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y., and Liang, Y · 2018
Later among the works it cites.
High dimensional robust sparse regression
Liu, L., Shen, Y., Li, T., and Caramanis, C · 2018
Later among the works it cites.
Learning from binary labels with instance-dependent noise
Menon, A. K., van Rooyen, B., and Natarajan, N · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Oymak, S., and Soltanolkotabi, M · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Cited alongside, same era.
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
Azizan, N., and Hassibi, B · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L., and Bach, F · 2018
Cited alongside, same era.
Sever: A robust meta-algorithm for stochastic optimization
Diakonikolas, I., Kamath, G., Kane, D. M., Li, J., Steinhardt, J., and Stewart, A · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Later among the works it cites.
Robust estimation via robust gradient estimation
Prasad, A., Suggala, A. S., Balakrishnan, S., and Ravikumar, P · 2018
Later among the works it cites.
Learning to reweight examples for robust deep learning
Ren, M., Zeng, W., Yang, B., and Urtasun, R · 2018
Later among the works it cites.
Iteratively learning from the best
Shen, Y., and Sanghavi, S · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Song, M., Montanari, A., and Nguyen, P · 2018
Later among the works it cites.
Spurious valleys in two-layer neural network optimization landscapes
Venturi, L., Bandeira, A., and Bruna, J · 2018
Later among the works it cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhang, Z., and Sabuncu, M. R · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian
Oymak, S., Fabian, Z., Li, M., and Soltanolkotabi, M · 2019
Closest in time.
Oymak, S., and Soltanolkotabi, M · 2019
Closest in time.