Fetching the paper…
Reading the bibliography…
Can a neural network minimizing cross-entropy learn linearly separable data? Despite progress in the theory of deep learning, this question remains unsolved.
Training a 3-node neural network is np-complete
Blum, A. L. and Rivest, R. L · 1992
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
Stochastic subgradient method converges on tame functions, 2018
Davis, D., Drusvyatskiy, D., Kakade, S., and Lee, J. D · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Allen-Zhu, Z. and Li, Y · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
Why do larger models generalize better? a theoretical perspective via the xor problem
Brutzkus, A. and Globerson, A · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Yehudai, G. and Shamir, O · 2019
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Later among the works it cites.
Learning parities with neural networks
Daniely, A. and Malach, E · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning, 2020
Ji, Z. and Telgarsky, M · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond NTK
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cao, Y. and Gu, Q · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Decoupling gating from linearity
Fiat, J., Malach, E., and Shalev-Shwartz, S · 2019
Cited alongside, same era.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Nacson, M. S., Gunasekar, S., Lee, J., Srebro, N., and Soudry, D · 2019
Cited alongside, same era.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
Wang, G., Giannakis, G. B., and Chen, J · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M
Cited in the paper.
Li, Y., Ma, T., and Zhang, H. R · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2020
Later among the works it cites.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Moroshko, E., Gunasekar, S., Woodworth, B., Lee, J. D., Srebro, N., and Soudry, D · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B. E., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
The inductive bias of relu networks on orthogonally separable data
Phuong, M. and Lampert, C · 2021
Closest in time.