Fetching the paper…
Reading the bibliography…
The accuracy of deep learning, i.e., deep neural networks, can be characterized by dividing the total error into three main types: approximation error, optimization error, and generalization error.
Arora, S., Du, S., Hu, W., Li, Z., Wang, R., 2019 · 1901
Earlier work this paper cites.
Generalization in deep networks: The role of distance from initialization
Nagarajan, V., Kolter, J., 2019b · 1901
Earlier work this paper cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Xu, Z., Zhang, Y., Luo, T., Xiao, Y., Ma, Z., 2019 · 1901
Earlier work this paper cites.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Cao, Y., Gu, Q., 2019 · 1902
Earlier work this paper cites.
Dying ReLU and initialization: Theory and numerical examples
Lu, L., Shin, Y., Su, Y., Karniadakis, G., 2019 · 1903
Earlier work this paper cites.
Data-dependent sample complexity of deep neural networks via Lipschitz augmentation
Wei, C., Ma, T., 2019 · 1905
Earlier work this paper cites.
Balls in ℝ k \mathbb{R}^{k} do not cut all subsets of k + 2 points
Dudley, R., 1979 · 1979
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G., 1989 · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., White, H., 1989 · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P., et al., 1998 · 1998
Earlier work this paper cites.
The elements of statistical learning. volume 1
Friedman, J., Hastie, T., Tibshirani, R., 2001 · 2001
Earlier work this paper cites.
Rademacher and Gaussian Complexities: Risk bounds and structural results
Bartlett, P., Mendelson, S., 2002 · 2002
Earlier work this paper cites.
The tradeoffs of large scale learning, in: Advances in neural information processing systems, pp. 161–168
Bottou, L., Bousquet, O., 2008 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., 2009 · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent, in: Proceedings of COMPSTAT’2010. Springer, pp. 177–186
Bottou, L., 2010 · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning, in: Advances in Neural Information Processing Systems
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., Ng, A., 2011 · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G., 2012 · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models, in: Proc. icml, p. 3
Maas, A., Hannun, A., Ng, A., 2013 · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., Srebro, N., 2014 · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., Singer, Y., 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D., Ba, J., 2015 · 2015
Cited alongside, same era.
Path-SGD: Path-normalized optimization in deep neural networks, in: Advances in Neural Information Processing Systems, pp. 2422–2430
Neyshabur, B., Salakhutdinov, R., Srebro, N., 2015 · 2015
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N., Mudigere, D., Nocedal, J., Smelyanskiy, M., Tang, P., 2016 · 2016
Opening the black box of deep neural networks via information
Shwartz-Ziv, R., Tishby, N., 2017 · 2017
Later among the works it cites.
Robust large margin deep neural networks
Sokolić, J., Giryes, R., Sapiro, G., Rodrigues, M., 2017 · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., Zhang, Y., 2018 · 2018
Later among the works it cites.
Data-dependent coresets for compressing neural networks with applications to generalization bounds
Baykal, C., Liebenwein, L., Gilitschenski, I., Feldman, D., Rus, D., 2018 · 2018
Later among the works it cites.
Reconciling modern machine learning and the bias-variance trade-off
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent converges to minimizers
Lee, J., Simchowitz, M., Jordan, M., Recht, B., 2016 · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al., 2016 · 2016
Cited alongside, same era.
Generalization error of invariant classifiers
Sokolic, J., Giryes, R., Sapiro, G., Rodrigues, M., 2016 · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O., 2016 · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets, in: Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org. pp. 1019–1028
Dinh, L., Pascanu, R., Bengio, S., Bengio, Y., 2017 · 2017
Cited alongside, same era.
Dziugaite, G., Roy, D., 2017 · 2017
Cited alongside, same era.
Fast rates for empirical risk minimization of strict saddle problems
Gonen, A., Shalev-Shwartz, S., 2017 · 2017
Cited alongside, same era.
Belkin, M., Hsu, D., Ma, S., Mandal, S., 2018 · 2018
Later among the works it cites.
Stability and convergence trade-off of iterative optimization algorithms
Chen, Y., Jin, C., Yu, B., 2018 · 2018
Later among the works it cites.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Cheng, Y., Wang, D., Zhou, P., Zhang, T., 2018 · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., Zhai, X., 2018 · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks, in: Advances in Neural Information Processing Systems, pp. 9461–9471
Gunasekar, S., Lee, J., Soudry, D., Srebro, N., 2018 · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data, in: Advances in Neural Information Processing Systems, pp. 8157–8166
Li, Y., Liang, Y., 2018 · 2018
Later among the works it cites.
Collapse of deep and narrow neural nets
Lu, L., Su, Y., Karniadakis, G., 2018 · 2018
Later among the works it cites.
On the spectral bias of deep neural networks
Rahaman, N., Arpit, D., Baratin, A., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., Courville, A., 2018 · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M., Gunasekar, S., Srebro, N., 2018 · 2018
Later among the works it cites.
Theory of deep learning IIb: Optimization properties of SGD
Zhang, C., Liao, Q., Rakhlin, A., Miranda, B., Golowich, N., Poggio, T., 2018 · 2018
Later among the works it cites.
Compressibility and generalization in large-scale deep learning
Zhou, W., Veitch, V., Austern, M., Adams, R., Orbanz, P., 2018 · 2018
Later among the works it cites.
The role of over-parametrization in generalization of neural networks, in: International Conference on Learning Representations
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., Srebro, N., 2019 · 2019
Closest in time.
On the information bottleneck theory of deep learning
Saxe, A.M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B.D., Cox, D.D., 2019 · 2019
Closest in time.