Fetching the paper…
Reading the bibliography…
We investigate the role of noise in optimization algorithms for learning over-parameterized models.
Towards understanding the importance of noise in training neural networks
Zhou, M., Liu, T., Li, Y., Lin, D., Zhou, E., and Zhao, T. (2019) · 1909
Earlier work this paper cites.
Revisiting landscape analysis in deep neural networks: Eliminating decreasing paths to infinity
Liang, S., Sun, R., and Srikant, R. (2019) · 1912
Earlier work this paper cites.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z., Wei, C., Lee, J. D., and Ma, T. (2020) · 2006
Earlier work this paper cites.
Deep learning and its applications to signal and information processing [exploratory dsp]
Yu, D. and Deng, L. (2010) · 2010
Earlier work this paper cites.
Randomized smoothing for stochastic optimization
Duchi, J. C., Bartlett, P. L., and Wainwright, M. J. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
On the inductive bias of dropout
Helmbold, D. P. and Long, P. M. (2015) · 2015
Earlier work this paper cites.
Global convergence of non-convex gradient descent for computing matrix squareroot
Jain, P., Jin, C., Kakade, S. M., and Netrapalli, P. (2015) · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T. (2015) · 2015
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2016) · 2016
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Ge, R., Jin, C., and Zheng, Y. (2017) · 2017
Cited alongside, same era.
Surprising properties of dropout in deep networks
Helmbold, D. P. and Long, P. M. (2017) · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I. (2017) · 2017
Cited alongside, same era.
Towards understanding regularization in batch normalization
Luo, P., Wang, X., Shao, W., and Peng, Z. (2018) · 2018
Later among the works it cites.
On the implicit bias of dropout
Mianjy, P., Arora, R., and Vidal, R. (2018) · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
Vershynin, R. (2018) · 2018
Later among the works it cites.
Recent trends in deep learning based natural language processing
Young, T., Hazarika, D., Poria, S., and Cambria, E. (2018) · 2018
Later among the works it cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Y., Ma, T., and Zhang, H. (2017) · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z. (2018) · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. and Lee, J. (2018) · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2018) · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Cited alongside, same era.
Adding one neuron can eliminate all bad local minima
Liang, S., Sun, R., Lee, J. D., and Srikant, R. (2018) · 2018
Cited alongside, same era.
Li, X., Lu, J., Arora, R., Haupt, J., Liu, H., Wang, Z., and Zhao, T. (2019) · 2019
Later among the works it cites.
Pa-gd: On the convergence of perturbed alternating gradient descent to second-order stationary points for structured nonconvex optimization
Lu, S., Hong, M., and Wang, Z. (2019) · 2019
Later among the works it cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Blanc, G., Gupta, N., Valiant, G., and Valiant, P. (2020) · 2020
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
E, W., Ma, C., and Wu, L. (2020) · 2020
Later among the works it cites.
Bounds on over-parameterization for guaranteed existence of descent paths in shallow relu networks
Sharifnassab, A., Salehkaleybar, S., and Golestani, S. J. (2020) · 2020
Later among the works it cites.