Fetching the paper…
Reading the bibliography…
It is believed that Gradient Descent (GD) induces an implicit bias towards good generalization in training machine learning models.
Perturbation bounds in connection with singular value decomposition
Wedin, P.-Å · 1972
Earlier work this paper cites.
Procrustes methods in the statistical analysis of shape
Goodall, C · 1991
Earlier work this paper cites.
Exact matrix completion via convex optimization
Candès, E. J. and Recht, B · 2009
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
Rudelson, M. and Vershynin, R · 2009
Earlier work this paper cites.
Matrix completion with noise
Candes, E. J. and Plan, Y · 2010
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Recht, B., Fazel, M., and Parrilo, P. A · 2010
Earlier work this paper cites.
Robust pca via outlier pursuit
Xu, H., Caramanis, C., and Sanghavi, S · 2010
Earlier work this paper cites.
Low rank matrix recovery for real-time cardiac mri
Zhao, B., Haldar, J. P., Brinegar, C., and Liang, Z.-P · 2010
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Earlier work this paper cites.
Low-rank matrix recovery via iteratively reweighted least squares minimization
Fornasier, M., Rauhut, H., and Ward, R · 2011
Earlier work this paper cites.
Scaled gradients on grassmann manifolds for matrix completion
Ngo, T. and Saad, Y · 2012
Earlier work this paper cites.
A unified approach to salient object detection via low rank matrix recovery
Shen, X. and Wu, Y · 2012
Earlier work this paper cites.
Matrix completion in colocated mimo radar: Recoverability, bounds & theoretical guarantees
Kalogerias, D. S. and Petropulu, A. P · 2013
Earlier work this paper cites.
Gradient methods for convex minimization: better rates under weaker conditions
Zhang, H. and Yin, W · 2013
Earlier work this paper cites.
Segmentation driven low-rank matrix recovery for saliency detection
Zou, W., Kpalma, K., Liu, Z., and Ronsin, J · 2013
Earlier work this paper cites.
Reweighted low-rank matrix recovery and its application in image restoration
Peng, Y., Suo, J., Dai, Q., and Xu, W · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
An overview of low-rank matrix recovery from incomplete observations
Davenport, M. A. and Romberg, J · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T · 2016
Earlier work this paper cites.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
Low-rank solutions of linear matrix equations via Procrustes flow
Tu, S., Boczar, R., Simchowitz, M., Soltanolkotabi, M., and Recht, B · 2016
Cited alongside, same era.
Guarantees of riemannian optimization for low rank matrix recovery
Wei, K., Cai, J.-F., Chan, T. F., and Leung, S · 2016
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Ge, R., Jin, C., and Zheng, Y · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Blanc, G., Gupta, N., Valiant, G., and Valiant, P · 2020
Later among the works it cites.
The surprising simplicity of the early-time learning dynamics of neural networks
Hu, W., Xiao, L., Adlam, B., and Pennington, J · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
Ji, Z. and Telgarsky, M · 2020
Later among the works it cites.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Li, Z., Luo, Y., and Lyu, K · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Global optimality in low-rank matrix optimization
Zhu, Z., Li, Q., Tang, G., and Wakin, M. B · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Cited alongside, same era.
Razin, N. and Cohen, N · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Damian, A., Ma, T., and Lee, J. D · 2021
Later among the works it cites.
Provable generalization of sgd-trained neural networks of any width in the presence of adversarial label noise
Frei, S., Cao, Y., and Gu, Q · 2021
Later among the works it cites.
Deep linear networks dynamics: Low-rank biases induced by initialization scale and l2 regularization
Jacot, A., Ged, F., Gabriel, F., Şimşek, B., and Hongler, C · 2021
Later among the works it cites.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Lyu, K., Li, Z., Wang, R., and Arora, S · 2021
Later among the works it cites.
Implicit regularization in tensor factorization
Razin, N., Maman, A., and Cohen, N · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Stöger, D. and Soltanolkotabi, M · 2021
Later among the works it cites.
Accelerating ill-conditioned low-rank matrix estimation via scaled gradient descent
Tong, T., Ma, C., and Chi, Y · 2021
Later among the works it cites.
The global optimization geometry of low-rank matrix optimization
Zhu, Z., Li, Q., Tang, G., and Wakin, M. B · 2021
Later among the works it cites.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Boursier, E., Pillaud-Vivien, L., and Flammarion, N · 2022
Later among the works it cites.
Algorithmic regularization in model-free overparametrized asymmetric matrix factorization
Jiang, L., Chen, Y., and Ding, L · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Lyu, K., Li, Z., and Arora, S · 2022
Later among the works it cites.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Razin, N., Maman, A., and Cohen, N · 2022
Later among the works it cites.
Implicit regularization towards rank minimization in relu networks
Timor, N., Vardi, G., and Shamir, O · 2022
Later among the works it cites.