Fetching the paper…
Reading the bibliography…
Recent work has highlighted the role of initialization scale in determining the structure of the solutions that gradient methods converge to.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 1906
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Approximate KKT points and a proximity measure for termination
Dutta, J., Deb, K., Tulshyan, R., and Arora, R · 2013
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Du, S. S., Hu, W., and Lee, J. D · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M. J · 2019
Cited alongside, same era.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Nacson, M. S., Gunasekar, S., Lee, J., Srebro, N., and Soudry, D · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Vaskevicius, T., Kanade, V., and Rebeschini, P · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Moroshko, E., Woodworth, B. E., Gunasekar, S., Lee, J. D., Srebro, N., and Soudry, D · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Razin, N. and Cohen, N · 2020
Later among the works it cites.
Implicit regularization in relu networks with the square loss
Vardi, G. and Shamir, O · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Li, Z., Luo, Y., and Lyu, K · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mirrorless mirror descent: A more natural discretization of riemannian gradient flow, 2020
Gunasekar, S., Woodworth, B., and Srebro, N · 2020
Cited alongside, same era.
Winnowing with gradient descent
Amid, E. and Warmuth, M. K
Cited in the paper.
Reparameterizing mirror descent as gradient descent
Amid, E. and Warmuth, M. K. K
Cited in the paper.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N
Cited in the paper.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N
Cited in the paper.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J
Cited in the paper.
On the proof of global convergence of gradient descent for deep relu networks with linear widths
Nguyen, Q · 2021
Closest in time.
A unifying view on implicit bias in training linear neural networks
Yun, C., Krishnan, S., and Mobahi, H · 2021
Closest in time.