Fetching the paper…
Reading the bibliography…
Adam has become one of the most favored optimizers in deep learning problems.
The geometry of sign gradient descent
Balles, L · 2002
Earlier work this paper cites.
On the equivalence of weak learnability and linear separability: new relaxations and efficient boosting algorithms
Shalev-Shwartz, S · 2010
Earlier work this paper cites.
Adam+: A stochastic method with adaptive variance reduction
Liu, M · 2011
Earlier work this paper cites.
Margins, shrinkage, and boosting
Telgarsky, M · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2015
Earlier work this paper cites.
Learning with incremental iterative regularization
Rosasco, L · 2015
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C · 2017
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L · 2018
Earlier work this paper cites.
SIGNSGD: compressed optimisation for non-convex problems
Bernstein, J · 2018
Earlier work this paper cites.
De, S · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y · 2018
Earlier work this paper cites.
On the convergence of adam and beyond
Reddi, S. J · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Arora, S · 2019
Earlier work this paper cites.
signsgd with majority vote is communication efficient and fault tolerant
Bernstein, J · 2019
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Chen, X · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ji, Z · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K · 2019
Cited alongside, same era.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Nacson, M. S · 2019
Cited alongside, same era.
The implicit bias of adagrad on separable data
Qian, Q · 2019
Cited alongside, same era.
Characterizing the implicit bias via a primal-dual analysis
Ji, Z · 2021
Later among the works it cites.
The implicit bias for adaptive optimization algorithms on homogeneous neural networks
Wang, B · 2021
Later among the works it cites.
Implicit bias in leaky relu networks trained on high-dimensional data
Frei, S · 2022
Later among the works it cites.
Toward understanding why adam converges faster than sgd for transformers
Pan, Y · 2022
Later among the works it cites.
Does momentum change the implicit regularization on separable data?
Wang, B · 2022
Later among the works it cites.
The implicit bias of batch normalization in linear models and two-layer linear convolutional neural networks
Cao, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The implicit regularization of stochastic gradient flow for least squares
Ali, A · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ji, Z · 2020
Cited alongside, same era.
The inductive bias of relu networks on orthogonally separable data
Phuong, M · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Razin, N · 2020
Cited alongside, same era.
Implicit regularization and convergence for weight normalization
Wu, X · 2020
Cited alongside, same era.
Lion secretly solves a constrained optimization: As lyapunov predicts
Chen, L · 2023
Later among the works it cites.
Implicit bias of gradient descent for mean squared error regression with two-layer wide neural networks
Jin, H · 2023
Later among the works it cites.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Kunstner, F · 2023
Later among the works it cites.
Implicit bias of gradient descent for logistic regression at the edge of stability
Wu, J · 2023
Later among the works it cites.
Understanding the generalization of adam in learning neural networks with proper regularization
Zou, D · 2023
Later among the works it cites.
Symbolic discovery of optimization algorithms
Chen, X · 2024
Closest in time.
Implicit bias of gradient descent for two-layer relu and leaky relu networks on nearly-orthogonal data
Kou, Y · 2024
Closest in time.
Implicit bias of adamw: ℓ ∞ \ell_{\infty} norm constrained optimization
Xie, S · 2024
Closest in time.
On the convergence of adaptive gradient methods for nonconvex optimization
Zhou, D · 2024
Closest in time.