Fetching the paper…
Reading the bibliography…
SGD performs worse than Adam by a significant margin on Transformers, but the reason remains unclear.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
J. Zhang, T. He, S. Sra, and A. Jadbabaie · 1905
Earlier work this paper cites.
Improving deep transformer with depth-scaled initialization and merged attention
B. Zhang, I. Titov, and R. Sennrich · 1908
Earlier work this paper cites.
An iteration method for the solution of the eigenvalue problem of linear differential and integral operators
C. Lanczos · 1950
Earlier work this paper cites.
On best conditioned matrices
G. E. Forsythe and E. G. Straus · 1955
Earlier work this paper cites.
Calculation of gauss quadrature rules
G. H. Golub and J. H. Welsch · 1969
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
M. F. Hutchinson · 1989
Earlier work this paper cites.
A direct proof of the christoffel-darboux identity and its equivalence to the recurrence relationship
C. Brezinski · 1990
Earlier work this paper cites.
Estimates in quadratic formulas
G. H. Golub and Z. Strakoš · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Bounds for the trace of the inverse and the determinant of symmetric positive definite matrices
Z. Bai and G. H. Golub · 1996
Earlier work this paper cites.
Some large-scale matrix computation problems
Z. Bai, G. Fahey, and G. Golub · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Lanczos algorithms for large symmetric eigenvalue computations: Vol. I: Theory
J. K. Cullum and R. A. Willoughby · 2002
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2002
Earlier work this paper cites.
Large scale machine learning
R. Collobert · 2004
Earlier work this paper cites.
Matrices, moments and quadrature with applications , volume 30
G. H. Golub and G. Meurant · 2009
Earlier work this paper cites.
Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix
H. Avron and S. Toledo · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Numerical methods for large eigenvalue problems: revised edition
Y. Saad · 2011
Earlier work this paper cites.
On the convergence of block coordinate descent type methods
A. Beck and L. Tetruashvili · 2013
Earlier work this paper cites.
An introduction to numerical methods and analysis
J. F. Epperson · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Approximating spectral densities of large matrices
L. Lin, Y. Saad, and C. Yang · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
L. Sagun, L. Bottou, and Y. LeCun · 2016
Earlier work this paper cites.
Why momentum really works
G. Goh · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2017
Earlier work this paper cites.
Fast estimation of tr(f(a)) via stochastic lanczos quadrature
S. Ubaru, J. Chen, and Y. Saad · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Estimating the spectral density of large implicit matrices
R. P. Adams, J. Pennington, M. J. Johnson, J. Smith, Y. Ovadia, B. Patton, and J. Saunderson · 2018
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Earlier work this paper cites.
The best of both worlds: Combining recent advances in neural machine translation
M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. Foster, L. Jones, N. Parmar, M. Schuster, Z. Chen, et al · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
G. Gur-Ari, D. A. Roberts, and E. Dyer · 2018
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2018
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Y. Dong, J.-B. Cordonnier, and A. Loukas · 2021
Later among the works it cites.
Hessian eigenspectra of more realistic nonlinear models
Z. Liao and M. W. Mahoney · 2021
Later among the works it cites.
A deeper look at the hessian eigenspectrum of deep neural networks and its applications to regularization
A. R. Sankar, Y. Khasbage, R. Vigneswaran, and V. N. Balasubramanian · 2021
Later among the works it cites.
Worst-case complexity of cyclic coordinate descent: O (nˆ 2) o (n 2) gap with randomized version
R. Sun and Y. Ye · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, et al · 2021
Later among the works it cites.
Early convolutions help transformers see better
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Papyan · 2018
Cited alongside, same era.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
Hessian-based analysis of large batch training and robustness to adversaries
Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Mahoney · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
D. Zhou, J. Chen, Y. Cao, Y. Tang, Z. Yang, and Q. Gu · 2018
Cited alongside, same era.
Non-convergence and limit cycles in the adam optimizer
S. Bock and M. Weiß · 2019
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2019
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2019
Cited alongside, same era.
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick · 2021
Later among the works it cites.
Towards practical adam: Non-convexity, convergence theory, and mini-batch acceleration
C. Chen, L. Shen, F. Zou, and W. Liu · 2022
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
M. Crawshaw, M. Liu, F. Orabona, W. Zhang, and Z. Zhuang · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Later among the works it cites.
A simple convergence proof of adam and adagrad
A. Défossez, L. Bottou, F. Bach, and N. Usunier · 2022
Later among the works it cites.
Asymptotic study of stochastic adaptive algorithms in non-convex landscape
S. Gadat and I. Gavra · 2022
Later among the works it cites.
Super-acceleration with cyclical step-sizes
B. Goujaud, D. Scieur, A. Dieuleveut, A. B. Taylor, and F. Pedregosa · 2022
Later among the works it cites.
G. Lavezzi, K. Guye, and M. Ciarcià · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
L. Noci, S. Anagnostidis, L. Biggio, A. Orvieto, S. P. Singh, and A. Lucchi · 2022
Later among the works it cites.
Adaptive inertia: Disentangling the effects of adaptive learning rate and momentum
Z. Xie, X. Wang, H. Zhang, I. Sato, and M. Sugiyama · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao · 2022
Later among the works it cites.
Glm-130b: An open bilingual pre-trained model
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, et al · 2022
Later among the works it cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
K. Ahn, X. Cheng, M. Song, C. Yun, A. Jadbabaie, and S. Sra · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Later among the works it cites.
How does adaptive optimization impact local neural network geometry?
K. Jiang, D. Malik, and Y. Li · 2023
Later among the works it cites.
F. Kunstner, J. Chen, J. W. Lavington, and M. Schmidt · 2023
Later among the works it cites.
Convergence of adam under relaxed assumptions
H. Li, A. Rakhlin, and A. Jadbabaie · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
H. Liu, Z. Li, D. Hall, P. Liang, and T. Ma · 2023
Later among the works it cites.
Full parameter fine-tuning for large language models with limited resources
K. Lv, Y. Yang, T. Liu, Q. Gao, Q. Guo, and X. Qiu · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
S. Malladi, T. Gao, E. Nichani, A. Damian, J. D. Lee, D. Chen, and S. Arora · 2023
Later among the works it cites.
A theory on adam instability in large-scale machine learning
I. Molybog, P. Albert, M. Chen, Z. DeVito, D. Esiobu, N. Goyal, P. S. Koura, S. Narang, A. Poulton, R. Silva, et al · 2023
Later among the works it cites.
Toward understanding why adam converges faster than sgd for transformers
Y. Pan and Y. Li · 2023
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
M. Wortsman, P. J. Liu, L. Xiao, K. Everett, A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, et al · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
A. Yang, B. Xiao, B. Wang, B. Zhang, C. Bian, C. Yin, C. Lv, D. Pan, D. Wang, D. Yan, et al · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
S. Zhai, T. Likhomanenko, E. Littwin, D. Busbridge, J. Ramapuram, Y. Zhang, J. Gu, and J. M. Susskind · 2023
Later among the works it cites.
Gaussian quadrature — Wikipedia, the free encyclopedia, 2023
Wikipedia · 2024
Closest in time.