Fetching the paper…
Reading the bibliography…
Adaptive gradient optimizers like Adam(W) are the default training algorithms for many deep learning architectures, such as transformers.
Stein’s Lemma for the Reparameterization Trick with Exponential-family Mixtures
Lin, W., Khan, M. E., and Schmidt, M · 1910
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Becker, S., Le Cun, Y., et al · 1988
Earlier work this paper cites.
Gaussian adaptation, an evolution-based efficient global optimizer
Taxén, L. and Kjellström, G · 1992
Earlier work this paper cites.
An overview of evolutionary algorithms for parameter optimization
Bäck, T. and Schwefel, H.-P · 1993
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming
Nesterov, Y. and Nemirovskii, A · 1994
Earlier work this paper cites.
Ensemble learning for multi-layer networks
Barber, D. and Bishop, C · 1997
Earlier work this paper cites.
Numerical optimization
Nocedal, J. and Wright, S. J · 1999
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S · 2006
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Roux, N., Manzagol, P.-A., and Bengio, Y · 2007
Earlier work this paper cites.
Natural evolution strategies
Wierstra, D., Schaul, T., Peters, J., and Schmidhuber, J · 2008
Earlier work this paper cites.
Adaptive regularization of weight vectors
Crammer, K., Kulesza, A., and Dredze, M · 2009
Earlier work this paper cites.
The variational Gaussian approximation revisited
Opper, M. and Archambeau, C · 2009
Earlier work this paper cites.
Exponential natural evolution strategies
Glasmachers, T., Schaul, T., Yi, S., Wierstra, D., and Schmidhuber, J · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Training deep and recurrent networks with hessian-free optimization
Martens, J. and Sutskever, I · 2012
Earlier work this paper cites.
Staines, J. and Barber, D · 2012
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
The multilinear normal distribution: Introduction and some basic properties
Ohlson, M., Ahmad, M. R., and Von Rosen, D · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Conjugate-computation variational inference: Converting variational inference in non-conjugate models to inferences in conjugate models
Khan, M. and Lin, W · 2017
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Hui, L. and Belkin, M · 2020
Later among the works it cites.
Handling the positive-definite constraint in the bayesian learning rule
Lin, W., Schmidt, M., and Khan, M. E · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
Martens, J · 2020
Later among the works it cites.
Sadam: A variant of adam for strongly convex functions
Wang, G., Lu, S., Cheng, Q., Tu, W., and Zhang, L · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khan, M. E., Lin, W., Tangkaratt, V., Liu, Z., and Nielsen, D · 2017
Cited alongside, same era.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al · 2017
Cited alongside, same era.
Variants of rmsprop and adagrad with logarithmic regret bounds
Mukkamala, M. C. and Hein, M · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Cited alongside, same era.
Shampoo: Preconditioned Stochastic Tensor Optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Cited alongside, same era.
Why are adaptive methods good for attention models?
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S · 2020
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Tensor normal training for deep learning models
Ren, Y. and Goldfarb, D · 2021
Later among the works it cites.
Analytic natural gradient updates for cholesky factor in gaussian variational approximation
Tan, L. S · 2022
Later among the works it cites.
Focal modulation networks
Yang, J., Li, C., Dai, X., and Gao, J · 2022
Later among the works it cites.
Graph attention multi-layer perceptron
Zhang, W., Yin, Z., Sheng, Z., Li, Y., Ouyang, W., Li, X., Tao, Y., Yang, Z., and Cui, B · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., et al · 2023
Later among the works it cites.
Convolutions through the lens of tensor networks
Dangel, F · 2023
Later among the works it cites.
Global context vision transformers
Hatamizadeh, A., Yin, H., Heinrich, G., Kautz, J., and Molchanov, P · 2023
Later among the works it cites.
The bayesian learning rule
Khan, M. E. and Rue, H · 2023
Later among the works it cites.
Kunstner, F., Chen, J., Lavington, J. W., and Schmidt, M · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Liu, H., Li, Z., Hall, D., Liang, P., and Ma, T · 2023
Later among the works it cites.
Shi, H.-J. M., Lee, T.-H., Iwasaki, S., Gallego-Posada, J., Li, Z., Rangadurai, K., Mudigere, D., and Rabbat, M · 2023
Later among the works it cites.
Vmamba: Visual state space model
Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., and Liu, Y · 2024
Closest in time.