Fetching the paper…
Reading the bibliography…
Training language models becomes increasingly expensive with scale, prompting numerous attempts to improve optimization efficiency.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Disentangling adaptive gradient methods from learning rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang · 2002
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
I.U.E. Nesterov · 2004
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients, 2018
Lukas Balles and Philipp Hennig · 2018
Earlier work this paper cites.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
signSGD with majority vote is communication efficient and fault tolerant
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar · 2019
Earlier work this paper cites.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, Yang Zhang, and Jonathan M Cohen · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Deepobs: A deep learning optimizer benchmark suite
Frank Schneider, Lukas Balles, and Philipp Hennig · 2019
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Cited alongside, same era.
Pytorch image models
Ross Wightman · 2019
Cited alongside, same era.
The geometry of sign gradient descent, 2020
Lukas Balles, Fabian Pedregosa, and Nicolas Le Roux · 2020
Cited alongside, same era.
Benchmarking in optimization: Best practice and open issues
Thomas Bartz-Beielstein, Carola Doerr, Daan van den Berg, Jakob Bossek, Sowmya Chandrasekaran, Tome Eftimov, Andreas Fischbach, Pascal Kerschke, William La Cava, Manuel Lopez-Ibanez, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Later among the works it cites.
How does adaptive optimization impact local neural network geometry?
Kaiqi Jiang, Dhruv Malik, and Yuanzhi Li · 2023
Later among the works it cites.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, and Suvrit Sra · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimizer benchmarking needs to account for hyperparameter tuning
Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, and François Fleuret · 2020
Cited alongside, same era.
Adasgd: Bridging the gap between sgd and adam
Jiaxuan Wang and Jenna Wiens · 2020
Cited alongside, same era.
Learning by turning: Neural architecture aware optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, and Yisong Yue · 2021
Cited alongside, same era.
Descending through a crowded valley-benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2021
Cited alongside, same era.
Tuning large neural networks via zero-shot hyperparameter transfer
Ge Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Cited alongside, same era.
Dissecting adaptive methods in gans, 2022
Samy Jelassi, David Dobre, Arthur Mensch, Yuanzhi Li, and Gauthier Gidel · 2022
Cited alongside, same era.
Toward understanding why adam converges faster than SGD for transformers
Yan Pan and Yuanzhi Li · 2022
Cited alongside, same era.
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Closest in time.
No train no gain: Revisiting efficient training algorithms for transformer-based language models
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt J Kusner · 2024
Closest in time.
How to fine-tune vision models with SGD
Ananya Kumar, Ruoqi Shen, Sebastien Bubeck, and Suriya Gunasekar · 2024
Closest in time.
Heavy-tailed class imbalance and why adam outperforms gradient descent on language models
Frederik Kunstner, Robin Yadav, Alan Milligan, Mark Schmidt, and Alberto Bietti · 2024
Closest in time.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang, and Tengyu Ma · 2024
Closest in time.
Resolving discrepancies in compute-optimal scaling of language models
Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, and Yair Carmon · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
Mitchell Wortsman, Peter J Liu, Lechao Xiao, Katie E Everett, Alexander A Alemi, Ben Adlam, John D Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, Jeffrey Pennington, Jascha Sohl-Dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2024
Closest in time.
Blockwise adaptivity: Faster training and better generalization in deep learning
Shuai Zheng and James T. Kwok · 2024
Closest in time.