Fetching the paper…
Reading the bibliography…
There is growing evidence of the effectiveness of Shampoo, a higher-order preconditioning method, over Adam in deep learning optimization tasks.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Distributed second-order optimization using kronecker-factored approximations
Jimmy Ba, Roger Grosse, and James Martens · 2017
Earlier work this paper cites.
Preconditioned stochastic gradient descent
Xi-Lin Li · 2017
Earlier work this paper cites.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Earlier work this paper cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E. Dahl, Christopher J. Shallue, and Roger B. Grosse · 2019
Earlier work this paper cites.
Scalable second order optimization for deep learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
A trace-restricted kronecker-factored approximation to natural gradient
Kaixin Gao, Xiaolei Liu, Zhenghai Huang, Min Wang, Zidong Wang, Dachuan Xu, and Fan Yu · 2021
Earlier work this paper cites.
Tensor normal training for deep learning models
Yi Ren and Donald Goldfarb · 2021
Cited alongside, same era.
8-bit optimizers via block-wise quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Fishy: Layerwise fisher approximation for higher-order neural network optimization
Abel Peirson, Ehsan Amid, Yatong Chen, Vladimir Feinberg, Manfred K Warmuth, and Rohan Anil · 2022
Cited alongside, same era.
Randomized k-facs: Speeding up k-fac with randomized numerical linear algebra
Constantin Octavian Puiu · 2022
Cited alongside, same era.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Cited alongside, same era.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V. Le · 2023
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Google Gemini Team · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Closest in time.
When does second-order optimization speed up training?
Satoki Ishikawa and Rio Yokota · 2024
Closest in time.
Stochastic hessian fittings with lie groups, 2024
Xi-Lin Li · 2024
Closest in time.
Can we remove the square-root in adaptive gradient methods? A second-order perspective
Wu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae, Richard E. Turner, and Alireza Makhzani · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Benchmarking neural network training algorithms, 2023
George E. Dahl, Frank Schneider, Zachary Nado, Naman Agarwal, Chandramouli Shama Sastry, Philipp Hennig, Sourabh Medapati, Runa Eschenhagen, Priya Kasimbeg, Daniel Suo, Juhan Bae, Justin Gilmer, Abel L. Peirson, Bilal Khan, Rohan Anil, Mike Rabbat, Shankar Krishnan, Daniel Snider, Ehsan Amid, Kongtao Chen, Chris J. Maddison, Rakshith Vasudev, Michal Badura, Ankush Garg, and Peter Mattson · 2023
Cited alongside, same era.
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Cited alongside, same era.
Kronecker-factored approximate curvature for modern neural network architectures
Runa Eschenhagen, Alexander Immer, Richard E Turner, Frank Schneider, and Philipp Hennig · 2023
Cited alongside, same era.
No train no gain: Revisiting efficient training algorithms for transformer-based language models
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt J. Kusner · 2023
Cited alongside, same era.
Brand new k-facs: Speeding up k-fac with online decomposition updates, 2023
Constantin Octavian Puiu · 2023
Cited alongside, same era.
Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki, Jose Gallego-Posada, Zhijing Li, Kaushik Rangadurai, Dheevatsa Mudigere, and Michael Rabbat · 2023
Cited alongside, same era.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang, and Tengyu Ma · 2024
Closest in time.
Adalomo: Low-memory optimization with adaptive learning rate
Kai Lv, Hang Yan, Qipeng Guo, Haijun Lv, and Xipeng Qiu · 2024
Closest in time.
Full parameter fine-tuning for large language models with limited resources
Kai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo, and Xipeng Qiu · 2024
Closest in time.
Mlc algoperf benchmark competition
MLCommons · 2024
Closest in time.
A new perspective on shampoo’s preconditioner
Depen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach, Sham Kakade, and Lucas Janson · 2024
Closest in time.
Curvature-informed SGD via general purpose lie-group preconditioners, 2024
Omead Pooladzandi and Xi-Lin Li · 2024
Closest in time.
Resolving discrepancies in compute-optimal scaling of language models
Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, and Yair Carmon · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Adamem: Memory efficient momentum for adafactor
Nikhil Vyas, Depen Morwani, and Sham M. Kakade · 2024
Closest in time.
4-bit shampoo for memory-efficient network training
Sike Wang, Jia Li, Pan Zhou, and Hua Huang · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
Mitchell Wortsman, Peter J Liu, Lechao Xiao, Katie E Everett, Alexander A Alemi, Ben Adlam, John D Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, Jeffrey Pennington, Jascha Sohl-Dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2024
Closest in time.