Fetching the paper…
Reading the bibliography…
Shampoo is an online and stochastic optimization algorithm belonging to the AdaGrad family of methods for training neural networks.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro. 1951 · 1951
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986 · 1986
Earlier work this paper cites.
Comparing Biases for Minimal Network Construction with Back-Propagation. In NeurIPS
Stephen J. Hanson and Lorien Y. Pratt. 1988 · 1988
Earlier work this paper cites.
Stochastic Gradient Learning in Neural Networks
Léon Bottou et al · 1991
Earlier work this paper cites.
cceleration of Stochastic Approximation by Averaging
Boris T Polyak and Anatoli B Juditsky. 1992 · 1992
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Shun-Ichi Amari. 1998 · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
On the Generalization Ability of On-Line Learning Algorithms
Nicolo Cesa-Bianchi, Alex Conconi, and Claudio Gentile. 2001 · 2001
Earlier work this paper cites.
Subgradient Methods
Stephen Boyd, Lin Xiao, and Almir Mutapcic. 2003 · 2003
Earlier work this paper cites.
A Stochastic Quasi-Newton Method for Online Convex Optimization. In AISTATS
Nicol N Schraudolph, Jin Yu, and Simon Günter. 2007 · 2007
Earlier work this paper cites.
Functions of Matrices: Theory and Computation
Nicholas J Higham. 2008 · 2008
Earlier work this paper cites.
SGD-QN: Careful Quasi-Newton Stochastic Gradient Descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari. 2009 · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database. In CVPR
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Large-Scale Machine Learning with Stochastic Gradient Descent. In COMPSTAT
Léon Bottou. 2010 · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization. In ICML
James Martens et al · 2010
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Learning Recurrent Neural Networks with Hessian-Free Optimization. In ICML
James Martens and Ilya Sutskever. 2011 · 2011
Earlier work this paper cites.
Training Deep and Recurrent Networks with Hessian-Free Optimization
James Martens and Ilya Sutskever. 2012 · 2012
Earlier work this paper cites.
Matrix Computations
Gene H Golub and Charles F Van Loan. 2013 · 2013
Earlier work this paper cites.
Training Highly Multiclass Classifiers
Maya R Gupta, Samy Bengio, and Jason Weston. 2014 · 2014
Earlier work this paper cites.
Compressing Neural Networks with the Hashing Trick. In ICML
Wenlin Chen, James Wilson, Stephen Tyree, Kilian Weinberger, and Yixin Chen. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In ICLR
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature. In ICML
James Martens and Roger Grosse. 2015 · 2015
Earlier work this paper cites.
A Multi-Batch L-BFGS Method for Machine Learning. In NeurIPS
Albert S Berahas, Jorge Nocedal, and Martin Takác. 2016 · 2016
Earlier work this paper cites.
Incorporating Nesterov Momentum into Adam
Timothy Dozat. 2016 · 2016
Cited alongside, same era.
A Kronecker-factored approximate Fisher matrix for convolution layers. In ICML
Roger Grosse and James Martens. 2016 · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Sub-sampled Newton Methods with Non-uniform Sampling
Peng Xu, Jiyan Yang, Fred Roosta, Christopher Ré, and Michael W Mahoney. 2016 · 2016
Cited alongside, same era.
Distributed Second-Order Optimization using Kronecker-Factored Approximations. In ICLR
Jimmy Ba, Roger Grosse, and James Martens. 2017 · 2017
Cited alongside, same era.
AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
Disentangling Adaptive Gradient Methods from Learning Rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang. 2020 · 2020
Later among the works it cites.
Scalable Second Order Optimization for Deep Learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. 2020 · 2020
Later among the works it cites.
An investigation of Newton-Sketch and subsampled Newton methods
Albert S Berahas, Raghu Bollapragada, and Jorge Nocedal. 2020 · 2020
Later among the works it cites.
Aaron Defazio. 2020 · 2020
Later among the works it cites.
Practical Quasi-Newton Methods for Training Deep Neural Networks
Donald Goldfarb, Yi Ren, and Achraf Bahamou. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Devarakonda, Maxim Naumov, and Michael Garland. 2017 · 2017
Cited alongside, same era.
Newton Sketch: A Near Linear-Time Optimization Algorithm with Linear-Quadratic Convergence
Mert Pilanci and Martin J Wainwright. 2017 · 2017
Cited alongside, same era.
Adaptive Sampling Strategies for Stochastic Optimization
Raghu Bollapragada, Richard Byrd, and Jorge Nocedal. 2018a · 2018
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal. 2018 · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018 · 2018
Cited alongside, same era.
Fast Approximate Natural Gradient Descent in a Kronecker Factored Eigenbasis. In NeurIPS
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent. 2018 · 2018
Cited alongside, same era.
Shampoo: Preconditioned Stochastic Tensor Optimization. In ICML
Vineet Gupta, Tomer Koren, and Yoram Singer. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
New Insights and Perspectives on the Natural Gradient Method
James Martens. 2020 · 2020
Later among the works it cites.
ZeRO: memory optimizations toward training trillion parameter models. In SC
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems. In SIGKDD
Hao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, and Jiyan Yang. 2020 · 2020
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact Hessian information
Peng Xu, Fred Roosta, and Michael W Mahoney. 2020a · 2020
Later among the works it cites.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes. In ICLR
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020 · 2020
Later among the works it cites.
Distributed Shampoo Implementation
Rohan Anil and Vineet Gupta. 2021 · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Later among the works it cites.
ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning. In SC
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Later among the works it cites.
Tensor Normal Training for Deep Learning Models
Yi Ren and Donald Goldfarb. 2021 · 2021
Later among the works it cites.
On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models
Rohan Anil, Sandra Gadanho, Da Huang, Nijith Jacob, Zhuoshu Li, Dong Lin, Todd Phillips, Cristina Pop, Kevin Regan, Gil I Shamir, et al · 2022
Later among the works it cites.
Quasi-Newton methods for machine learning: forget the past, just sample
Albert S Berahas, Majid Jahani, Peter Richtárik, and Martin Takáč. 2022 · 2022
Later among the works it cites.
TorchRec: a PyTorch Domain Library for Recommendation Systems. In RecSys
Dmytro Ivchenko, Dennis Van Der Staay, Colin Taylor, Xing Liu, Will Feng, Rahul Kindi, Anirudh Sudarshan, and Shahin Sefati. 2022 · 2022
Later among the works it cites.
Software-hardware co-design for fast and scalable training of deep learning recommendation models. In ISCA
Dheevatsa Mudigere, Yuchen Hao, Jianyu Huang, Zhihao Jia, Andrew Tulloch, Srinivas Sridharan, Xing Liu, Mustafa Ozdal, Jade Nie, Jongsoo Park, et al · 2022
Later among the works it cites.
Low-Rank Updates of Matrix Square Roots
Shany Shumeli, Petros Drineas, and Haim Avron. 2022 · 2022
Later among the works it cites.
Fast Differentiable Matrix Square Root
Yue Song, Nicu Sebe, and Wei Wang. 2022 · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, and Francesco Orabona. 2022 · 2022
Later among the works it cites.
Computing the Square Root of a Low-Rank Perturbation of the Scaled Identity Matrix
Massimiliano Fasi, Nicholas J Higham, and Xiaobo Liu. 2023 · 2023
Closest in time.
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al · 2023
Closest in time.