Fetching the paper…
Reading the bibliography…
Deep learning have achieved promising results on a wide spectrum of AI applications.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020 · 1904
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O (1/kˆ 2). In Dokl. akad. nauk Sssr , Vol. 269. 543–547
Yurii E Nesterov. 1983 · 1983
Earlier work this paper cites.
First- and Second-Order Methods for Learning: Between Steepest Descent and Newton’s Method
Roberto Battiti. 1992 · 1992
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Shun-ichi Amari. 1998 · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian. 1999 · 1999
Earlier work this paper cites.
Adaptive Method of Realizing Natural Gradient Learning for Multilayer Perceptrons
Shun-ichi Amari, Hyeyoung Park, and Kenji Fukumizu. 2000 · 2000
Earlier work this paper cites.
Trust Region Methods
Andrew R. Conn, Nicholas I. M. Gould, and Philippe L. Toint. 2000 · 2000
Earlier work this paper cites.
The Tradeoffs of Large Scale Learning. In Proceedings of the 20th International Conference on Neural Information Processing Systems . 161–168
Léon Bottou and Olivier Bousquet. 2007 · 2007
Earlier work this paper cites.
A stochastic quasi-Newton method for online convex optimization. In Artificial intelligence and statistics . PMLR, 436–443
Nicol N Schraudolph, Jin Yu, and Simon Günter. 2007 · 2007
Earlier work this paper cites.
SGD-QN: Careful Quasi-Newton Stochastic Gradient Descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari. 2009 · 2009
Earlier work this paper cites.
Sharpness-Aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021 · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization. In Proceedings of the 27th International Conference on Machine Learning (ICML-10), June 21-24, 2010, Haifa, Israel . 735–742
James Martens. 2010 · 2010
Earlier work this paper cites.
Distributed training strategies for the structured perceptron. In Human language technologies: The 2010 annual conference of the North American chapter of the association for computational linguistics . 456–464
Ryan McDonald, Keith Hall, and Gideon Mann. 2010 · 2010
Earlier work this paper cites.
Martin Zinkevich, Markus Weimer, Alexander J. Smola, and Lihong Li. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Richard H. Byrd, Gillian M. Chin, Jorge Nocedal, and Yuchen Wu. 2012 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton. 2012 · 2012
Earlier work this paper cites.
ADADELTA: An Adaptive Learning Rate Method
Matthew D. Zeiler. 2012 · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013 (JMLR Workshop and Conference Proceedings, Vol. 28) . 1139–1147
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton. 2013a · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013 (JMLR Workshop and Conference Proceedings, Vol. 28) . 1139–1147
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton. 2013b · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky. 2014 · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu. 2014a · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In Fifteenth Annual Conference of the International Speech Communication Association . Citeseer
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 2014 · 2014
Earlier work this paper cites.
Understanding Machine Learning - From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition. In Conference on learning theory . PMLR, 797–842
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan. 2015 · 2015
Earlier work this paper cites.
Fast r-cnn. In Proceedings of the IEEE international conference on computer vision . 1440–1448
Ross Girshick. 2015 · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3431–3440
Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning . PMLR, 2408–2417
James Martens and Roger Grosse. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing. In INTERSPEECH 2015, 16th Annual Conference of the International Speech Communication Association, Dresden, Germany, September 6-10, 2015 . 1488–1492
Nikko Strom. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015 . IEEE Computer Society, 1–9
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning. In OSDI
Martín Abadi, Paul Barham, Jianmin Chen, Z. Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqian Zhang. 2016 · 2016
Earlier work this paper cites.
Second-order stochastic optimization in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan. 2016a · 2016
Earlier work this paper cites.
Finding Approximate Local Minima for Nonconvex Optimization in Linear Time
Naman Agarwal, Zeyuan Allen Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma. 2016b · 2016
Earlier work this paper cites.
Exact and Inexact Subsampled Newton Methods for Optimization
Raghu Bollapragada, Richard Byrd, and Jorge Nocedal. 2016 · 2016
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat. 2016 · 2016
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks. In 2016 2nd Workshop on Machine Learning in HPC Environments (MLHPC) . IEEE, 1–8
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen. 2016 · 2016
Earlier work this paper cites.
A Kronecker-factored approximate Fisher matrix for convolution layers. In ICML , Vol. 48. 573–582
Roger B Grosse and James Martens. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder. 2016 · 2016
Cited alongside, same era.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Parallel SGD: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré. 2016 · 2016
Cited alongside, same era.
Sparse Communication for Distributed Gradient Descent. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017 . 440–445
Alham Fikri Aji and Kenneth Heafield. 2017 · 2017
Memory Efficient Adaptive Optimization. In NeurIPS
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer. 2019 · 2019
Later among the works it cites.
Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis
Tal Ben-Nun and Torsten Hoefler. 2019 · 2019
Later among the works it cites.
On Empirical Comparisons of Optimizers for Deep Learning
Dami Choi, Christopher J. Shallue, Zachary Nado, Jaehoon Lee, Chris J. Maddison, and George E. Dahl. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1)
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Natural Compression for Distributed Deep Learning
Samuel Horvath, Chen-Yu Ho, Ludovit Horvath, Atal Narayan Sahu, Marco Canini, and Peter Richtárik. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding. In NIPS . 1709–1720
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning. In International Conference on Machine Learning . PMLR, 557–565
Aleksandar Botev, Hippolyt Ritter, and David Barber. 2017 · 2017
Cited alongside, same era.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina. 2017 · 2017
Cited alongside, same era.
AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
Aditya Devarakonda, Maxim Naumov, and Michael Garland. 2017 · 2017
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets. In International Conference on Machine Learning . PMLR, 1019–1028
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Neural Collaborative Filtering. In WWW . 173–182
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017 · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems . 1729–1739
Elad Hoffer, Itay Hubara, and Daniel Soudry. 2017 · 2017
Cited alongside, same era.
Densely Connected Convolutional Networks. In CVPR . 2261–2269
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes. In International Conference on Machine Learning . PMLR, 3252–3261
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi. 2019 · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Later among the works it cites.
On the Variance of the Adaptive Learning Rate and Beyond. In International Conference on Learning Representations
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. 2019 · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12359–12367
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka. 2019 · 2019
Later among the works it cites.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. In Interspeech 2019 . 2613–2617
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin Dogus Cubuk, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
A Survey on Deep Learning: Algorithms, Techniques, and Applications
Samira Pouyanfar, Saad Sadiq, Yilin Yan, Haiman Tian, Yudong Tao, Maria E. Presa Reyes, Mei-Ling Shyu, Shu-Ching Chen, and S. S. Iyengar. 2019 · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar. 2019 · 2019
Later among the works it cites.
DeepOBS: A Deep Learning Optimizer Benchmark Suite. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Frank Schneider, Lukas Balles, and Philipp Hennig. 2019 · 2019
Later among the works it cites.
Measuring the Effects of Data Parallelism on Neural Network Training
Christopher J. Shallue, Jaehoon Lee, Joseph M. Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. 2019 · 2019
Later among the works it cites.
Local SGD Converges Fast and Communicates Little. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Sebastian U. Stich. 2019 · 2019
Later among the works it cites.
Optimization for deep learning: theory and algorithms
Ruoyu Sun. 2019 · 2019
Later among the works it cites.
Yet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds
Masafumi Yamazaki, Akihiko Kasagi, Akihiro Tabuchi, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, and Kohta Nakashima. 2019 · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse. 2019 · 2019
Later among the works it cites.
Scalable second order optimization for deep learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. 2020 · 2020
Later among the works it cites.
A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 1597–1607
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020 · 2020
Later among the works it cites.
Practical Quasi-Newton Methods for Training Deep Neural Networks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020
Donald Goldfarb, Yi Ren, and Achraf Bahamou. 2020 · 2020
Later among the works it cites.
Bootstrap Your Own Latent: A new approach to self-supervised learning. In Neural Information Processing Systems
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Pires, Zhaohan Guo, Mohammad Azar, et al · 2020
Later among the works it cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick. 2020 · 2020
Later among the works it cites.
Don’t Use Large Mini-Batches, Use Local SGD
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi. 2020 · 2020
Later among the works it cites.
New Insights and Perspectives on the Natural Gradient Method
James Martens. 2020 · 2020
Later among the works it cites.
Convolutional neural network training with distributed K-FAC. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–12
J Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu, and Ian T Foster. 2020 · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 3505–3506
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
On the Generalization Benefit of Noise in Stochastic Gradient Descent. In International Conference on Machine Learning . PMLR, 9058–9067
Samuel Smith, Erich Elsen, and Soham De. 2020 · 2020
Later among the works it cites.
Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li. 2020 · 2020
Later among the works it cites.
A survey on large-scale machine learning
Meng Wang, Weijie Fu, Xiangnan He, Shijie Hao, and Xindong Wu. 2020 · 2020
Later among the works it cites.
On Layer Normalization in the Transformer Architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . 10524–10533
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Second-order Optimization for Non-convex Machine Learning: an Empirical Study. In Proceedings of the 2020 SIAM International Conference on Data Mining, SDM 2020, Cincinnati, Ohio, USA, May 7-9, 2020 . 199–207
Peng Xu, Fred Roosta, and Michael W. Mahoney. 2020 · 2020
Later among the works it cites.
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB’s Convergence Speed
Conglong Li, Ammar Ahmad Awan, Hanlin Tang, Samyam Rajbhandari, and Yuxiong He. 2021 · 2021
Closest in time.
Sparse-MLP: A Fully-MLP Architecture with Conditional Computation
Yuxuan Lou, Fuzhao Xue, Zangwei Zheng, and Yang You. 2021 · 2021
Closest in time.
A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes
Zachary Nado, Justin Gilmer, Christopher J. Shallue, Rohan Anil, and George E. Dahl. 2021 · 2021
Closest in time.
NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization
Ali Ramezani-Kebrya, Fartash Faghri, Ilia Markov, Vitaly Aksenov, Dan Alistarh, and Daniel M. Roy. 2021 · 2021
Closest in time.
Kronecker-factored Quasi-Newton Methods for Convolutional Neural Networks
Yi Ren and Donald Goldfarb. 2021 · 2021
Closest in time.
1-bit Adam: Communication Efficient Large-Scale Training with Adam’s Convergence Speed. In ICML
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He. 2021 · 2021
Closest in time.
Fuzhao Xue, Ziji Shi, Yuxuan Lou, Yong Liu, and Yang You. 2021 · 2021
Closest in time.