Fetching the paper…
Reading the bibliography…
Training deep neural networks on large datasets can often be accelerated by using multiple compute nodes.
Application of programs with maximin objective functions to problems of optimal resource allocation
Seymour Kaplan · 1974
Earlier work this paper cites.
Problems in decentralized decision making and computation
John Nikolas Tsitsiklis · 1984
Earlier work this paper cites.
A bridging model for parallel computation
Leslie G Valiant · 1990
Earlier work this paper cites.
Laplacian matrices of graphs: a survey
Russell Merris · 1994
Earlier work this paper cites.
The mosek interior point optimizer for linear programming: An implementation of the homogeneous algorithm
Erling D. Andersen and Knud D. Andersen · 2000
Earlier work this paper cites.
Kademlia: A peer-to-peer information system based on the xor metric
Petar Maymounkov and David Mazieres · 2002
Earlier work this paper cites.
Reversible markov chains and random walks on graphs, 2002. unfinished monograph, recompiled 2014, 2002
David Aldous and James Allen Fill · 2002
Earlier work this paper cites.
Looking up data in p2p systems
Hari Balakrishnan, M Frans Kaashoek, David Karger, Robert Morris, and Ion Stoica · 2003
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Lin Xiao and Stephen Boyd · 2004
Earlier work this paper cites.
Randomized gossip algorithms
Stephen Boyd, Arpita Ghosh, Balaji Prabhakar, and Devavrat Shah · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Pitch Patarasuk and Xin Yuan · 2009
Earlier work this paper cites.
Asynchronous gossip algorithms for stochastic optimization
S Sundhar Ram, A Nedić, and Venugopal V Veeravalli · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc' aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc Le, and Andrew Ng · 2012
Earlier work this paper cites.
Distributed autonomous online learning: Regrets and intrinsic privacy-preserving properties
Feng Yan, Shreyas Sundaram, SVN Vishwanathan, and Yuan Qi · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li · 2014
Earlier work this paper cites.
Distributed optimization over time-varying directed graphs
Angelia Nedić and Alex Olshevsky · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
On network throughput variability in microsoft azure cloud
V. Persico, P. Marchetta, A. Botta, and A. Pescape · 2015
Earlier work this paper cites.
Measuring network throughput in the cloud: The case of amazon ec2
Valerio Persico, Pietro Marchetta, Alessio Botta, and Antonio Pescapè · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Collective algorithms for multiported torus networks
Paul Sack and William Gropp · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Earlier work this paper cites.
Stochastic gradient-push for strongly convex functions on time-varying directed graphs
Angelia Nedić and Alex Olshevsky · 2016
Earlier work this paper cites.
On the convergence of decentralized gradient descent
Kun Yuan, Qing Ling, and Wotao Yin · 2016
Earlier work this paper cites.
Generating a function for network delay
Andrei M Sukhov, MA Astrakhantseva, AK Pervitsky, SS Boldyrev, and AA Bukatov · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Fast asynchronous parallel stochastic gradient descent: A lock-free approach with convergence guarantee
Shen-Yi Zhao and Wu-Jun Li · 2016
Earlier work this paper cites.
Arock: an algorithmic framework for asynchronous parallel coordinate updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2016
Earlier work this paper cites.
An asynchronous mini-batch algorithm for regularized stochastic optimization
Hamid Reza Feyzmahdavian, Arda Aytekin, and Mikael Johansson · 2016
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2017
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Practical secure aggregation for privacy-preserving machine learning
Aaron Segal, Antonio Marcedone, Benjamin Kreuter, Daniel Ramage, H. Brendan McMahan, Karn Seth, K. A. Bonawitz, Sarvar Patel, and Vladimir Ivanov · 2017
Cited alongside, same era.
Proteus: Agile ml elasticity through tiered reliability in dynamic resource markets
Aaron Harlap, Alexey Tumanov, Andrew Chung, Gregory R. Ganger, and Phillip B. Gibbons · 2017
Cited alongside, same era.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2017
Cited alongside, same era.
Qsgd: communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Z Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, X Yu Felix, Sanjiv Kumar, and H Brendan McMahan · 2017
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Demystifying gpt-3 language model: A technical overview, 2020
Chuan Li · 2020
Later among the works it cites.
MLPerf Training Benchmark
Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debojyoti Dutta, Udit Gupta, Kim Hazelwood, Andrew Hock, Xinyuan Huang, Bill Jia, Daniel Kang, David Kanter, Naveen Kumar, Jeffery Liao, Guokai Ma, Deepak Narayanan, Tayo Oguntebi, Gennady Pekhimenko, Lillian Pentecost, Vijay Janapa Reddi, Taylor Robie, Tom St. John, Carole-Jean Wu, Lingjie Xu, Cliff Young, and Matei Zaharia · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Terngrad: ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Asaga: asynchronous parallel saga
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2017
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Cited alongside, same era.
Federated learning for mobile keyboard prediction, 2018
Andrew Hard, Chloé M Kiddon, Daniel Ramage, Francoise Beaufays, Hubert Eichner, Kanishka Rao, Rajiv Mathews, and Sean Augenstein · 2018
Cited alongside, same era.
Applied federated learning: Improving google keyboard query suggestions, 2018
Timothy Yang, Galen Andrew, Hubert Eichner, Haicheng Sun, Wei Li, Nicholas Kong, Daniel Ramage, and Françoise Beaufays · 2018
Cited alongside, same era.
A hybrid gpu cluster and volunteer computing platform for scalable deep learning
Ekasit Kijsipongse, Apivadee Piyatumrong, and Suriya U-ruekolan · 2018
Cited alongside, same era.
Optimal algorithms for non-smooth distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Laurent Massoulié, and Yin Tat Lee · 2018
Cited alongside, same era.
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
A survey on distributed machine learning
Joost Verbraeken, Matthijs Wolting, Jonathan Katzy, Jeroen Kloppenburg, Tim Verbelen, and Jan S. Rellermeyer · 2020
Later among the works it cites.
A unified architecture for accelerating distributed DNN training in heterogeneous gpu/cpu clusters
Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo · 2020
Later among the works it cites.
Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data
Micah J. Sheller, Brandon Edwards, G. Anthony Reina, Jason Martin, Sarthak Pati, Aikaterini Kotrotsou, Mikhail Milchenko, Weilin Xu, Daniel Marcus, Rivka R. Colen, and Spyridon Bakas · 2020
Later among the works it cites.
Towards crowdsourced training of large neural networks using decentralized mixture-of-experts
Max Ryabinin and Anton Gusev · 2020
Later among the works it cites.
A dual approach for optimal algorithms in distributed optimization over networks
César A Uribe, Soomin Lee, Alexander Gasnikov, and Angelia Nedić · 2020
Later among the works it cites.
Scaffold: Stochastic controlled averaging for federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh · 2020
Later among the works it cites.
Tighter theory for local sgd on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Later among the works it cites.
Is local sgd better than minibatch sgd?
Blake Woodworth, Kumar Kshitij Patel, Sebastian Stich, Zhen Dai, Brian Bullins, Brendan Mcmahan, Ohad Shamir, and Nathan Srebro · 2020
Later among the works it cites.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian Stich · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Later among the works it cites.
Multi-node bert-pretraining: Cost-efficient approach, 2020
Jiahuang Lin, Xin Li, and Gennady Pekhimenko · 2020
Later among the works it cites.
Toward communication efficient adaptive gradient method
Xiangyi Chen, Xiaoyun Li, and Ping Li · 2020
Later among the works it cites.
Distributed algorithms for composite optimization: Unified and tight convergence analysis
Jinming Xu, Ye Tian, Ying Sun, and Gesualdo Scutari · 2020
Later among the works it cites.
Optimal and practical algorithms for smooth and strongly convex decentralized optimization
Dmitry Kovalev, Adil Salim, and Peter Richtárik · 2020
Later among the works it cites.
Adaptive gradient quantization for data-parallel sgd
Fartash Faghri, Iman Tabrizian, Ilia Markov, Dan Alistarh, Daniel M Roy, and Ali Ramezani-Kebrya · 2020
Later among the works it cites.
On biased compression for distributed learning
Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Later among the works it cites.
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtarik · 2020
Later among the works it cites.
Linearly converging error compensated sgd
Eduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, and Peter Richtarik · 2020
Later among the works it cites.
Artemis: tight convergence guarantees for bidirectional compression in federated learning
Constantin Philippenko and Aymeric Dieuleveut · 2020
Later among the works it cites.
A unified analysis of stochastic gradient methods for nonconvex federated optimization
Zhize Li and Peter Richtárik · 2020
Later among the works it cites.
Federated learning with compression: Unified analysis and sharp guarantees
Farzin Haddadpour, Mohammad Mahdi Kamani, Aryan Mokhtari, and Mehrdad Mahdavi · 2020
Later among the works it cites.
Improved convergence rates for non-convex federated learning with compression
Rudrajit Das, Abolfazl Hashemi, Sujay Sanghavi, and Inderjit S Dhillon · 2020
Later among the works it cites.
Error compensated distributed sgd can be accelerated
Xun Qian, Peter Richtárik, and Tong Zhang · 2020
Later among the works it cites.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova, Tao Lin, Sebastian U Stich, and Martin Jaggi · 2020
Later among the works it cites.
Don’t use large mini-batches, use local SGD
Tao Lin, Sebastian Urban Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Later among the works it cites.
Minibatch vs local sgd for heterogeneous distributed learning
Blake Woodworth, Kumar Kshitij Patel, and Nathan Srebro · 2020
Later among the works it cites.
Federated accelerated stochastic gradient descent
Honglin Yuan and Tengyu Ma · 2020
Later among the works it cites.
Federated composite optimization
Honglin Yuan, Manzil Zaheer, and Sashank Reddi · 2020
Later among the works it cites.
Advances in asynchronous parallel and distributed optimization
Mahmoud Assran, Arda Aytekin, Hamid Reza Feyzmahdavian, Mikael Johansson, and Michael G Rabbat · 2020
Later among the works it cites.
A tight convergence analysis for stochastic gradient descent with delayed updates
Yossi Arjevani, Ohad Shamir, and Nathan Srebro · 2020
Later among the works it cites.
Local sgd: Unified theory and new efficient methods
Eduard Gorbunov, Filip Hanzely, and Peter Richtarik · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Nvidia data center deep learning product performance
NVIDIA · 2021
Closest in time.
Adaptive federated optimization
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan · 2021
Closest in time.
Nuqsgd: Provably communication-efficient data-parallel sgd via nonuniform quantization
Ali Ramezani-Kebrya, Fartash Faghri, Ilya Markov, Vitalii Aksenov, Dan Alistarh, and Daniel M Roy · 2021
Closest in time.
Marina: Faster non-convex distributed learning with compression
Eduard Gorbunov, Konstantin P. Burlachenko, Zhize Li, and Peter Richtarik · 2021
Closest in time.
A linearly convergent algorithm for decentralized optimization: Sending less bits for free!
Dmitry Kovalev, Anastasia Koloskova, Martin Jaggi, Peter Richtarik, and Sebastian Stich · 2021
Closest in time.