Fetching the paper…
Reading the bibliography…
In distributed machine learning, where agents collaboratively learn from diverse private data sets, there is a fundamental tension between consensus and optimality.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Distributed strongly convex optimization
Konstantinos I Tsianos and Michael G Rabbat · 2012
Earlier work this paper cites.
Internet of things (iot): A vision, architectural elements, and future directions
Jayavardhana Gubbi, Rajkumar Buyya, Slaven Marusic, and Marimuthu Palaniswami · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Earlier work this paper cites.
Stochastic proximal gradient descent with acceleration techniques
Atsushi Nitanda · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep learning for detecting robotic grasps
Ian Lenz, Honglak Lee, and Ashutosh Saxena · 2015
Earlier work this paper cites.
An early resource characterization of deep learning on wearables, smartphones and internet-of-things devices
Nicholas D Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, and Fahim Kawsar · 2015
Earlier work this paper cites.
Can deep learning revolutionize mobile sensing?
Nicholas D Lane and Petko Georgiev · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Cited alongside, same era.
Learning from data with heterogeneous noise using sgd
Shuang Song, Kamalika Chaudhuri, and Anand Sarwate · 2015
Cited alongside, same era.
Keras, 2015
François Chollet et al · 2015
Cited alongside, same era.
How to scale distributed deep learning?
Peter H Jin, Qiaochu Yuan, Forrest Iandola, and Kurt Keutzer · 2016
Cited alongside, same era.
Distributed deep learning using synchronous stochastic gradient descent
Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere, Karthikeyan Vaidynathan, Srinivas Sridharan, Dhiraj Kalamkar, Bharat Kaul, and Pradeep Dubey · 2016
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Later among the works it cites.
Balancing communication and computation in distributed optimization
Albert S Berahas, Raghu Bollapragada, Nitish Shirish Keskar, and Ermin Wei · 2017
Later among the works it cites.
Generalised gossip-based subgradient method for distributed optimisation
Zhanhong Jiang, Kushal Mukherjee, and Soumik Sarkar · 2017
Later among the works it cites.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2017
Later among the works it cites.
Accelerated distributed nesterov gradient descent
Guannan Qu and Na Li · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Cited alongside, same era.
Gossip training for deep learning
Michael Blot, David Picard, Matthieu Cord, and Nicolas Thome · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Collaborative deep learning in fixed topology networks
Zhanhong Jiang, Aditya Balu, Chinmay Hegde, and Soumik Sarkar · 2017
Cited alongside, same era.
Later among the works it cites.
Adaptive consensus admm for distributed optimization
Zheng Xu, Gavin Taylor, Hao Li, Mario Figueiredo, Xiaoming Yuan, and Tom Goldstein · 2017
Later among the works it cites.
Asynchronous stochastic gradient descent with delay compensation for distributed deep learning
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Later among the works it cites.
Projection-free distributed online learning in networks
Wenpeng Zhang, Peilin Zhao, Wenwu Zhu, Steven CH Hoi, and Tong Zhang · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.