Fetching the paper…
Reading the bibliography…
Training machine learning models in parallel is an increasingly important workload.
Distributed Learning with Compressed Gradient Differences
K. Mishchenko, E. Gorbunov, M. Takáč, and P. Richtárik · 1901
Earlier work this paper cites.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovski and D. B. Yudin · 1983
Earlier work this paper cites.
Interprocessor collective communication library (InterCom)
M. Barnett, L. Shuler, R. van de Geijn, S. Gupta, D. G. Payne, and J. Watts · 1994
Earlier work this paper cites.
Optimization of Collective Communication Operations in MPICH
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
MPI Collective Communications on The Blue Gene/P Supercomputer: Algorithms and Optimizations
A. Faraj, S. Kumar, B. Smith, A. Mamidala, and J. Gunnels · 2009
Earlier work this paper cites.
Robust Stochastic Approximation Approach to Stochastic Programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Bandwidth Optimal All-reduce Algorithms for Clusters of Workstations
P. Patarasuk and X. Yuan · 2009
Earlier work this paper cites.
Camdoop: Exploiting In-network Aggregation for Big Data Applications
P. Costa, A. Donnelly, A. Rowstron, and G. O’Shea · 2012
Earlier work this paper cites.
Large Scale Distributed Deep Networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng · 2012
Earlier work this paper cites.
Forwarding Metamorphosis: Fast Programmable Match-action Processing in Hardware for SDN
P. Bosshart, G. Gibb, H.-S. Kim, G. Varghese, N. McKeown, M. Izzard, F. Mujica, and M. Horowitz · 2013
Earlier work this paper cites.
One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Earlier work this paper cites.
P4: Programming Protocol-independent Packet Processors
P. Bosshart, D. Daly, G. Gibb, M. Izzard, N. McKeown, J. Rexford, C. Schlesinger, D. Talayco, A. Vahdat, G. Varghese, and D. Walker · 2014
Earlier work this paper cites.
Project Adam: Building an Efficient and Scalable Deep Learning Training System
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Earlier work this paper cites.
Fast distributed coordinate descent for minimizing non-strongly convex losses
O. Fercoq, Z. Qu, P. Richtárik, and M. Takáč · 2014
Earlier work this paper cites.
Scaling Distributed Machine Learning with the Parameter Server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Doll’a r, and C. L. Zitnick · 2014
Earlier work this paper cites.
NetAgg: Using Middleboxes for Application-Specific On-path Aggregation in Data Centres
L. Mai, L. Rupprecht, A. Alim, P. Costa, M. Migliavacca, P. Pietzuch, and A. L. Wolf · 2014
Earlier work this paper cites.
1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
NetPaxos: Consensus at Network Speed
H. T. Dang, D. Sciascia, M. Canini, F. Pedone, and R. Soulé · 2015
Earlier work this paper cites.
The MovieLens Datasets: History and Context
F. M. Harper and J. A. Konstan · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Managed Communication and Consistency for Fast Data-Parallel Iterative Analytics
J. Wei, W. Dai, A. Qiao, Q. Ho, H. Cui, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Earlier work this paper cites.
MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2016
Cited alongside, same era.
Paxos Made Switch-y
H. T. Dang, M. Canini, F. Pedone, and R. Soulé · 2016
Cited alongside, same era.
Scalable Hierarchical Aggregation Protocol (SHArP): A Hardware Architecture for Efficient Data Reduction
R. L. Graham, D. Bureddy, P. Lui, H. Rosenstock, G. Shainer, G. Bloch, D. Goldenerg, M. Dubman, S. Kotchubievsky, V. Koushnir, L. Levi, A. Margolin, T. Ronen, A. Shpiner, O. Wertheim, and E. Zahavi · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
FireCaffe: Near-Linear Acceleration of Deep Neural Network Training on Compute Clusters
F. N. Iandola, M. W. Moskewicz, K. Ashraf, and K. Keutzer · 2016
Cited alongside, same era.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.
Deep Learning on V100
R. Xu, F. Han, and N. Dandapanthula · 2017
Later among the works it cites.
signSGD: Compressed Optimisation for Non-Convex Problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Later among the works it cites.
signSGD with Majority Vote is Communication Efficient And Byzantine Fault Tolerant
J. Bernstein, J. Zhao, K. Azizzadenesheli, and A. Anandkumar · 2018
Later among the works it cites.
Volta: Performance and Programmability
J. Choquette, O. Giroux, and D. Foley · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Cited alongside, same era.
SSD: Single Shot MultiBox Detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Cited alongside, same era.
Distributed Coordinate Descent Method for Learning with Big Data
P. Richtárik and M. Takáč · 2016
Cited alongside, same era.
CNTK: Microsoft’s Open-Source Deep-Learning Toolkit
F. Seide and A. Agarwal · 2016
Cited alongside, same era.
Packet Transactions: High-Level Programming for Line-Rate Switches
A. Sivaraman, A. Cheung, M. Budiu, C. Kim, M. Alizadeh, H. Balakrishnan, G. Varghese, N. McKeown, and S. Licking · 2016
Cited alongside, same era.
DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
S. Zhou, Z. Ni, X. Zhou, H. Wen, Y. Wu, and Y. Zou · 2016
Cited alongside, same era.
Sparse Communication for Distributed Gradient Descent
A. F. Aji and K. Heafield · 2017
Cited alongside, same era.
Later among the works it cites.
Training DNNs with Hybrid Block Floating Point
M. Drumond, T. Lin, M. Jaggi, and B. Falsafi · 2018
Later among the works it cites.
Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang · 2018
Later among the works it cites.
NetChain: Scale-Free Sub-RTT Coordination
X. Jin, X. Li, H. Zhang, N. Foster, J. Lee, R. Soulé, C. Kim, and I. Stoica · 2018
Later among the works it cites.
A Network-Centric Hardware/Algorithm Co-Design to Accelerate Distributed Training of Deep Neural Networks
Y. Li, J. Park, M. Alian, Y. Yuan, Z. Qu, P. Pan, R. Wang, A. Gerhard Schwing, H. Esmaeilzadeh, and N. Sung Kim · 2018
Later among the works it cites.
Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
Y. Lin, S. Han, H. Mao, Y. Wang, and B. Dally · 2018
Later among the works it cites.
PHub: Rack-Scale Parameter Server for Distributed Deep Neural Network Training
L. Luo, J. Nelson, L. Ceze, A. Phanishayee, and A. Krishnamurthy · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
D. Masters and C. Luschi · 2018
Later among the works it cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in TensorFlow
A. Sergeev and M. D. Balso · 2018
Later among the works it cites.
Speeding up ImageNet Training on Supercomputers
Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer · 2018
Later among the works it cites.
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
M. Jeon, S. Venkataraman, A. Phanishayee, J. Qian, W. Xiao, and F. Yang · 2019
Closest in time.
Beyond Data and Model Parallelism for Deep Neural Networks
Z. Jia, M. Zaharia, and A. Aiken · 2019
Closest in time.
Accelerating Distributed Reinforcement Learning with In-Switch Computing
Y. Li, I.-J. Liu, Y. Yuan, D. Chen, A. Schwing, and J. Huang · 2019
Closest in time.
When Should The Network Be The Computer?
D. R. K. Ports and J. Nelson · 2019
Closest in time.
DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving
W. Deng, J. Pan, T. Zhou, D. Kong, A. Flores, and G. Lin · 2020
Closest in time.
PANAMA: Network Architecture for Machine Learning Workloads in the Cloud
N. Gebara, T. Ukyab, P. Costa, and M. Ghobadi · 2020
Closest in time.
U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image Translation
J. Kim, M. Kim, H. Kang, and K. H. Lee · 2020
Closest in time.
NetReduce: RDMA-Compatible In-Network Reduction for Distributed DNN Training Acceleration
S. Liu, Q. Wang, J. Zhang, Q. Lin, Y. Liu, M. Xu, R. C. Chueng, and J. He · 2020
Closest in time.
Compressed Communication for Distributed Deep Learning: Survey and Quantitative Evaluation
H. Xu, C.-Y. Ho, A. M. Abdelmoniem, A. Dutta, E. H. Bergou, K. Karatsenidis, M. Canini, and P. Kalnis · 2020
Closest in time.