Fetching the paper…
Reading the bibliography…
This paper presents the design, implementation, and evaluation of the PyTorch distributed data parallel module.
A theoretical framework for back-propagation
Y. LeCun, D. Touresky, G. Hinton, and T. Sejnowski · 1988
Earlier work this paper cites.
The MNIST Database
Y. LeCun, C. Cortes, and C. Burges · 1999
Earlier work this paper cites.
Deep content-based music recommendation
A. Van den Oord, S. Dieleman, and B. Schrauwen · 2013
Earlier work this paper cites.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Deeplog: Anomaly detection and diagnosis from system logs through deep learning
M. Du, F. Li, G. Zheng, and V. Srikumar · 2017
Earlier work this paper cites.
Deepart: Learning joint representations of visual arts
H. Mao, M. Cheung, and J. She · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Tictac: Accelerating distributed deep learning with communication scheduling
S. H. Hashemi, S. A. Jyothi, and R. H. Campbell · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in TensorFlow
A. Sergeev and M. D. Balso · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young, et al · 2018
Earlier work this paper cites.
https://github.com/facebookincubator/gloo , 2019
Gloo: a collective communications library · 2019
Cited alongside, same era.
https://developer.nvidia.com/nccl , 2019
NVIDIA Collective Communications Library (NCCL) · 2019
Cited alongside, same era.
https://www.nvidia.com/en-us/data-center/nvlink/ , 2019
NVLINK AND NVSWITCH: The Building Blocks of Advanced Multi-GPU Communication · 2019
Cited alongside, same era.
https://www.open-mpi.org/ , 2019
Open MPI: A High Performance Message Passing Library · 2019
Cited alongside, same era.
https://pybind11.readthedocs.io/ , 2019
Pybind11: Seamless operability between C++11 and Python · 2019
Cited alongside, same era.
https://pytorch.org/docs/master/rpc.html , 2019
PyTorch Distributed RPC Framework · 2019
Cited alongside, same era.
Parity models: Erasure-coded resilience for prediction serving systems
J. Kosaian, K. V. Rashmi, and S. Venkataraman · 2019
Later among the works it cites.
Pipedream: generalized pipeline parallelism for dnn training
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Later among the works it cites.
A generic communication scheduler for distributed dnn training acceleration
Y. Peng, Y. Zhu, Y. Chen, Y. Bao, B. Yi, C. Lan, C. Wu, and C. Guo · 2019
Later among the works it cites.
Zero: Memory optimization towards training a trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
https://pytorch.org/docs/stable/nn.html#torch.nn.Module.forward , 2019
PyTorch Module forward · 2019
Cited alongside, same era.
https://docs.scipy.org/ , 2019
SciPy: open-source software for mathematics, science, and engineering · 2019
Cited alongside, same era.
Blueconnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy
M. Cho, U. Finkler, M. Serrano, D. Kung, and H. Hunter · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave, and A. Joulin · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al · 2019
Cited alongside, same era.
Massively Scale Your Deep Learning Training with NCCL 2.4
S. Jeaugey · 2019
Cited alongside, same era.
Deep learning for the life sciences: applying deep learning to genomics, microscopy, drug discovery, and more
B. Ramsundar, P. Eastman, P. Walters, and V. Pande · 2019
Later among the works it cites.
Gradientflow: Optimizing network performance for large-scale distributed dnn training
P. Sun, Y. Wen, R. Han, W. Feng, and S. Yan · 2019
Later among the works it cites.
Blink: Fast and generic collectives for distributed ml
G. Wang, S. Venkataraman, A. Phanishayee, J. Thelin, N. Devanur, and I. Stoica · 2019
Later among the works it cites.
Slowmo: Improving communication-efficient distributed sgd with slow momentum
J. Wang, V. Tantia, N. Ballas, and M. Rabbat · 2019
Later among the works it cites.
https://pytorch.org/docs/stable/nn.html#torch.nn.parallel.DistributedDataParallel , 2020
PyTorch DistributedDataParallel · 2020
Closest in time.
https://www.tensorflow.org/guide/distributed_training#multiworkermirroredstrategy , 2020
TensorFlow Distributed Training MultiWorkerMirroredStrategy · 2020
Closest in time.
https://www.tensorflow.org/guide/distributed_training#parameterserverstrategy , 2020
TensorFlow Distributed Training ParameterServerStrategy · 2020
Closest in time.
Preemptive all-reduce scheduling for expediting distributed dnn training
Y. Bao, Y. Peng, Y. Chen, and C. Wu · 2020
Closest in time.