Fetching the paper…
Reading the bibliography…
Deep learning (DL) applications are increasingly being deployed on HPC systems, to leverage the massive parallelism and computing power of those systems for DL model training.
A. Krizhevsky, G. Hinton
2009
Earlier work this paper cites.
L. Bautista-Gomez, S. Tsuboi, D. Komatitsch, F. Cappello, N. Maruyama, and S. Matsuoka, “Fti: High performance fault tolerance interface for hybrid systems,” in
2011
Earlier work this paper cites.
I. Icke and J. C. Bongard, “Improving genetic programming based symbolic regression using deterministic machine learning,” in
2013
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in
2014
Earlier work this paper cites.
J. Schmidhuber, “Deep learning in neural networks: An overview,”
2014
Earlier work this paper cites.
A. Krizhevsky, “One weird trick for parallelizing convolutional neural networks,” 2014
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” 2014
2014
Earlier work this paper cites.
S. Tokui and K. Oono, “Chainer:a next-generation open source framework for deep learning,” 2015
2015
Earlier work this paper cites.
——, “Tensorflow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
I. Ilievski, T. Akhtar, J. Feng, and C. A. Shoemaker, “Efficient hyperparameter optimization for deep learning algorithms using deterministic rbf surrogates,” 2016
2016
Earlier work this paper cites.
M. Abadi and et al., “Tensorflow: A system for large-scale machine learning,” in
2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” 2016
2016
Cited alongside, same era.
V. Amatya, A. Vishnu, C. Siegel, and J. Daily, “What does fault tolerant deep learning need from mpi?” in
2017
Cited alongside, same era.
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup, “Reproducibility of benchmarked deep reinforcement learning tasks for continuous control,” 2017
2017
Cited alongside, same era.
W. Liu, Z. Wang, X. Liu, N. Zeng, Y. Liu, and F. E. Alsaadi, “A survey of deep neural network architectures and their applications,”
2017
Cited alongside, same era.
Pytorch.org, “Getting started with distributed data parallel,” 2017. [Online]. Available: https://pytorch.org/tutorials/intermediate/ddp_tutorial.html
2017
Cited alongside, same era.
C.-C. Chen, C.-L. Yang, and H.-Y. Cheng, “Efficient and robust parallel dnn training through model parallelism on multi-gpu platform,” 2018
2018
Later among the works it cites.
A. Qiao, B. Aragam, B. Zhang, and E. Xing, “Fault tolerance in iterative-convergent machine learning,” in
2019
Later among the works it cites.
B. Nicolae, A. Moody, E. Gonsiorowski, K. Mohror, and F. Cappello, “Veloc: Towards high performance adaptive asynchronous checkpointing at large scale,” in
2019
Later among the works it cites.
O. Beaumont, J. Herrmann, L. Eyraud-Dubois, J. Hermann, A. Joly, and A. Shilova, “Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory,” 2019
2019
Later among the works it cites.
S. Tokui, R. Okuta, T. Akiba, Y. Niitani, T. Ogawa, S. Saito, S. Suzuki, K. Uenishi, B. Vogel, and H. Y. Vincent, “Chainer: A deep learning framework for accelerating the research cycle,” 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Sergeev and M. D. Balso, “Horovod: fast and easy distributed deep learning in TensorFlow,”
2018
Cited alongside, same era.
O. E. Gundersen and S. Kjensmo, “State of the art: Reproducibility in artificial intelligence,” 2018
2018
Cited alongside, same era.
B. Reagen, U. Gupta, L. Pentecost, P. Whatmough, S. K. Lee, N. Mulholland, D. Brooks, and G. Wei, “Ares: A framework for quantifying the resilience of deep neural networks,” in
2018
Cited alongside, same era.
P. Nagarajan, G. Warnell, and P. Stone, “Deterministic implementations for reproducibility in deep reinforcement learning,” 2018
2018
Cited alongside, same era.
A. Harlap, D. Narayanan, A. Phanishayee, V. Seshadri, N. Devanur, G. Ganger, and P. Gibbons, “Pipedream: Fast and efficient pipeline parallel dnn training,” 2018
2018
Cited alongside, same era.
“ABCI.” [Online]. Available: https://abci.ai/
Cited in the paper.
“Summit.” [Online]. Available: https://www.olcf.ornl.gov/summit/
Cited in the paper.
2019
Later among the works it cites.
A. Paszke and et al., “Pytorch: An imperative style, high-performance deep learning library,” 2019
2019
Later among the works it cites.
T. H. Authors, “Horovod api,” 2019. [Online]. Available: https://horovod.readthedocs.io/en/latest/api.html
2019
Later among the works it cites.
B. Nicolae, J. Li, J. M. Wozniak, G. Bosilca, M. Dorier, and F. Cappello, “Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models,” in
2020
Closest in time.
N. Corporation, “Nvidia cudnn,” 2020. [Online]. Available: https://developer.nvidia.com/cudnn
2020
Closest in time.
N. Corporation, “Nvidia tesla v100,” 2020. [Online]. Available: https://www.nvidia.com/es-es/data-center/tesla-v100/
2020
Closest in time.