Fetching the paper…
Reading the bibliography…
The performance of Deep-Learning (DL) computing frameworks rely on the performance of data ingestion and checkpointing.
J. D. McCalpin et al. , “Memory bandwidth and machine balance in current high performance computers,” IEEE computer society technical committee on computer architecture (TCCA) newsletter , vol. 2, no. 19–25, 1995
1995
Earlier work this paper cites.
R. Thakur, W. Gropp, and E. Lusk, “A case for using MPI’s derived datatypes to improve I/O performance,” in Supercomputing, 1998. SC98. IEEE/ACM Conference on . IEEE, 1998, pp. 1–1
1998
Earlier work this paper cites.
S. A. Weil, S. A. Brandt, E. L. Miller, D. D. Long, and C. Maltzahn, “Ceph: A scalable, high-performance distributed file system,” in Proceedings of the 7th symposium on Operating systems design and implementation . USENIX Association, 2006, pp. 307–320
2006
Earlier work this paper cites.
H. Shan and J. Shalf, “Using IOR to Analyze the I/O Performance for HPC Platforms,” in Cray User Group Conference (CUG07) , 2007
2007
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
K. Sato, N. Maruyama, K. Mohror, A. Moody, T. Gamblin, B. R. de Supinski, and S. Matsuoka, “Design and modeling of a non-blocking checkpointing system,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis . IEEE Computer Society Press, 2012, p. 19
2012
Earlier work this paper cites.
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and N. Andrew, “Deep learning with COTS HPC systems,” in Proceedings of the 30th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, S. Dasgupta and D. McAllester, Eds., vol. 28, no. 3. Atlanta, Georgia, USA: PMLR, 17–19 Jun 2013, pp. 1337–1345. [Online]. Available: http://proceedings.mlr.press/v28/coates13.html
2013
Earlier work this paper cites.
M. Snir, R. W. Wisniewski, J. A. Abraham, S. V. Adve, S. Bagchi, P. Balaji, J. Belak, P. Bose, F. Cappello, B. Carlson, A. A. Chien, P. Coteus, N. A. DeBardeleben, P. C. Diniz, C. Engelmann, M. Erez, S. Fazzari, A. Geist, R. Gupta, F. Johnson, S. Krishnamoorthy, S. Leyffer, D. Liberty, S. Mitra, T. Munson, R. Schreiber, J. Stearley, and E. V. Hensbergen, “Addressing failures in exascale computing,” The International Journal of High Performance Computing Applications , vol. 28, no. 2, pp. 129–173, 2014. [Online]. Available: https://doi.org/10.1177/1094342014522573
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the 22nd ACM international conference on Multimedia . ACM, 2014, pp. 675–678
2014
Earlier work this paper cites.
M. Dorier, G. Antoniu, R. Ross, D. Kimpe, and S. Ibrahim, “CALCioM: Mitigating I/O Interference in HPC Systems through Cross-Application Coordination,” in 2014 IEEE 28th International Parallel and Distributed Processing Symposium , May 2014, pp. 155–164
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision (IJCV) , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: a system for large-scale machine learning,” in OSDI , vol. 16, 2016, pp. 265–283
2016
Cited alongside, same era.
I. B. Peng, S. Markidis, E. Laure, G. Kestor, and R. Gioiosa, “Exploring application performance on emerging hybrid-memory supercomputers,” in 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC/SmartCity/DSS) , Dec 2016, pp. 473–480
2016
Cited alongside, same era.
W. Bhimji, D. Bard, M. Romanus, D. Paul, A. Ovsyannikov, B. Friesen, M. Bryson, J. Correa, G. K. Lockwood, V. Tsulaia et al. , “Accelerating science with the nersc burst buffer early user program,” CUG16 , 2016
D. Harnie, M. Saey, A. E. Vapirev, J. K. Wegner, A. Gedich, M. Steijaert, H. Ceulemans, R. Wuyts, and W. D. Meuter, “Scaling machine learning for target prediction in drug discovery using Apache Spark,” Future Generation Computer Systems , vol. 67, pp. 409 – 417, 2017. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0167739X1630111X
2017
Later among the works it cites.
S. Pumma, M. Si, W. Feng, and P. Balaji, “Towards Scalable Deep Learning via I/O Analysis and Optimization,” in 2017 IEEE 19th International Conference on High Performance Computing and Communications; IEEE 15th International Conference on Smart City; IEEE 3rd International Conference on Data Science and Systems (HPCC/SmartCity/DSS) , Dec 2017, pp. 223–230
2017
Later among the works it cites.
S. Pumma, M. Si, W.-c. Feng, and P. Balaji, “Parallel I/O Optimizations for Scalable Deep Learning,” in Parallel and Distributed Systems (ICPADS), 2017 IEEE 23rd International Conference on . IEEE, 2017, pp. 720–729
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
D. Henseler, B. Landsteiner, D. Petesch, C. Wright, and N. J. Wright, “Architecture and design of Cray DataWarp,” Cray User Group CUG , 2016
2016
Cited alongside, same era.
F. Seide and A. Agarwal, “CNTK: Microsoft’s open-source deep-learning toolkit,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . ACM, 2016, pp. 2135–2135
2016
Cited alongside, same era.
S. Shi, Q. Wang, P. Xu, and X. Chu, “Benchmarking State-of-the-Art Deep Learning Software Tools,” in 2016 7th International Conference on Cloud Computing and Big Data (CCBD) , Nov 2016, pp. 99–104
2016
Cited alongside, same era.
M. Zaharia, R. S. Xin, P. Wendell, T. Das, M. Armbrust, A. Dave, X. Meng, J. Rosen, S. Venkataraman, M. J. Franklin, A. Ghodsi, J. Gonzalez, S. Shenker, and I. Stoica, “Apache Spark: A unified engine for big data processing,” Commun. ACM , vol. 59, no. 11, pp. 56–65, Oct. 2016. [Online]. Available: http://doi.acm.org/10.1145/2934664
2016
Cited alongside, same era.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers et al. , “In-datacenter performance analysis of a tensor processing unit,” in Computer Architecture (ISCA), 2017 ACM/IEEE 44th Annual International Symposium on . IEEE, 2017, pp. 1–12
2017
Cited alongside, same era.
W. Schenck, S. El Sayed, M. Foszczynski, W. Homberg, and D. Pleiter, “Evaluation and performance modeling of a burst buffer solution,” ACM SIGOPS Operating Systems Review , vol. 50, no. 2, pp. 12–26, 2017
2017
Cited alongside, same era.
H. Kim, H. Nam, W. Jung, and J. Lee, “Performance analysis of CNN frameworks for GPUs,” in 2017 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , April 2017, pp. 55–64
2017
Cited alongside, same era.
A. A. Awan, K. Hamidouche, J. M. Hashmi, and D. K. Panda, “S-Caffe: co-designing MPI runtimes and Caffe for scalable deep learning on modern GPU clusters,” Acm Sigplan Notices , vol. 52, no. 8, pp. 193–205, 2017
2017
Later among the works it cites.
H. Ma, F. Mao, and G. W. Taylor, “Theano-MPI: A theano-based distributed training framework,” in Euro-Par 2016: Parallel Processing Workshops , F. Desprez, P.-F. Dutot, C. Kaklamanis, L. Marchal, K. Molitorisz, L. Ricci, V. Scarano, M. A. Vega-Rodríguez, A. L. Varbanescu, S. Hunold, S. L. Scott, S. Lankes, and J. Weidendorfer, Eds. Cham: Springer International Publishing, 2017, pp. 800–813
2017
Later among the works it cites.
S. Markidis, S. W. D. Chien, E. Laure, I. B. Peng, and J. S. Vetter, “NVIDIA Tensor Core Programmability, Performance & Precision,” in 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) , May 2018, pp. 522–531
2018
Closest in time.
2018
Closest in time.
Y. Oyama, T. Ben-Nun, T. Hoefler, and S. Matsuoka, “Accelerating Deep Learning Frameworks with Micro-batches.” IEEE, Sep. 2018, to appear in IEEE International Conference on Cluster Computing (Cluster’18)
2018
Closest in time.
Y. Zhu, F. Chowdhury, H. Fu, A. Moody, K. Mohror, K. Sato, and W. Yu, “Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC Systems,” in IEEE International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS 2018) , 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.