Fetching the paper…
Reading the bibliography…
The training process of Deep Neural Network (DNN) is compute-intensive, often taking days to weeks to train a DNN model.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M. Z., Ranzato, M., Senior, A. W., Tucker, P. A., Yang, K., and Ng, A. Y · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A.-r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Solving the straggler problem with bounded staleness
Cipar, J., Ho, Q., Kim, J. K., Lee, S., Ganger, G. R., Gibson, G., Keeton, K., and Xing, E. P · 2013
Earlier work this paper cites.
Fast training of convolutional networks through ffts
Mathieu, M., Henaff, M., and LeCun, Y · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Chilimbi, T. M., Suzue, Y., Apacible, J., and Kalyanaraman, K · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Large scale recurrent neural network on GPU
Li, B., Zhou, E., Huang, B., Duan, J., Wang, Y., Xu, N., Zhang, J., and Yang, H · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Lian, X., Huang, Y., Li, Y., and Liu, J · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Strom, N · 2015
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, W. J · 2017
Later among the works it cites.
Device placement optimization with reinforcement learning
Mirhoseini, A., Pham, H., Le, Q. V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., Norouzi, M., Bengio, S., and Dean, J · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Later among the works it cites.
DPS: A dsm-based parameter server for machine learning
Sun, C., Zhang, Y., Yu, W., Zhang, R., Bhuiyan, M. Z. A., and Li, J · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
QSGD: randomized quantization for communication-optimal stochastic gradient descent
Alistarh, D., Li, J., Tomioka, R., and Vojnovic, M · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
Chen, J., Monga, R., Bengio, S., and Józefowicz, R · 2016
Cited alongside, same era.
Firecaffe: Near-linear acceleration of deep neural network training on compute clusters
Iandola, F. N., Moskewicz, M. W., Ashraf, K., and Keutzer, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Fast algorithms for convolutional neural networks
Lavin, A. and Gray, S · 2016
Cited alongside, same era.
Theano-mpi: A theano-based distributed training framework
Ma, H., Mao, F., and Taylor, G. W · 2016
Cited alongside, same era.
Staleness-aware async-sgd for distributed deep learning
Zhang, W., Gupta, S., Lian, X., and Liu, J · 2016
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Wu, L., Zhu, Z., and E, W · 2017
Later among the works it cites.
Pipedream: Fast and efficient pipeline parallel DNN training
Harlap, A., Narayanan, D., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., and Gibbons, P. B · 2018
Closest in time.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2018
Closest in time.
Exploring hidden dimensions in parallelizing convolutional neural networks
Jia, Z., Lin, S., Qi, C. R., and Aiken, A · 2018
Closest in time.
MXNET-MPI: embedding MPI parallelism in parameter server task model for scaling deep learning
Mamidala, A. R., Kollias, G., Ward, C., and Artico, F · 2018
Closest in time.
Massively multilingual neural machine translation in the wild: Findings and challenges
Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G., Cherry, C., Macherey, W., Chen, Z., and Wu, Y · 2019
Closest in time.
Improving ml applications in shared computing environments
Harlap, A · 2019
Closest in time.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2019
Closest in time.
Regularized evolution for image classifier architecture search
Real, E., Aggarwal, A., Huang, Y., and Le, Q. V · 2019
Closest in time.