Fetching the paper…
Reading the bibliography…
Use of Deep Learning (DL) in commercial applications such as image classification, sentiment analysis and speech recognition is increasing.
D. P. Anderson, “Boinc: A system for public-resource computing and storage,” in Fifth IEEE/ACM international workshop on grid computing . IEEE, 2004, pp. 4–10
2004
Earlier work this paper cites.
D. Kondo, B. Javadi, P. Malecot, F. Cappello, and D. P. Anderson, “Cost-benefit analysis of cloud computing versus desktop grids,” in 2009 IEEE International Symposium on Parallel & Distributed Processing . IEEE, 2009, pp. 1–12
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on Knowledge and Data Engineering , vol. 22, no. 10, pp. 1345–1359, 2010
2010
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild!: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in Neural Information Processing Systems , J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Q. Weinberger, Eds., vol. 24. Curran Associates, Inc., 2011, pp. 693–701
2011
Earlier work this paper cites.
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng, “Large scale distributed deep networks,” in Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1 , ser. NIPS’12. Red Hook, NY, USA: Curran Associates Inc., 2012, pp. 1223–1231
2012
Earlier work this paper cites.
Q. Ho, J. Cipar, H. Cui, J. K. Kim, S. Lee, P. B. Gibbons, G. A. Gibson, G. R. Ganger, and E. P. Xing, “More effective distributed ml via a stale synchronous parallel parameter server,” Advances in neural information processing systems , pp. 1223–1231, 2013
2013
Earlier work this paper cites.
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman, “Project adam: Building an efficient and scalable deep learning training system,” in 11th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 14) , 2014, pp. 571–582
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14) . Broomfield, CO: USENIX Association, Oct. 2014, pp. 583–598
2014
Earlier work this paper cites.
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “Overfeat: Integrated recognition, localization and detection using convolutional networks,” CoRR , vol. 1312.6229, 2014
2014
Earlier work this paper cites.
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu, “Petuum: A new platform for distributed machine learning on big data,” IEEE Transactions on Big Data , vol. 1, no. 2, pp. 49–67, 2015
2015
Cited alongside, same era.
X. Lian, Y. Huang, Y. Li, and J. Liu, “Asynchronous parallel stochastic gradient for nonconvex optimization,” in Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015
2015
Cited alongside, same era.
S. Zhang, A. E. Choromanska, and Y. LeCun, “Deep learning with elastic averaging sgd,” in Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015, pp. 685–693
2015
Cited alongside, same era.
E. Kijsipongse, A. Piyatumrong, U. Suriya et al. , “A hybrid gpu cluster and volunteer computing platform for scalable deep learning,” The Journal of Supercomputing , vol. 74, no. 7, pp. 3236–3263, 2018
2018
Later among the works it cites.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” Siam Reviews , vol. 60, no. 2, pp. 223–311, 2018
2018
Later among the works it cites.
2019
Later among the works it cites.
J. J. Dai, Y. Wang, X. Qiu, D. Ding, Y. Zhang, Y. Wang, X. Jia, C. L. Zhang, Y. Wan, Z. Li et al. , “Bigdl: A distributed deep learning framework for big data,” in Proceedings of the ACM Symposium on Cloud Computing , 2019, pp. 50–60
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng, “Tensorflow: A system for large-scale machine learning,” in Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation . USA: USENIX Association, 2016, pp. 265–283
2016
Cited alongside, same era.
M. Hardt, B. Recht, and Y. Singer, “Train faster, generalize better: Stability of stochastic gradient descent,” in Proceedings of The 33rd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. F. Balcan and K. Q. Weinberger, Eds., vol. 48. New York, USA: PMLR, 20–22 Jun 2016, pp. 1225–1234
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 06 2016, pp. 770–778
2016
Cited alongside, same era.
T. Desell, “Developing a volunteer computing project to evolve convolutional neural networks and their hyperparameters,” in 2017 IEEE 13th International Conference on e-Science (e-Science) . IEEE, 2017, pp. 19–28
2017
Cited alongside, same era.
B. Bhattacharjee, S. Boag, C. Doshi, P. Dube, B. Herta, V. Ishakian, K. Jayaram, R. Khalaf, A. Krishna, Y. B. Li et al. , “Ibm deep learning service,” IBM Journal of Research and Development , vol. 61, no. 4/5, pp. 10:1–10:11, 2017
2017
Cited alongside, same era.
S. Zheng, Q. Meng, T. Wang, W. Chen, N. Yu, Z.-M. Ma, and T.-Y. Liu, “Asynchronous stochastic gradient descent with delay compensation,” in International Conference on Machine Learning . PMLR, 2017, pp. 4120–4129
2017
Cited alongside, same era.
2018
Cited alongside, same era.
J. Á. Morell, A. Camero, and E. Alba, “Jsdoop and tensorflow. js: Volunteer distributed web browser-based neural network training,” IEEE Access , vol. 7, pp. 158 671–158 684, 2019
2019
Later among the works it cites.
(2020, Feb.) Turing-nlg: A 17-billion-parameter language model by microsoft. [Online]. Available: https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/
2020
Later among the works it cites.
M. Ryabinin and A. Gusev, “Towards crowdsourced training of large neural networks using decentralized mixture-of-experts,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 3659–3672
2020
Later among the works it cites.
M. I. Zimmerman, J. R. Porter, M. D. Ward, S. Singh, N. Vithani, A. Meller, U. L. Mallimadugula, C. E. Kuhn, J. H. Borowsky, R. P. Wiewiora et al. , “Citizen scientists create an exascale computer to combat covid-19,” BioRxiv , 2020
2020
Later among the works it cites.