Fetching the paper…
Reading the bibliography…
Deep Learning (DL) has had an immense success in the recent past, leading to state-of-the-art results in various domains such as image recognition and natural language processing.
Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex
Hubel, D. H., and Wiesel, T. N · 1962
Earlier work this paper cites.
Minds, brains, and programs
Searle, J. R · 1980
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
A bridging model for parallel computation
Valiant, L. G · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E., and Salakhutdinov, R. R · 2006
Earlier work this paper cites.
Mapreduce: Simplified data processing on large clusters
Dean, J., and Ghemawat, S · 2008
Earlier work this paper cites.
The Datacenter As a Computer: An Introduction to the Design of Warehouse-Scale Machines
Barroso, L. A., and Hoelzle, U · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
The unreasonable effectiveness of data
Halevy, A., Norvig, P., and Pereira, F · 2009
Earlier work this paper cites.
A view of cloud computing
Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R., Konwinski, A., Lee, G., Patterson, D., Rabkin, A., Stoica, I., and Zaharia, M · 2010
Earlier work this paper cites.
Deep, big, simple neural nets for handwritten digit recognition
Cireşan, D. C., Meier, U., Gambardella, L. M., and Schmidhuber, J · 2010
Earlier work this paper cites.
Weka-A Machine Learning Workbench for Data Mining
Frank, E., Hall, M., Holmes, G., Kirkby, R., Pfahringer, B., Witten, I. H., and Trigg, L · 2010
Earlier work this paper cites.
Debunking the 100x gpu vs. cpu myth: An evaluation of throughput computing on cpu and gpu
Lee, V. W., Kim, C., Chhugani, J., Deisher, M., Kim, D., Nguyen, A. D., Satish, N., Smelyanskiy, M., Chennupaty, S., Hammarlund, P., Singhal, R., and Dubey, P · 2010
Earlier work this paper cites.
Pregel: A system for large-scale graph processing
Malewicz, G., Austern, M. H., Bik, A. J., Dehnert, J. C., Horn, I., Leiser, N., and Czajkowski, G · 2010
Earlier work this paper cites.
The quest for artificial intelligence: A history of ideas and achievements
Nilsson, N · 2010
Earlier work this paper cites.
An architecture for parallel topic models
Smola, A., and Narayanamurthy, S · 2010
Earlier work this paper cites.
Theano: Deep learning on gpus with python
Bergstra, J., Bastien, F., Breuleux, O., Lamblin, P., Pascanu, R., Delalleau, O., Desjardins, G., Warde-Farley, D., Goodfellow, I., Bergeron, A., et al · 2011
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C., and Lin, C.-J · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Neuflow: A runtime reconfigurable dataflow processor for vision
Farabet, C., Martini, B., Corda, B., Akselrod, P., Culurciello, E., and LeCun, Y · 2011
Earlier work this paper cites.
Dominant resource fairness: Fair allocation of multiple resource types
Ghodsi, A., Zaharia, M., Hindman, B., Konwinski, A., Shenker, S., and Stoica, I · 2011
Earlier work this paper cites.
Mesos: A platform for fine-grained resource sharing in the data center
Hindman, B., Konwinski, A., Zaharia, M., Ghodsi, A., Joseph, A. D., Katz, R., Shenker, S., and Stoica, I · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F · 2011
Earlier work this paper cites.
Improving the speed of neural networks on cpus
Vanhoucke, V., Senior, A., and Mao, M. Z · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J., and Bengio, Y · 2012
Earlier work this paper cites.
Multi-column deep neural network for traffic sign classification
CireşAn, D., Meier, U., Masci, J., and Schmidhuber, J · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., and Kingsbury, B · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Learning to label aerial images from noisy data
Mnih, V., and Hinton, G. E · 2012
Earlier work this paper cites.
Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing
Zaharia, M., Chowdhury, M., Das, T., Dave, A., Ma, J., McCauly, M., Franklin, M. J., Shenker, S., and Stoica, I · 2012
Earlier work this paper cites.
Solving the straggler problem with bounded staleness
Cipar, J., Ho, Q., Kim, J. K., Lee, S., Ganger, G. R., Gibson, G., Keeton, K., and Xing, E · 2013
Earlier work this paper cites.
Deep learning with cots hpc systems
Coates, A., Huval, B., Wang, T., Wu, D. J., Ng, A. Y., and Catanzaro, B · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R., and Zhang, T · 2013
Earlier work this paper cites.
Reasoning with neural tensor networks for knowledge base completion
Socher, R., Chen, D., Manning, C. D., and Ng, A · 2013
Earlier work this paper cites.
Apache hadoop yarn: Yet another resource negotiator
Vavilapalli, V. K., Murthy, A. C., Douglas, C., Agarwal, S., Konar, M., Evans, R., Graves, T., Lowe, J., Shah, H., Seth, S., Saha, B., Curino, C., O’Malley, O., Radia, S., Reed, B., and Baldeschwieler, E · 2013
Earlier work this paper cites.
Butterfly mixing: Accelerating incremental-update algorithms on clusters
Zhao, H., and Canny, J · 2013
Earlier work this paper cites.
A reliable effective terascale linear learning system
Agarwal, A., Chapelle, O., Dudík, M., and Langford, J · 2014
Earlier work this paper cites.
Big data deep learning: Challenges and perspectives
Chen, X., and Lin, X · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Chilimbi, T., Suzue, Y., Apacible, J., and Kalyanaraman, K · 2014
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
Cui, H., Cipar, J., Ho, Q., Kim, J. K., Lee, S., Kumar, A., Wei, J., Dai, W., Ganger, G. R., Gibbons, P. B., Gibson, G. A., and Xing, E. P · 2014
Earlier work this paper cites.
A tutorial survey of architectures, algorithms, and applications for deep learning
Deng, L · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Multi-resource packing for cluster schedulers
Grandl, R., Ananthanarayanan, G., Kandula, S., Rao, S., and Akella, A · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Graves, A., and Jaitly, N · 2014
Earlier work this paper cites.
A historical perspective of speech recognition
Huang, X., Baker, J., and Reddy, R · 2014
Earlier work this paper cites.
An efficient approach for assessing hyperparameter importance
Hutter, F., Hoos, H., and Leyton-Brown, K · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R. B., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Dogwild!-distributed hogwild for cpu & gpu
Noel, C., and Osindero, S · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Learning from noisy labels with deep neural networks
Sukhbaatar, S., and Fergus, R · 2014
Earlier work this paper cites.
Minerva: A scalable and highly efficient training platform for deep learning
Wang, M., Xiao, T., Li, J., Zhang, J., Hong, C., and Zhang, Z · 2014
Earlier work this paper cites.
A scalable and topology configurable protocol for distributed parameter synchronization
Wang, M., Zhou, H., Guo, M., and Zhang, Z · 2014
Earlier work this paper cites.
Mariana: Tencent deep learning platform and its applications
Zou, Y., Jin, X., Li, Y., Guo, Z., Wang, E., and Xiao, B · 2014
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2015
Earlier work this paper cites.
High-performance distributed ml at scale through parameter server consistency models
Dai, W., Kumar, A., Wei, J., Ho, Q., Gibson, G., and Xing, E. P · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Djinn and tonic: Dnn as a service and its implications for future warehouse scale computers
Hauswald, J., Kang, Y., Laurenzano, M. A., Chen, Q., Li, C., Mudge, T., Dreslinski, R. G., Mars, J., and Tang, L · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Deep learning for detecting robotic grasps
Lenz, I., Lee, H., and Saxena, A · 2015
Earlier work this paper cites.
Malt: Distributed data-parallelism for existing ml applications
Li, H., Kadav, A., Kruus, E., and Ungureanu, C · 2015
Earlier work this paper cites.
Optimizing network performance in distributed machine learning
Mai, L., Hong, C., and Costa, P · 2015
Earlier work this paper cites.
Singa: A distributed deep learning platform
Ooi, B. C., Tan, K.-L., Wang, S., Wang, W., Cai, Q., Chen, G., Gao, J., Luo, Z., Tung, A. K., Wang, Y., Xie, Z., Zhang, M., and Zheng, K · 2015
Earlier work this paper cites.
Accelerating deep convolutional neural networks using specialized hardware, February 2015
Ovtcharov, K., Ruwase, O., Kim, J.-Y., Fowers, J., Strauss, K., and Chung, E · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J · 2015
Cited alongside, same era.
Privacy-preserving deep learning
Shokri, R., and Shmatikov, V · 2015
Cited alongside, same era.
Automating model search for large scale machine learning
Sparks, E. R., Talwalkar, A., Haas, D., Franklin, M. J., Jordan, M. I., and Kraska, T · 2015
Cited alongside, same era.
Scalable distributed dnn training using commodity gpu cloud computing
Strom, N · 2015
Cited alongside, same era.
Chainer: a next-generation open source framework for deep learning
Tokui, S., Oono, K., Hido, S., and Clayton, J · 2015
Cited alongside, same era.
Large-scale cluster management at google with borg
Verma, A., Pedrosa, L., Korupolu, M., Oppenheimer, D., Tune, E., and Wilkes, J · 2015
Cited alongside, same era.
Neuromorphic computing with multi-memristive synapses
Boybat, I., Le Gallo, M., Nandakumar, S., Moraitis, T., Parnell, T., Tuma, T., Rajendran, B., Leblebici, Y., Sebastian, A., and Eleftheriou, E · 2018
Later among the works it cites.
The streaming rollout of deep networks - towards fully model-parallel execution
Fischer, V., Koehler, J., and Pfeil, T · 2018
Later among the works it cites.
Bringing HPC Techniques to Deep Learning
Gibiansky, A · 2018
Later among the works it cites.
Pipedream: Fast and efficient pipeline parallel DNN training
Harlap, A., Narayanan, D., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., and Gibbons, P. B · 2018
Later among the works it cites.
Pipedream: Pipeline parallelism for dnn training
Harlap, A., Narayanan, D., Phanishayee, A., Seshadri, V., Ganger, G. R., and Gibbons, P. B · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Managed communication and consistency for fast data-parallel iterative analytics
Wei, J., Dai, W., Qiao, A., Ho, Q., Cui, H., Ganger, G. R., Gibbons, P. B., Gibson, G. A., and Xing, E. P · 2015
Cited alongside, same era.
Reef: Retainable evaluator execution framework
Weimer, M., Chen, Y., Chun, B.-G., Condie, T., Curino, C., Douglas, C., Lee, Y., Majestro, T., Malkhi, D., Matusevych, S., Myers, B., Narayanamurthy, S., Ramakrishnan, R., Rao, S., Sears, R., Sezgin, B., and Wang, J · 2015
Cited alongside, same era.
Learning from massive noisy labeled data for image classification
Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
Xing, E. P., Ho, Q., Dai, W., Kim, J. K., Wei, J., Lee, S., Zheng, X., Xie, P., Kumar, A., and Yu, Y · 2015
Cited alongside, same era.
Performance modeling and scalability optimization of distributed deep learning systems
Yan, F., Ruwase, O., He, Y., and Chilimbi, T · 2015
Cited alongside, same era.
Optimizing fpga-based accelerator design for deep convolutional neural networks
Zhang, C., Li, P., Sun, G., Guan, Y., Xiao, B., and Cong, J · 2015
Cited alongside, same era.
Hazelwood, K., Bird, S., Brooks, D., Chintala, S., Diril, U., Dzhulgakov, D., Fawzy, M., Jia, B., Jia, Y., Kalro, A., Law, J., Lee, K., Lu, J., Noordhuis, P., Smelyanskiy, M., Xiong, L., and Wang, X · 2018
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2018
Later among the works it cites.
Flexps: Flexible parallelism control in parameter server architecture
Huang, Y., Jin, T., Wu, Y., Cai, Z., Yan, X., Yang, F., Li, J., Guo, Y., and Cheng, J · 2018
Later among the works it cites.
Serving deep learning models in a serverless platform
Ishakian, V., Muthusamy, V., and Slominski, A · 2018
Later among the works it cites.
Parallelized training of deep nn: Comparison of current concepts and frameworks
Jäger, S., Zorn, H.-P., Igel, S., and Zirpins, C · 2018
Later among the works it cites.
Improving the expressiveness of deep learning frameworks with recursion
Jeong, E., Jeong, J. S., Kim, S., Yu, G.-I., and Chun, B.-G · 2018
Later among the works it cites.
Exploring hidden dimensions in accelerating convolutional neural networks
Jia, Z., Lin, S., Qi, C. R., and Aiken, A · 2018
Later among the works it cites.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2018
Later among the works it cites.
Ease.ml: Towards multi-tenant resource sharing for machine learning workloads
Li, T., Zhong, J., Liu, J., Wu, W., and Zhang, C · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B · 2018
Later among the works it cites.
Usability study of distributed deep learning frameworks for convolutional neural networks
Liu, J., Dutta, J., Li, N., Kurup, U., and Shah, M · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Masters, D., and Luschi, C · 2018
Later among the works it cites.
A hierarchical model for device placement
Mirhoseini, A., Goldie, A., Pham, H., Steiner, B., Le, Q. V., and Dean, J · 2018
Later among the works it cites.
Ray: A distributed framework for emerging AI applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., and Stoica, I · 2018
Later among the works it cites.
Accelerating deep learning workloads through efficient multi-model execution
Narayanan, D., Santhanam, K., Phanishayee, A., and Zaharia, M · 2018
Later among the works it cites.
https://github.com/NervanaSystems/neon
Neon · 2018
Later among the works it cites.
A performance evaluation of federated learning algorithms
Nilsson, A., Smith, S., Ulm, G., Gustavsson, E., and Jirstrand, M · 2018
Later among the works it cites.
Object storage for deep learning frameworks
Ozeri, O., Ofer, E., and Kat, R · 2018
Later among the works it cites.
Optimus: An efficient dynamic resource scheduler for deep learning clusters
Peng, Y., Bao, Y., Chen, Y., Wu, C., and Guo, C · 2018
Later among the works it cites.
Hoard: A distributed data caching system to accelerate deep learning training on the cloud
Pinto, C., Gkoufas, Y., Reale, A., Seelam, S., and Eliuk, S · 2018
Later among the works it cites.
A survey on deep learning: Algorithms, techniques, and applications
Pouyanfar, S., Sadiq, S., Yan, Y., Tian, H., Tao, Y., Reyes, M. P., Shyu, M.-L., Chen, S.-C., and Iyengar, S. S · 2018
Later among the works it cites.
Litz: Elastic framework for high-performance distributed machine learning
Qiao, A., Aghayev, A., Yu, W., Chen, H., Ho, Q., Gibson, G. A., and Xing, E. P · 2018
Later among the works it cites.
Numa-caffe: Numa-aware deep learning neural networks
Roy, P., Song, S. L., Krishnamoorthy, S., Vishnu, A., Sengupta, D., and Liu, X · 2018
Later among the works it cites.
A generic framework for privacy preserving deep learning
Ryffel, T., Trask, A., Dahl, M., Wagner, B., Mancuso, J., Rueckert, D., and Passerat-Palmbach, J · 2018
Later among the works it cites.
High-accuracy low-precision training
Sa, C. D., Leszczynski, M., Zhang, J., Marzoev, A., Aberger, C. R., Olukotun, K., and Ré, C · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in tensorflow
Sergeev, A., and Balso, M. D · 2018
Later among the works it cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Later among the works it cites.
esgd: Communication efficient distributed deep learning on the edge
Tao, Z., and Li, Q · 2018
Later among the works it cites.
Aggressive synchronization with partial processing for iterative ml jobs on clusters
Wang, S., Chen, W., Pi, A., and Zhou, X · 2018
Later among the works it cites.
Gandiva: Introspective cluster scheduling for deep learning
Xiao, W., Bhardwaj, R., Ramjee, R., Sivathanu, M., Kwatra, N., Han, Z., Patel, P., Peng, X., Zhao, H., Zhang, Q., Yang, F., and Zhou, L · 2018
Later among the works it cites.
Recent trends in deep learning based natural language processing [review article]
Young, T., Hazarika, D., Poria, S., and Cambria, E · 2018
Later among the works it cites.
Dynamic control flow in large-scale machine learning
Yu, Y., Abadi, M., Barham, P., Brevdo, E., Burrows, M., Davis, A., Dean, J., Ghemawat, S., Harley, T., Hawkins, P., Isard, M., Kudlur, M., Monga, R., Murray, D., and Zheng, X · 2018
Later among the works it cites.
Stay fresh: Speculative synchronization for fast distributed machine learning
Zhang, C., Tian, H., Wang, W., and Yan, F · 2018
Later among the works it cites.
https://developer.nvidia.com/nccl
NVIDIA Collective Communications Library (NCCL) · 2019
Closest in time.
https://www.nvidia.com/en-us/data-center/dgx-station/
NVIDIA DGX Station · 2019
Closest in time.
https://onnx.ai/
ONNX · 2019
Closest in time.
Towards federated learning at scale: System design
Bonawitz, K., Eichner, H., Grieskamp, W., Huba, D., Ingerman, A., Ivanov, V., Kiddon, C. M., Konečný, J., Mazzocchi, S., McMahan, B., Overveldt, T. V., Petrou, D., Ramage, D., and Roselander, J · 2019
Closest in time.
https://docs.microsoft.com/en-us/cognitive-toolkit/multiple-gpus-and-machines
Multiple GPUs and Machines – Cognitive Toolkit – CNTK | Microsoft Docs · 2019
Closest in time.
https://developer.apple.com/machine-learning/
Apple CoreML · 2019
Closest in time.
https://deeplearning4j.org/docs/latest/deeplearning4j-scaleout-technicalref
Deeplearning4j on Spark: Technical Explanation · 2019
Closest in time.
Tictac: Accelerating distributed deep learning with communication scheduling
Hashemi, S. H., Jyothi, S. A., and Campbell, R. H · 2019
Closest in time.
A comprehensive survey of deep learning for image captioning
Hossain, M. Z., Sohel, F., Shiratuddin, M. F., and Laga, H · 2019
Closest in time.
Analysis of large-scale multi-tenant gpu clusters for dnn training workloads
Jeon, M., Venkataraman, S., Phanishayee, A., Qian, J., Xiao, W., and Yang, F · 2019
Closest in time.
CROSSBOW: scaling deep learning with small batch sizes on multi-gpu servers
Koliousis, A., Watcharapichat, P., Weidlich, M., Mai, L., Costa, P., and Pietzuch, P. R · 2019
Closest in time.
https://www.intel.ai/kubernetes-volume-controller-kvc-data-management-tailored-for-machine-learning-workloads-in-kubernetes
Data Management Tailored for Machine Learning Workloads in Kubernetes · 2019
Closest in time.
Evaluating modern GPU interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect
Li, A., Song, S. L., Chen, J., Li, J., Liu, X., Tallent, N. R., and Barker, K. J · 2019
Closest in time.
https://mxnet.incubator.apache.org/versions/master/faq/distributed_training.html
Distributed Training in MXNet · 2019
Closest in time.
https://github.com/apache/incubator-mxnet/issues/841
Does mxnet support Stale Synchronous Parallel (aka. SSP) · 2019
Closest in time.
https://mxnet.incubator.apache.org/versions/master/faq/gradient_compression.html
Gradient Compression — mxnet documentation · 2019
Closest in time.
https://mxnet.apache.org/
MXNet · 2019
Closest in time.
https://cwiki.apache.org/confluence/display/MXNET/MXNet+Graph+Optimization+and+Quantization+based+on+subgraph+and+MKL-DNN
MXNet Graph Optimization and Quantization based on subgraph and MKL-DNN · 2019
Closest in time.
https://github.com/Tencent/ncnn
Tencent ncnn · 2019
Closest in time.
http://paddlepaddle.org/
PaddlePaddle · 2019
Closest in time.
Park, J. H., Kim, S., Lee, J., Jeon, M., and Noh, S. H · 2019
Closest in time.
https://github.com/OpenMined/PySyft
OpenMined/PySyft: A library for encrypted, privacy preserving deep learning · 2019
Closest in time.
https://pytorch.org/docs/stable/notes/multiprocessing.html
Multiprocessing Best Practies – PyTorch Master Documentation · 2019
Closest in time.
https://code.fb.com/ml-applications/qnnpack/
QNNPACK: Open source library for optimized mobile deep learning · 2019
Closest in time.
https://www.sas.com/
SAS · 2019
Closest in time.
What is the Parameter Server?
Smola, A · 2019
Closest in time.
https://www.tensorflow.org/guide/distribute_strategy
Distributed Training in TensorFlow · 2019
Closest in time.
https://www.tensorflow.org/lite/performance/post_training_quantization
Post-training quantization · 2019
Closest in time.
https://www.tensorflow.org/federated
TensorFlow Federated · 2019
Closest in time.
https://github.com/Waikato/wekaDeeplearning4j
wekaDeeplearning4j · 2019
Closest in time.
A comprehensive survey on graph neural networks
Wu, Z., Pan, S., Chen, F., Long, G., Zhang, C., and Yu, P. S · 2019
Closest in time.
https://xgboost.ai/
XGBoost · 2019
Closest in time.