Television by pulse code modulation
W. M. Goodall · 1951
Earlier work this paper cites.
Picture coding using pseudo-random noise
L. Roberts · 1962
Earlier work this paper cites.
A bridging model for parallel computation
L. G. Valiant · 1990
Earlier work this paper cites.
Probabilistic rounding in neural network learning with limited precision
M. Höhfeld and S. E. Fahlman · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Bulk synchronous parallel computing-a paradigm for transportable software
T. Cheatham, A. Fahmy, D. C. Stefanescu, and L. G. Valiant · 1995
Earlier work this paper cites.
Bulk synchronous parallel: practical experience with a model for parallel computing
R. Krizanc and A. Saarimaki · 1996
Earlier work this paper cites.
Efficient algorithms for all-to-all communications in multiport message-passing systems
J. Bruck, Ching-Tien Ho, S. Kipnis, E. Upfal, and D. Weathersby · 1997
Earlier work this paper cites.
A tonotopic artificial neural network architecture for phoneme probability estimation
N. Strom · 1997
Earlier work this paper cites.
Sparse connection and pruning in large dynamic artificial neural networks
N. Ström · 1997
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Focusing Solutions for Data Mining: Analytical Studies and Experimental Results in Real-world Domains
T. Reinartz · 1999
Earlier work this paper cites.
Connection-level analysis and modeling of network traffic
S. Sarvotham, R. Riedi, and R. Baraniuk · 2001
Earlier work this paper cites.
Finding frequent items in data streams
M. Charikar, K. Chen, and M. Farach-Colton · 2002
Earlier work this paper cites.
Gossip-based computation of aggregate information
D. Kempe, A. Dobra, and J. Gehrke · 2003
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Lin Xiao and S. Boyd · 2003
Earlier work this paper cites.
Optimization of collective reduction operations
R. Rabenseifner · 2004
Earlier work this paper cites.
Gossip algorithms: design, analysis and applications
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah · 2005
Earlier work this paper cites.
Optimization of collective communication operations in mpich
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Randomized gossip algorithms
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah · 2006
Earlier work this paper cites.
Consensus and cooperation in networked multi-agent systems
R. Olfati-Saber, J. A. Fax, and R. M. Murray · 2007
Earlier work this paper cites.
A distributed consensus protocol for clock synchronization in wireless sensor network
L. Schenato and G. Gamba · 2007
Earlier work this paper cites.
Mpi for python: Performance improvements and mpi-2 extensions
L. Dalcín, R. Paz, M. Storti, and J. D’Elía · 2008
Earlier work this paper cites.
Randomized consensus algorithms over large scale networks
F. Fagnani and S. Zampieri · 2008
Earlier work this paper cites.
Distributed stochastic subgradient projection algorithms for convex optimization
S. S. Ram, A. Nedic, and V. V. Veeravalli · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
BCube: a high performance, server-centric network architecture for modular data centers
C. Guo, G. Lu, D. Li, H. Wu, X. Zhang, Y. Shi, C. Tian, Y. Zhang, and S. Lu · 2009
Earlier work this paper cites.
Sparse online learning via truncated gradient
J. Langford, L. Li, and T. Zhang · 2009
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
R. Mcdonald, M. Mohri, N. Silberman, D. Walker, and G. S. Mann · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
A. Nedic and A. Ozdaglar · 2009
Earlier work this paper cites.
Two-tree algorithms for full bandwidth broadcast, reduction and scan
P. Sanders, J. Speck, and J. L. Träff · 2009
Earlier work this paper cites.
Gossip consensus algorithms via quantized communication
R. Carli, F. Fagnani, P. Frasca, and S. Zampieri · 2010
Earlier work this paper cites.
Toward performance models of mpi implementations for understanding application scaling issues
T. Hoefler, W. Gropp, R. Thakur, and J. L. Träff · 2010
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
A. Krizhevsky, V. Nair, and G. Hinton · 2010
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
R. T. McDonald, K. B. Hall, and G. Mann · 2010
Earlier work this paper cites.
Piccolo: Building fast, distributed programs with partitioned tables
R. Power and J. Li · 2010
Earlier work this paper cites.
An architecture for parallel topic models
A. Smola and S. Narayanamurthy · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. A. Zinkevich, M. Weimer, A. Smola, and L. Li · 2010
Earlier work this paper cites.
Parallel coordinate descent for l1-regularized loss minimization
J. K. Bradley, A. Kyrola, D. Bickson, and C. Guestrin · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Re, and S. J. Wright · 2011
Earlier work this paper cites.
Improving the speed of neural networks on cpus
V. Vanhoucke, A. Senior, and M. Z. Mao · 2011
Earlier work this paper cites.
Scalable inference in latent variable models
A. Ahmed, M. Aly, J. Gonzalez, S. Narayanamurthy, and A. J. Smola · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, et al · 2012
Earlier work this paper cites.
Distributed graphlab: A framework for machine learning and data mining in the cloud
Y. Low, D. Bickson, J. Gonzalez, C. Guestrin, A. Kyrola, and J. M. Hellerstein · 2012
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, J. K. Kim, S. Lee, P. B. Gibbons, G. A. Gibson, G. R. Ganger, and E. P. Xing · 2013
Earlier work this paper cites.
Asynchronous distributed optimization using a randomized alternating direction method of multipliers
F. Iutzeler, P. Bianchi, P. Ciblat, and W. Hachem · 2013
Earlier work this paper cites.
Parameter server for distributed machine learning
M. Li, L. Zhou, Z. Yang, A. Q. Li, F. Xia, D. G. Andersen, and A. J. Smola · 2013
Earlier work this paper cites.
Distributed delayed proximal gradient methods
D. G. A. M. Li and A. Smola · 2013
Earlier work this paper cites.
D-admm: A communication-efficient distributed algorithm for separable optimization
J. F. C. Mota, J. M. F. Xavier, P. M. Q. Aguiar, and M. Püschel · 2013
Earlier work this paper cites.
A survey of methods for distributed machine learning
D. Peteiro-Barral and B. Guijarro-Berdiñas · 2013
Earlier work this paper cites.
On the o(1=k) convergence of asynchronous distributed alternating direction method of multipliers
E. Wei and A. E. Ozdaglar · 2013
Earlier work this paper cites.
A block coordinate descent method for regularized multiconvex optimization with applications to nonnegative tensor factorization and completion
Y. Xu and W. Yin · 2013
Earlier work this paper cites.
Distributed stochastic gradient mcmc
S. Ahn, B. Shahbaba, and M. Welling · 2014
Earlier work this paper cites.
cuDNN: Efficient primitives for deep learning
Original
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
H. Cui, J. Cipar, Q. Ho, J. K. Kim, S. Lee, A. Kumar, J. Wei, W. Dai, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2014
Earlier work this paper cites.
Exploiting iterative-ness for parallel ml computations
H. Cui, A. Tumanov, J. Wei, L. Xu, W. Dai, J. Haber-Kucharsky, Q. Ho, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Original
A. Krizhevsky · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
M. Li, D. G. Andersen, A. Smola, and K. Yu · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Dimmwitted: A study of main-memory statistical analytics
C. Zhang and C. Ré · 2014
Earlier work this paper cites.
Asynchronous distributed admm for consensus optimization
R. Zhang and J. T. Kwok · 2014
Earlier work this paper cites.
Improving deep neural network acoustic models using generalized maxout networks
X. Zhang, J. Trmal, D. Povey, and S. Khudanpur · 2014
Earlier work this paper cites.
Mariana: Tencent deep learning platform and its applications
Y. Zou, X. Jin, Y. Li, Z. Guo, E. Wang, and B. Xiao · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
Earlier work this paper cites.
Passcode: Parallel asynchronous stochastic dual co-ordinate descent
C.-J. Hsieh, H.-F. Yu, and I. S. Dhillon · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Earlier work this paper cites.
Singa: A distributed deep learning platform
B. C. Ooi, K.-L. Tan, S. Wang, W. Wang, Q. Cai, G. Chen, J. Gao, Z. Luo, A. K. Tung, Y. Wang, Z. Xie, M. Zhang, and K. Zheng · 2015
Earlier work this paper cites.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Z. Qu, P. Richtárik, and T. Zhang · 2015
Earlier work this paper cites.
Multi-agent mirror descent for decentralized stochastic optimization
M. Rabbat · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
Taming the wild: A unified analysis of hogwild-style algorithms
C. D. Sa, C. Zhang, K. Olukotun, and C. Ré · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
N. Strom · 2015
Earlier work this paper cites.
Petuum: A new platform for distributed machine learning on big data
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
S. Zhang, A. Choromanska, and Y. LeCun · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, et al · 2016
Earlier work this paper cites.
Qsgd: Randomized quantization for communication-optimal stochastic gradient descent
Original
D. Alistarh, J. Li, R. Tomioka, and M. Vojnovic · 2016
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
D. Amodei, S. Ananthanarayanan, R. Anubhai, et al · 2016
Earlier work this paper cites.
Layer normalization
Original
J. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
On data dependence in distributed stochastic optimization
A. S. Bijral, A. D. Sarwate, and N. Srebro · 2016
Earlier work this paper cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Earlier work this paper cites.
Revisiting distributed synchronous sgd
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Earlier work this paper cites.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
K. Chen and Q. Huo · 2016
Earlier work this paper cites.
Gossip dual averaging for decentralized optimization of pairwise functions
I. Colin, A. Bellet, J. Salmon, and S. Clémençon · 2016
Earlier work this paper cites.
Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing · 2016
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks
N. Dryden, S. A. Jacobs, T. Moon, and B. Van Essen · 2016
Earlier work this paper cites.