Fetching the paper…
Reading the bibliography…
Deep Neural Networks (DNNs) are becoming an important tool in modern computing applications.
Asynchronous Methods for Deep Reinforcement Learning. In
V. Mnih et al · 1937
Earlier work this paper cites.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro. 1951 · 1951
Earlier work this paper cites.
A Method for the Construction of Minimum-Redundancy Codes
D. A. Huffman. 1952 · 1952
Earlier work this paper cites.
Gaussian Elimination is Not Optimal
V. Strassen. 1969 · 1969
Earlier work this paper cites.
The Parallel Evaluation of General Arithmetic Expressions
R. P. Brent. 1974 · 1974
Earlier work this paper cites.
Arithmetic Complexity of Computations
S. Winograd. 1980 · 1980
Earlier work this paper cites.
The Byzantine Generals Problem
L. Lamport, R. Shostak, and M. Pease. 1982 · 1982
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate
Y. Nesterov. 1983 · 1983
Earlier work this paper cites.
An asynchronous algorithm for scattering information between the active nodes of a multicomputer system
Z. Drezner and A. Barak. 1986 · 1986
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
J. Tsitsiklis, D. Bertsekas, and M. Athans. 1986 · 1986
Earlier work this paper cites.
Backpropagation Applied to Handwritten Zip Code Recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. 1989 · 1989
Earlier work this paper cites.
Finding Structure in Time
J. L. Elman. 1990 · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
P. J. Werbos. 1990 · 1990
Earlier work this paper cites.
An Efficient Implementation of the Back-propagation Algorithm on the Connection Machine CM-2
X. Zhang, M. McKenna, J. P. Mesirov, and D. L. Waltz. 1990 · 1990
Earlier work this paper cites.
A Comparative Analysis of Selection Schemes Used in Genetic Algorithms
D. E. Goldberg and K. Deb. 1991 · 1991
Earlier work this paper cites.
Acceleration of Stochastic Approximation by Averaging
B. T. Polyak and A. B. Juditsky. 1992 · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams. 1992 · 1992
Earlier work this paper cites.
LogP: Towards a Realistic Model of Parallel Computation. In
D. Culler, R. Karp, D. Patterson, A. Sahay, K. E. Schauser, E. Santos, R. Subramonian, and T. von Eicken. 1993 · 1993
Earlier work this paper cites.
Architectures for Neuro-Computers: Review and Performance Evaluation
P. Ienne. 1993 · 1993
Earlier work this paper cites.
Efficient computation of the DFT with only a subset of input or output points
H. V. Sorensen and C. S. Burrus. 1993 · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi. 1994 · 1994
Earlier work this paper cites.
Neural net simulation on parallel computers. In
U. A. Muller and A. Gunzinger. 1994 · 1994
Earlier work this paper cites.
Rounding Errors in Algebraic Processes
J. H. Wilkinson. 1994 · 1994
Earlier work this paper cites.
Parallel neural network training on Multi-Spert. In
P. Farber and K. Asanovic. 1997 · 1997
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Thread Scheduling for Multiprogrammed Multiprocessors. In
N. S. Arora, R. D. Blumofe, and C. G. Plaxton. 1998 · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. 1998 · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits
Y. LeCun and C. Cortes. 1998 · 1998
Earlier work this paper cites.
Scheduling Multithreaded Computations by Work Stealing
R. D. Blumofe and C. E. Leiserson. 1999 · 1999
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian. 1999 · 1999
Earlier work this paper cites.
Ensemble Methods in Machine Learning. In
T. G. Dietterich. 2000 · 2000
Earlier work this paper cites.
Optimization of collective reduction operations. In
R. Rabenseifner. 2004 · 2004
Earlier work this paper cites.
Gossip algorithms: design, analysis and applications. In
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah. 2005 · 2005
Earlier work this paper cites.
High Performance Convolutional Neural Networks for Document Processing. In
K. Chellapilla, S. Puri, and P. Simard. 2006 · 2006
Earlier work this paper cites.
A Fast Learning Algorithm for Deep Belief Nets
G. E. Hinton, S. Osindero, and Y. W. Teh. 2006 · 2006
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. Wright. 2006 · 2006
Earlier work this paper cites.
Greedy Layer-Wise Training of Deep Networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle. 2007 · 2007
Earlier work this paper cites.
Collective Communication: Theory, Practice, and Experience: Research Articles
E. Chan, M. Heimlich, A. Purkayastha, and R. van de Geijn. 2007 · 2007
Earlier work this paper cites.
Map-Reduce for Machine Learning on Multicore
C. Chu, S. K. Kim, Y. Lin, Y. Yu, G. Bradski, K. Olukotun, and A. Y. Ng. 2007 · 2007
Earlier work this paper cites.
MapReduce: Simplified Data Processing on Large Clusters
J. Dean and S. Ghemawat. 2008 · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database. In
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Sparse Collective Operations for MPI. In
T. Hoefler and J. L. Traeff. 2009 · 2009
Earlier work this paper cites.
Intel Math Kernel Library. Reference Manual
Intel. 2009 · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky. 2009 · 2009
Earlier work this paper cites.
Unsupervised feature learning for audio classification using convolutional deep belief networks
H. Lee, P. Pham, Y. Largman, and A. Y. Ng. 2009 · 2009
Earlier work this paper cites.
Robust Stochastic Approximation Approach to Stochastic Programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro. 2009 · 2009
Earlier work this paper cites.
Large-scale Deep Unsupervised Learning Using Graphics Processors. In
R. Raina, A. Madhavan, and A. Y. Ng. 2009 · 2009
Earlier work this paper cites.
Asynchronous gossip algorithms for stochastic optimization. In
S. Sundhar Ram, A. Nedic, and V. V. Veeravalli. 2009 · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks. In
X. Glorot and Y. Bengio. 2010 · 2010
Earlier work this paper cites.
Deep Learning via Hessian-free Optimization. In
J. Martens. 2010 · 2010
Earlier work this paper cites.
Tiled convolutional neural networks
J. Ngiam, Z. Chen, D. Chia, P. W. Koh, Q. V. Le, and A. Y. Ng. 2010 · 2010
Earlier work this paper cites.
A Survey on Transfer Learning
S. J. Pan and Q. Yang. 2010 · 2010
Earlier work this paper cites.
Parallelized Stochastic Gradient Descent. In
M. A. Zinkevich, M. Weimer, A. Smola, and L. Li. 2010 · 2010
Earlier work this paper cites.
Distributed Delayed Stochastic Optimization
A. Agarwal and J. C. Duchi. 2011 · 2011
Earlier work this paper cites.
Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein. 2011 · 2011
Earlier work this paper cites.
Torch7: A Matlab-like Environment for Machine Learning. In
R. Collobert, K. Kavukcuoglu, and C. Farabet. 2011 · 2011
Earlier work this paper cites.
Shallow vs. Deep Sum-Product Networks
O. Delalleau and Y. Bengio. 2011 · 2011
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
J. Duchi, E. Hazan, and Y. Singer. 2011 · 2011
Earlier work this paper cites.
Hybrid Deterministic-Stochastic Methods for Data Fitting
M. P. Friedlander and M. W. Schmidt. 2011 · 2011
Earlier work this paper cites.
Generic Topology Mapping Strategies for Large-scale Parallel Architectures. In
T. Hoefler and M. Snir. 2011 · 2011
Earlier work this paper cites.
On Optimization Methods for Deep Learning. In
Q. V. Le, J. Ngiam, A. Coates, A. Lahiri, B. Prochnow, and A. Y. Ng. 2011 · 2011
Earlier work this paper cites.
Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
B. Recht, C. Re, S. Wright, and F. Niu. 2011 · 2011
Earlier work this paper cites.
Improving the speed of neural networks on CPUs. In
V. Vanhoucke, A. Senior, and M. Z. Mao. 2011 · 2011
Earlier work this paper cites.
Large Scale Distributed Deep Networks. In
J. Dean et al · 2012
Earlier work this paper cites.
Optimal Distributed Online Prediction Using Mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao. 2012 · 2012
Earlier work this paper cites.
Scalable stacking and learning for building deep architectures. In
L. Deng, D. Yu, and J. Platt. 2012 · 2012
Earlier work this paper cites.
Neural Networks for Machine Learning, Lecture 6a: Overview of Mini-batch Gradient Descent
G. Hinton. 2012 · 2012
Earlier work this paper cites.
Optimization Principles for Collective Neighborhood Communications. In
T. Hoefler and T. Schneider. 2012 · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
A. Krizhevsky, I. Sutskever, and G. Hinton. 2012 · 2012
Earlier work this paper cites.
Building High-level Features Using Large Scale Unsupervised Learning. In
Q. V. Le, M. Ranzato, R. Monga, M. Devin, K. Chen, G. S. Corrado, J. Dean, and A. Y. Ng. 2012 · 2012
Earlier work this paper cites.
Practical Bayesian Optimization of Machine Learning Algorithms
J. Snoek, H. Larochelle, and R. P Adams. 2012 · 2012
Earlier work this paper cites.
Deep Learning of Representations: Looking Forward. In
Y. Bengio. 2013 · 2013
Earlier work this paper cites.
Mitosis Detection in Breast Cancer Histology Images with Deep Neural Networks. In
D. C. Cireşan, A. Giusti, L. M. Gambardella, and J. Schmidhuber. 2013 · 2013
Earlier work this paper cites.
Deep Learning with COTS HPC Systems. In
A. Coates, B. Huval, T. Wang, D. J. Wu, A. Y. Ng, and B. Catanzaro. 2013 · 2013
Earlier work this paper cites.
More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server. In
Q. Ho et al · 2013
Earlier work this paper cites.
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
R. Johnson and T. Zhang. 2013 · 2013
Earlier work this paper cites.
GPU Asynchronous Stochastic Gradient Descent to Speed Up Neural Network Training
T. Paine, H. Jin, J. Yang, Z. Lin, and T. S. Huang. 2013 · 2013
Earlier work this paper cites.
Multi-GPU Training of ConvNets
O. Yadan, K. Adams, Y. Taigman, and M. Ranzato. 2013 · 2013
Earlier work this paper cites.
Asynchronous stochastic gradient descent for DNN training. In
S. Zhang, C. Zhang, Z. You, R. Zheng, and B. Xu. 2013 · 2013
Earlier work this paper cites.
Do Deep Nets Really Need to be Deep?
J. Ba and R. Caruana. 2014 · 2014
Earlier work this paper cites.
DianNao: A Small-footprint High-throughput Accelerator for Ubiquitous Machine-learning. In
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam. 2014 · 2014
Earlier work this paper cites.
cuDNN: Efficient Primitives for Deep Learning
S. Chetlur et al · 2014
Earlier work this paper cites.
Project Adam: Building an Efficient and Scalable Deep Learning Training System. In
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman. 2014 · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In
K. Cho et al · 2014
Earlier work this paper cites.
Minimizing Computation in Convolutional Neural Networks. In
J. Cong and B. Xiao. 2014 · 2014
Earlier work this paper cites.
Generative Adversarial Nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. 2014 · 2014
Earlier work this paper cites.
Using Advanced MPI: Modern Features of the Message-Passing Interface
W. Gropp, T. Hoefler, R. Thakur, and E. Lusk. 2014 · 2014
Cited alongside, same era.
Energy, Memory, and Runtime Tradeoffs for Implementing Collective Communication Operations
T. Hoefler and D. Moor. 2014 · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding. In
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. 2014 · 2014
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky. 2014 · 2014
Cited alongside, same era.
Scaling Distributed Machine Learning with the Parameter Server. In
M. Li et al · 2014
Cited alongside, same era.
Network In Network. In
Multi-Scale Context Aggregation by Dilated Convolutions. In
F. Yu and V. Koltun. 2016 · 2016
Later among the works it cites.
Using Supercomputer to Speed up Neural Network Training. In
Y. Yu, J. Jiang, and X. Chi. 2016 · 2016
Later among the works it cites.
DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
S. Zhou, Z. Ni, X. Zhou, H. Wen, Y. Wu, and Y. Zou. 2016 · 2016
Later among the works it cites.
ZNNi: Maximizing the Inference Throughput of 3D Convolutional Networks on CPUs and GPUs. In
A. Zlateski, K. Lee, and H. S. Seung. 2016 · 2016
Later among the works it cites.
Sparse Communication for Distributed Gradient Descent
A. F. Aji and K. Heafield. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Lin, Q. Chen, and S. Yan. 2014 · 2014
Cited alongside, same era.
Fast Training of Convolutional Networks through FFTs
M. Mathieu, M. Henaff, and Y. LeCun. 2014 · 2014
Cited alongside, same era.
Distributed learning of multilingual DNN feature extractors using GPUs. In
Y. Miao, H. Zhang, and F. Metze. 2014 · 2014
Cited alongside, same era.
Dogwild!-Distributed Hogwild for CPU & GPU. In
C. Noel and S. Osindero. 2014 · 2014
Cited alongside, same era.
Parallel training of Deep Neural Networks with Natural Gradient and Parameter Averaging
D. Povey, X. Zhang, and S. Khudanpur. 2014 · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David. 2014 · 2014
Cited alongside, same era.
Large-Scale Deep Belief Nets With MapReduce
K. Zhang and X. W. Chen. 2014 · 2014
Cited alongside, same era.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic. 2017 · 2017
Later among the works it cites.
S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters. In
A. A. Awan, K. Hamidouche, J. M. Hashmi, and D. K. Panda. 2017 · 2017
Later among the works it cites.
Distributed Second-Order Optimization using Kronecker-Factored Approximations. In
J. Ba, R. Grosse, and J. Martens. 2017 · 2017
Later among the works it cites.
Practical Neural Network Performance Prediction for Early Stopping
B. Baker, O. Gupta, R. Raskar, and N. Naik. 2017b · 2017
Later among the works it cites.
End-to-end optimized image compression. In
J. Ballé, V. Laparra, and E. P. Simoncelli. 2017 · 2017
Later among the works it cites.
Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent
P. Blanchard, E. M. El Mhamdi, R. Guerraoui, and J. Stainer. 2017 · 2017
Later among the works it cites.
SMASH: One-Shot Model Architecture Search through HyperNetworks
A. Brock, T. Lim, J. M. Ritchie, and N. Weston. 2017 · 2017
Later among the works it cites.
AdaComp : Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
C.-Y. Chen, J. Choi, D. Brand, A. Agrawal, W. Zhang, and K. Gopalakrishnan. 2017 · 2017
Later among the works it cites.
Dual Path Networks
Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, and J. Feng. 2017 · 2017
Later among the works it cites.
Parallel Deep Neural Network Training for Big Data on Blue Gene/Q
I. H. Chung et al · 2017
Later among the works it cites.
Simple And Efficient Architecture Search for Convolutional Neural Networks
T. Elsken, J.-H. Metzen, and F. Hutter. 2017 · 2017
Later among the works it cites.
On the Performance of Network Parallel Training in Artificial Neural Networks
L. Ericson and R. Mbuvha. 2017 · 2017
Later among the works it cites.
Meta Learning Shared Hierarchies
K. Frans, J. Ho, X. Chen, P. Abbeel, and J. Schulman. 2017 · 2017
Later among the works it cites.
AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
A. Gaunt, M. Johnson, M. Riechert, D. Tarlow, R. Tomioka, D. Vytiniotis, and S. Webster. 2017 · 2017
Later among the works it cites.
Bringing HPC Techniques to Deep Learning
A. Gibiansky. 2017 · 2017
Later among the works it cites.
TensorFlow XLA Overview
Google. 2017 · 2017
Later among the works it cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He. 2017 · 2017
Later among the works it cites.
Distributed Hessian-Free Optimization for Deep Neural Network. In
X. He, D. Mudigere, M. Smelyanskiy, and M. Takac. 2017 · 2017
Later among the works it cites.
Corrected Gossip Algorithms for Fast Reliable Broadcast on Unreliable Systems. In
T. Hoefler, A. Barak, A. Shiloh, and Z. Drezner. 2017 · 2017
Later among the works it cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry. 2017 · 2017
Later among the works it cites.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. 2017 · 2017
Later among the works it cites.
Gaia: Geo-distributed Machine Learning Approaching LAN Speeds. In
K. Hsieh, A. Harlap, N. Vijaykumar, D. Konomis, G. R. Ganger, P. B. Gibbons, and O. Mutlu. 2017 · 2017
Later among the works it cites.
Densely connected convolutional networks. In
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger. 2017 · 2017
Later among the works it cites.
Population Based Training of Neural Networks
M. Jaderberg et al · 2017
Later among the works it cites.
Heterogeneity-aware Distributed Parameter Servers. In
J. Jiang, B. Cui, C. Zhang, and L. Yu. 2017 · 2017
Later among the works it cites.
In-Datacenter Performance Analysis of a Tensor Processing Unit. In
N. P. Jouppi et al · 2017
Later among the works it cites.
L. Kaiser, A. N. Gomez, N. Shazeer, A. Vaswani, N. Parmar, L. Jones, and J. Uszkoreit. 2017 · 2017
Later among the works it cites.
Progressive Growing of GANs for Improved Quality, Stability, and Variation
T. Karras, T. Aila, S. Laine, and J. Lehtinen. 2017 · 2017
Later among the works it cites.
Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks
U. Köster et al · 2017
Later among the works it cites.
Deep Learning at 15PF: Supervised and Semi-supervised Classification for Scientific Data. In
T. Kurth et al · 2017
Later among the works it cites.
DeepRebirth: Accelerating Deep Neural Network Execution on Mobile Devices
D. Li, X. Wang, and D. Kong. 2017 · 2017
Later among the works it cites.
Ease.ml: Towards Multi-tenant Resource Sharing for Machine Learning Workloads
T. Li, J. Zhong, J. Liu, W. Wu, and C. Zhang. 2017 · 2017
Later among the works it cites.
Deep Reinforcement Learning: An Overview
Y. Li. 2017 · 2017
Later among the works it cites.
Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
X. Lian et al · 2017
Later among the works it cites.
Progressive Neural Architecture Search
C. Liu, B. Zoph, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy. 2017 · 2017
Later among the works it cites.
Hyper-parameter Selection in Deep Neural Networks Using Parallel Particle Swarm Optimization. In
P. R. Lorenzo, J. Nalepa, L. S. Ramos, and J. R. Pastor. 2017 · 2017
Later among the works it cites.
SGDR: Stochastic Gradient Descent with Warm Restarts. In
I. Loshchilov and F. Hutter. 2017 · 2017
Later among the works it cites.
R. Miikkulainen et al · 2017
Later among the works it cites.
DeepArchitect: Automatically Designing and Training Deep Architectures
R. Negrinho and G. Gordon. 2017 · 2017
Later among the works it cites.
Basic Linear Algebra Subprograms (BLAS)
Netlib. 2017 · 2017
Later among the works it cites.
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks?. In
E. Nurvitadhi et al · 2017
Later among the works it cites.
CUBLAS Library Documentation
NVIDIA. 2017a · 2017
Later among the works it cites.
Programming Tensor Cores in CUDA 9
NVIDIA. 2017b · 2017
Later among the works it cites.
Elastic Deep Learning
PaddlePaddle. 2017 · 2017
Later among the works it cites.
F. Petroski Such et al · 2017
Later among the works it cites.
Paleo: A Performance Model for Deep Neural Networks. In
H. Qi, E. R. Sparks, and A. Talwalkar. 2017 · 2017
Later among the works it cites.
Reflections on Random Kitchen Sinks
A. Rahimi and B. Recht. 2017 · 2017
Later among the works it cites.
Large-Scale Evolution of Image Classifiers. In
E. Real, S. Moore, A. Selle, S. Saxena, Y. L. Suematsu, J. Tan, Q. V. Le, and A. Kurakin. 2017 · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
S. L. Smith, P. Kindermans, and Q. V. Le. 2017 · 2017
Later among the works it cites.
Towards Pervasive and User Satisfactory CNN across GPU Microarchitectures. In
M. Song, Y. Hu, H. Chen, and T. Li. 2017 · 2017
Later among the works it cites.
Efficient Processing of Deep Neural Networks: A Tutorial and Survey
V. Sze, Y. H. Chen, T. J. Yang, and J. S. Emer. 2017 · 2017
Later among the works it cites.
Parallel Multi Channel Convolution using General Matrix Multiplication
A. Vasudevan, A. Anderson, and D. Gregg. 2017 · 2017
Later among the works it cites.
CHAOS: a parallelization scheme for training convolutional neural networks on Intel Xeon Phi
A. Viebke, S. Memeti, S. Pllana, and A. Abraham. 2017 · 2017
Later among the works it cites.
On the Origin of Deep Learning
H. Wang, B. Raj, and E. P. Xing. 2017 · 2017
Later among the works it cites.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li. 2017 · 2017
Later among the works it cites.
Genetic CNN. In
L. Xie and A. Yuille. 2017 · 2017
Later among the works it cites.
Large Batch Training of Convolutional Networks
Y. You, I. Gitman, and B. Ginsburg. 2017b · 2017
Later among the works it cites.
100-epoch ImageNet Training with AlexNet in 24 Minutes
Y. You, Z. Zhang, C. Hsieh, and J. Demmel. 2017c · 2017
Later among the works it cites.
Evolving Deep Networks Using HPC. In
S. R. Young et al · 2017
Later among the works it cites.
Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters. In
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing. 2017 · 2017
Later among the works it cites.
YellowFin and the Art of Momentum Tuning
J. Zhang, I. Mitliagkas, and C. Ré. 2017 · 2017
Later among the works it cites.
Practical Network Blocks Design with Q-Learning
Z. Zhong, J. Yan, and C.-L. Liu. 2017 · 2017
Later among the works it cites.
Neural Architecture Search with Reinforcement Learning. In
B. Zoph and Q. V. Le. 2017 · 2017
Later among the works it cites.
Learning Transferable Architectures for Scalable Image Recognition
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le. 2017 · 2017
Later among the works it cites.
TVM: End-to-End Optimization Stack for Deep Learning
T. Chen et al · 2018
Closest in time.
GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
J. Daily et al · 2018
Closest in time.
Communication-Optimal Convolutional Neural Nets
J. Demmel and G. Dinh. 2018 · 2018
Closest in time.
Integrated Model, Batch, and Domain Parallelism in Training Neural Networks. In
A. Gholami, A. Azad, P. H. Jin, K. Keutzer, and A. Buluç. 2018 · 2018
Closest in time.
Hyperparameter optimization: a spectral approach. In
E. Hazan, A. Klivans, and Y. Yuan. 2018 · 2018
Closest in time.
Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks. In
S. Krishnan, Y. Xiao, and R. A. Saurous. 2018 · 2018
Closest in time.
Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training. In
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally. 2018 · 2018
Closest in time.
Hierarchical Representations for Efficient Architecture Search. In
H. Liu, K. Simonyan, O. Vinyals, C. Fernando, and K. Kavukcuoglu. 2018 · 2018
Closest in time.
DARTS: Differentiable Architecture Search
H. Liu, K. Simonyan, and Y. Yang. 2018 · 2018
Closest in time.
Efficient Sparse-Winograd Convolutional Neural Networks
X. Liu, J. Pool, S. Han, and W. J. Dally. 2018 · 2018
Closest in time.
Accelerating Deep Learning Frameworks with Micro-batches. In
Y. Oyama, T. Ben-Nun, T. Hoefler, and S. Matsuoka. 2018 · 2018
Closest in time.
Efficient Neural Architecture Search via Parameter Sharing
H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, and J. Dean. 2018 · 2018
Closest in time.
Regularized Evolution for Image Classifier Architecture Search
E. Real, A. Aggarwal, Y. Huang, and Q. V Le. 2018 · 2018
Closest in time.
SparCML: High-Performance Sparse Communication for Machine Learning
C. Renggli, D. Alistarh, and T. Hoefler. 2018 · 2018
Closest in time.
Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
N. Vasilache et al · 2018
Closest in time.
Persistent RNNs: Stashing Recurrent Weights On-Chip. In
G. Diamos et al · 2033
Closest in time.