Understand
We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior.
- We perform a detailed analysis of scaling, and identify optimal design points for different networks.
- We demonstrate scaling of CNNs on 100s of nodes, and present what we believe to be record training throughputs.
- A 512 minibatch VGG-A CNN training run is scaled 90X on 128 nodes.
Built on
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Conversational speech transcription using context-dependent deep neural networks
Seide, Frank, Li, Gang, and Yu, Dong · 2011
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Sermanet, Pierre, Eigen, David, Zhang, Xiang, Mathieu, Michaël, Fergus, Rob, and LeCun, Yann · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Chilimbi, Trishul M., Suzue, Yutaka, Apacible, Johnson, and Kalyanaraman, Karthik · 2014
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J · 2014
Earlier work this paper cites.
Similar
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Seide, Frank, Fu, Hao, Droppo, Jasha, Li, Gang, and Yu, Dong · 2014
Cited alongside, same era.
On parallelizability of stochastic gradient descent for speech dnns
Seide, Frank, Fu, Hao, Droppo, Jasha, Li, Gang, and Yu, Dong · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2014
Cited alongside, same era.
Deep learning with elastic averaging SGD
Zhang, Sixin, Choromanska, Anna, and LeCun, Yann · 2014
Cited alongside, same era.
Then
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S., Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Goodfellow, Ian, Harp, Andrew, Irving, Geoffrey, Isard, Michael, Jia, Yangqing, Jozefowicz, Rafal, Kaiser, Lukasz, Kudlur, Manjunath, Levenberg, Josh, Mané, Dan, Monga, Rajat, Moore, Sherry, Murray, Derek, Olah, Chris, Schuster, Mike, Shlens, Jonathon, Steiner, Benoit, Sutskever, Ilya, Talwar, Kunal, Tucker, Paul, Vanhoucke, Vincent, Vasudevan, Vijay, Viégas, Fernanda, Vinyals, Oriol, Warden, Pete, Wattenberg, Martin, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2015
Later among the works it cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Iandola, Forrest N., Ashraf, Khalid, Moskewicz, Matthew W., and Keutzer, Kurt · 2015
Later among the works it cites.
Improving concurrency and asynchrony in multithreaded mpi applications using software offloading
Vaidyanathan, K., Kalamkar, D. D., Pamnany, K., Hammond, J. R., Balaji, P., Das, D., Park, J., and Joo, Balint · 2015
Later among the works it cites.
Deep image: Scaling up image recognition
Wu, Ren, Yan, Shengen, Shan, Yi, Dang, Qingqing, and Sun, Gang · 2015
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…