Fetching the paper…
Reading the bibliography…
Deep learning thrives with large neural networks and large datasets.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Interprocessor collective communication library (intercom)
M. Barnett, L. Shuler, R. van De Geijn, S. Gupta, D. G. Payne, and J. Watts · 1994
Earlier work this paper cites.
Using MPI: Portable Parallel Programming with the Message-Passing Interface
W. Gropp, E. Lusk, and A. Skjellum · 1999
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2004
Earlier work this paper cites.
Optimization of collective reduction operations
R. Rabenseifner · 2004
Earlier work this paper cites.
Optimization of collective comm. operations in MPICH
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
L. Bottou · 2009
Earlier work this paper cites.
Natural language processing (almost) from scratch
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural nets
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun · 2014
Cited alongside, same era.
Visualizing and understanding convolutional neural networks
M. D. Zeiler and R. Fergus · 2014
Cited alongside, same era.
Fast R-CNN
R. Girshick · 2015
Cited alongside, same era.
Why random reshuffling beats stochastic gradient descent
M. Gürbüzbalaban, A. Ozdaglar, and P. Parrilo · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Revisiting Distributed Synchronous SGD
J. Chen, X. Pan, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Later among the works it cites.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
K. Chen and Q. Huo · 2016
Later among the works it cites.
Training and investigating Residual Nets
S. Gross and M. Wilber · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Quantized neural networks: Training neural networks with low precision weights and activations
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Later among the works it cites.
The Microsoft 2016 Conversational Speech Recognition System
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig · 2016
Later among the works it cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Closest in time.
Introducing Big Basin: Our next-generation AI hardware
K. Lee · 2017
Closest in time.
Scaling Distributed Machine Learning with System and Algorithm Co-design
M. Li · 2017
Closest in time.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Closest in time.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Closest in time.