Fetching the paper…
Reading the bibliography…
In the twilight of Moore's law, GPUs and other specialized hardware accelerators have dramatically sped up neural network training.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) {O}(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A. J · 2010
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G. E., and Hinton, G. E · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
SSD: Single shot multibox detector
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A. C · 2016
Cited alongside, same era.
Parallel SGD: When does averaging help?
Zhang, J., De Sa, C., Mitliagkas, I., and Ré, C · 2016
Characterizing deep-learning I/O workloads in TensorFlow
Chien, S. W., Markidis, S., Sishtla, C. P., Santos, L., Herman, P., Narasimhamurthy, S., and Laure, E · 2018
Later among the works it cites.
Faster SGD training by minibatch persistency
Fischetti, M., Mandatelli, I., and Salvagnin, D · 2018
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Later among the works it cites.
Efficient training of convolutional neural nets on large distributed systems
Kumar, S., Sreedhar, D., Saxena, V., Sabharwal, Y., and Verma, A · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Critical hyper-parameters: No random, no cry
Bousquet, O., Gelly, S., Kurach, K., Teytaud, O., and Vincent, D · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Later among the works it cites.
Image classification at supercomputer scale
Ying, C., Kumar, S., Chen, D., Wang, T., and Cheng, Y · 2018
Later among the works it cites.
Augment your batch: better training with larger batches
Hoffer, E., Ben-Nun, T., Hubara, I., Giladi, N., Hoefler, T., and Soudry, D · 2019
Closest in time.
Accelerating data loading in deep neural network training
Yang, C.-C. and Cong, G · 2019
Closest in time.