Fetching the paper…
Reading the bibliography…
Large-batch SGD is important for scaling training of deep neural networks.
Scheduling multithreaded computations by work stealing
Blumofe, R. D. and Leiserson, C. E · 1999
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., LeCun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Zoneout: Regularizing rnns by randomly preserving hidden activations
Krueger, D., Maharaj, T., Kramár, J., Pezeshki, M., Ballas, N., Ke, N. R., Goyal, A., Bengio, Y., Courville, A., and Pal, C · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, K · 2016
Earlier work this paper cites.
Data augmentation generative adversarial networks
Antoniou, A., Storkey, A., and Edwards, H · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
DeVries, T. and Taylor, G. W · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Deep learning at 15pf: Supervised and semi-supervised classification for scientific data
Kurth, T., Zhang, J., Satish, N., Racah, E., Mitliagkas, I., Patwary, M. M. A., Malas, T., Sundaram, N., Bhimji, W., Smorkalov, M., Deslippe, J., Shiryaev, M., Sridharan, S., Prabhat, and Dubey, P · 2017
Darts: Differentiable architecture search
Liu, H., Simonyan, K., and Yang, Y · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Masters, D. and Luschi, C · 2018
Later among the works it cites.
Imagenet/resnet-50 training in 224 seconds
Mikami, H., Suganuma, H., U.-Chupala, P., Tanaka, Y., and Kageyama, Y · 2018
Later among the works it cites.
Step size matters tep size matters in deep learning deep learning
Nar, K. and Sastry, S. S · 2018
Later among the works it cites.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., and Matsuoka, S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R · 2017
Cited alongside, same era.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E · 2017
Cited alongside, same era.
A bayesian data augmentation approach for learning deep models
Tran, T., Pham, T., Carneiro, G., Palmer, L., and Reid, I · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Ben-Nun, T. and Hoefler, T · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2018
Cited alongside, same era.
Later among the works it cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Sergeev, A. and Balso, M. D · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J. M., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Later among the works it cites.
Rendergan: Generating realistic labeled data
Sixt, L., Wild, B., and Landgraf, T · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V · 2018
Later among the works it cites.
How SGD Selects the Global Minima in Over-parameterized Learning : A Dynamical Stability Perspective
Wu, L., Ma, C., and E, W · 2018
Later among the works it cites.
Image classification at supercomputer scale
Ying, C., Kumar, S., Chen, D., Wang, T., and Cheng, Y · 2018
Later among the works it cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Later among the works it cites.
On the computational inefficiency of large batch sizes for stochastic gradient descent, 2019
Golmant, N., Vemuri, N., Yao, Z., Feinberg, V., Gholami, A., Rothauge, K., Mahoney, M., and Gonzalez, J · 2019
Closest in time.