Fetching the paper…
Reading the bibliography…
Large batch size training in deep neural networks (DNNs) possesses a well-known 'generalization gap' that remarkably induces generalization performance degradation.
Neural Networks: Tricks of the Trade
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K · 2012
Earlier work this paper cites.
The Effects of Hyperparameters on SGD Training of Neural Networks
Breuel, T. M · 2015
Earlier work this paper cites.
Optimizing Neural Networks with Kronecker-Factored Approximate Curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Persistent RNNs: Stashing Recurrent Weights on-Chip
Diamos, G., Sengupta, S., Catanzaro, B., Chrzanowski, M., Coates, A., Elsen, E., Engel, J., Hannun, A., and Satheesh, S · 2016
Earlier work this paper cites.
Distributed Training of Deep Neural Networks: Theoretical and Practical Limits of Parallel Scalability
Keuper, J. and Preundt, F · 2016
Earlier work this paper cites.
Proportionate Gradient Updates with PercentDelta
Abuelhaija, S · 2017
Earlier work this paper cites.
AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
Devarakonda, A., Naumov, M., and Garland, M · 2017
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Train Longer, Generalize Better: Closing the Generalization Gap in Large Batch Training of Neural Networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Batch Size Matters: A Diffusion Approximation Framework on Nonconvex Stochastic Gradient Descent
Li, C. J., Li, L., Qian, J., and Liu, J · 2017
Cited alongside, same era.
Exploring Generalization in Deep Learning
Neyshabur, B., Bhojanapalli, S., Mcallester, D., and Srebro, N · 2017
Cited alongside, same era.
MegDet: A Large Mini-Batch Object Detector
Peng, C., Xiao, T., Li, Z., Jiang, Y., Zhang, X., Jia, K., Yu, G., and Sun, J · 2017
Cited alongside, same era.
Super-Convergence: Very Fast Training of Residual Networks Using Large Learning Rates
Smith, L. N. and Topin, N · 2017
Cited alongside, same era.
Nonlinear Conjugate Gradients For Scaling Synchronous Distributed DNN Training
ImageNet/ResNet-50 Training in 224 Seconds
Mikami, H., Suganuma, H., Chupala, P. U., Tanaka, Y., and Kageyama, Y · 2018
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Smith, S. L., Kindermans, P., and Le, Q. V · 2018
Later among the works it cites.
SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning
Wen, W., Wang, Y., Yan, F., Xu, C., Wu, C., Chen, Y., and Li, H · 2018
Later among the works it cites.
Large Batch Size Training of Neural Networks with Adversarial Training and Second-Order Information
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W · 2018
Later among the works it cites.
ImageNet Training in Minutes
You, Y., Zhang, Z., Hsieh, C., Demmel, J., and Keutzer, K · 2018
Later among the works it cites.
Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD
Haruki, K., Suzuki, T., Hamakawa, Y., Toda, T., Sakai, R., Ozawa, M., and Kimura, M · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adya, S., Palakkode, V., and Tuzel, O · 2018
Cited alongside, same era.
Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent
Dai, X. and Zhu, Y · 2018
Cited alongside, same era.
Don’t Use Large Mini-Batches, Use Local SGD
Lin, T., Stich, S. U., and Jaggi, M · 2018
Cited alongside, same era.
An Empirical Model of Large-Batch Training
Mccandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Cited alongside, same era.
Finding Flatter Minima with SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A
Cited in the paper.
Three Factors Influencing Minima in SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Storkey, A., and Bengio, Y
Cited in the paper.
Large Batch Training of Convolutional Networks
You, Y., Gitman, I., and Ginsburg, B
Cited in the paper.
Later among the works it cites.
Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
Li, Y., Wei, C., and Ma, T · 2019
Later among the works it cites.
Interplay Between Optimization and Generalization of Stochastic Gradient Descent with Covariance Noise
Wen, Y., Luk, K., Gazeau, M., Zhang, G., Chan, H., and Ba, J · 2019
Later among the works it cites.