Fetching the paper…
Reading the bibliography…
The stochastic gradient descent (SGD) method is most widely used for deep neural network (DNN) training.
The design for the Wall Street Journal-based CSR corpus
Paul, D. B. and Baker, J. M · 1992
Earlier work this paper cites.
Synaptic weight noise during multilayer perceptron training: fault tolerance and training improvements
Murray, A. F. and Edwards, P. J · 1993
Earlier work this paper cites.
The Penn Treebank: annotating predicate argument structure
Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B · 1994
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
An, G · 1996
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
On weight-noise-injection training
Ho, K., Leung, C.-s., and Sum, J · 2008
Earlier work this paper cites.
Zur, R. M., Jiang, Y., Pesce, L. L., and Drukker, K · 2009
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Krizhevsky, A., Nair, V., and Hinton, G · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Regularization of neural networks using DropConnect
Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Deeply-supervised nets
Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Cited alongside, same era.
TensorFlow: A System for Large-scale Machine Learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., Ying, C., and Le, Q. V · 2017
Later among the works it cites.
Scaling SGD batch size to 32k for ImageNet training
You, Y., Gitman, I., and Ginsburg, B · 2017
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Later among the works it cites.
DNN’s Sharpest Directions Along the SGD Trajectory
Jastrzebski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2018
Later among the works it cites.
Finding flatter minima with SGD
Jastrzkebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2017
Cited alongside, same era.
How to Escape Saddle Points Efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
An Alternative View: When Does SGD Escape Local Minima?
Kleinberg, R., Li, Y., and Yuan, Y · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Later among the works it cites.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R · 2018
Later among the works it cites.
Nvidia dgx-1
NVIDIA · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Sergeev, A. and Del Balso, M · 2018
Later among the works it cites.
SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning
Wen, W., Wang, Y., Yan, F., Xu, C., Wu, C., Chen, Y., and Li, H · 2018
Later among the works it cites.
Image classification at supercomputer scale
Ying, C., Kumar, S., Chen, D., Wang, T., and Cheng, Y · 2018
Later among the works it cites.
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2018
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Later among the works it cites.
Simple Gated ConvNet for Small Footprint Acoustic Modeling
Lee, L., Park, J., and Sung, W · 2019
Later among the works it cites.