Fetching the paper…
Reading the bibliography…
Descent methods for deep networks are notoriously capricious: they require careful tuning of step size, momentum and weight decay, and which method will work best on a new benchmark is a priori unclear.
Tests on a cell assembly theory of the action of the brain, using a large digital computer
Rochester, N., Holland, J., Haibt, L., and Duda, W · 1956
Earlier work this paper cites.
The perceptron: A probabilistic model for information storage and organization in the brain
Rosenblatt, F · 1958
Earlier work this paper cites.
Self-organization of orientation sensitive cells in the striate cortex
von der Malsburg, C · 1973
Earlier work this paper cites.
The role of constraints in Hebbian learning
Miller, K. and MacKay, D · 1994
Earlier work this paper cites.
An elementary introduction to modern convex geometry
Ball, K · 1997
Earlier work this paper cites.
Some PAC-Bayesian theorems
McAllester, D. A · 1998
Earlier work this paper cites.
Bounds for averaging classifiers
Langford, J. and Seeger, M · 2001
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Nesterov, Y · 2004
Earlier work this paper cites.
The self-tuning neuron: Synaptic scaling of excitatory synapses
Turrigiano, G · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Lecture 6.5—RMSprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. E · 2012
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Homeostatic role of heterosynaptic plasticity: models and experiments
Chistiakova, M., Bannon, N., Chen, J.-Y., Bazhenov, M., and Volgushev, M · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Decoupled networks
Liu, W., Liu, Z., Yu, Z., Dai, B., Lin, R., Wang, Y., Rehg, J. M., and Song, L · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Later among the works it cites.
On the convergence of Adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Later among the works it cites.
Proxquant: Quantized neural networks via proximal operators
Bai, Y., Wang, Y.-X., and Liberty, E · 2019
Later among the works it cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2019
Later among the works it cites.
Non-convex projected gradient descent for generalized low-rank tensor regression
Chen, H., Raskutti, G., and Yuan, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Cited alongside, same era.
Centered weight normalization in accelerating training of deep neural networks
Huang, L., Liu, X., Liu, Y., Lang, B., and Tao, D · 2017
Cited alongside, same era.
Deep hyperspherical learning
Liu, W., Zhang, Y.-M., Li, X., Yu, Z., Dai, B., Zhao, T., and Song, L · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Scaling SGD batch size to 32K for ImageNet training
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Micro-batch training with batch-channel normalization and weight standardization
Qiao, S., Wang, H., Liu, C., Shen, W., and Yuille, A · 2019
Later among the works it cites.
Optimization for deep learning: theory and algorithms
Sun, R · 2019
Later among the works it cites.
Deep learning generalizes because the parameter–function map is biased towards simple functions
Valle-Perez, G., Camargo, C. Q., and Louis, A. A · 2019
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S., Saxe, A. M., and Sompolinsky, H · 2020
Later among the works it cites.
A correspondence between normalization strategies in artificial and biological neural networks
Shen, Y., Wang, J., and Navlakha, S · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Later among the works it cites.
Characterizing signal propagation to close the performance gap in unnormalized ResNets
Brock, A., De, S., and Smith, S. L · 2021
Closest in time.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Closest in time.
Orthogonal over-parameterized training
Liu, W., Lin, R., Liu, Z., Rehg, J. M., Paull, L., Xiong, L., Song, L., and Weller, A · 2021
Closest in time.