Fetching the paper…
Reading the bibliography…
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay.
Some methods of speeding up the convergence of iteration methods
Polyak, B · 1964
Earlier work this paper cites.
Minimization methods for nonsmooth convex and quasiconvex functions
Nesterov, Y. E · 1984
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K., and Shalev-Shwartz, S · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Layer-specific adaptive learning rates for deep networks
Singh, B., De, S., Zhang†, Y., Goldstein, T., and Taylor, G · 2015
Earlier work this paper cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Codreanu, V., Podareanu, D., and Saletore, V · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Improving generalization performance by switching from adam to SGD
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Later among the works it cites.
100-epoch imagenet training with alexnet in 24 minutes
You, Y., Zhang, Z., Hsieh, C., and Demmel, J · 2018
Later among the works it cites.
Block-normalized gradient method: An empirical study for training deep neural network
Yu, A. W., Lin, Q., Salakhutdinov, R., and Carbonell, J · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keskar, N. S. and Socher, R · 2017
Cited alongside, same era.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G. F., Elsen, E., García, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Bag of tricks for image classification with convolutional neural networks
He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M · 2018
Cited alongside, same era.
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Cited alongside, same era.
Closest in time.
Jasper: An end-to-end convolutional neural acoustic model
Li, J., Lavrukhin, V., Ginsburg, B., Leary, R., Kuchaiev, O., Cohen, J. M., Nguyen, H., and Gaddei, R. T · 2019
Closest in time.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Closest in time.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M. A., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Closest in time.
Towards a human-like open-domain chatbot
Adiwardana, D. D. F., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., and Le, Q. V · 2020
Closest in time.