Fetching the paper…
Reading the bibliography…
BERT has recently attracted a lot of attention in natural language understanding (NLU) and achieved state-of-the-art results in various NLU tasks.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Y. E. Nesterov · 1983
Earlier work this paper cites.
Minimization methods for nonsmooth convex and quasiconvex functions
Y. E. Nesterov · 1984
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
L. B. Almeida, T. Langlois, J. D. Amaral, and A. Plakhov · 1999
Earlier work this paper cites.
Introductory Lectures on Convex Optimization
Y. Nesterov · 2004
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G.S. Corrado, R. Monga, K. Chen, M. Devin, Q.V. Le, and A. Ng · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Earlier work this paper cites.
Lecture 6.5 - RMSProp, COURSERA: Neural networks for machine learning, 2012
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
M. Li, T. Zhang, Y. Chen, and A. J. Smola · 2014
Cited alongside, same era.
Beyond convexity: Stochastic quasi-convex optimization
E. Hazan, K. Levy, and S. Shalev-Shwartz · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y Chen, T. Lillicrap, F. Hui, L. Sifre, G. Driessche, T. Graepel, and D. Hassabis · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Large batch training of convolutional networks
Y. You, I. Gitman, and B. Ginsburg · 2017
Later among the works it cites.
Block-normalized gradient method: An empirical study for training deep neural network
A. W. Yu, L. Huang, Q. Lin, R. Salakhutdinov, and J. Carbonell · 2017
Later among the works it cites.
Normalized gradient with adaptive stepsize method for deep neural network training
A. W. Yu, Q. Lin, R. Salakhutdinov, and J. Carbonell · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Layer-specific adaptive learning rates for deep networks
B. Singh, S. De, Y. Zhang, T. Goldstein, and G. Taylor · 2015
Cited alongside, same era.
Incorporating nesterov momentum into adam
T. Dozat · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Cited alongside, same era.
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Z. Zhou, Q. Zhang, G. Lu, H. Wang, W. Zhang, and Y. Yu · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, Ming-W. Chang, K. Lee, and K. Toutanova · 2019
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2019
Later among the works it cites.
Reducing bert pre-training time from 3 days to 76 minutes
Y. You, J. Li, J. Hseu, X. Song, J. Demmel, and C. Hsieh · 2020
Closest in time.