Fetching the paper…
Reading the bibliography…
In several recently proposed stochastic optimization methods (e.g.
The approximation of one matrix by another of lower rank
Eckart, C. and Young, G · 1936
Earlier work this paper cites.
Learning the parts of objects by nonnegative matrix factorization
Lee, Daniel D. and Seung, H. Sebastian · 1999
Earlier work this paper cites.
Nonnegative matrix factorization and I-divergence alternating minimization
Finesso, Lorenzo and Spreij, Peter · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John C., Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
Zeiler, Matthew D · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Cited alongside, same era.
Training highly multiclass classifiers
Gupta, Maya R., Bengio, Samy, and Weston, Jason · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2015
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, Priya, Dollár, Piotr, Girshick, Ross B., Noordhuis, Pieter, Wesolowski, Lukasz, Kyrola, Aapo, Tulloch, Andrew, Jia, Yangqing, and He, Kaiming · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, Noam, Mirhoseini, Azalia, Maziarz, Krzysztof, Davis, Andy, Le, Quoc, Hinton, Geoffrey, and Dean, Jeff · 2017
Later among the works it cites.
Attention is all you need
Vaswani, Ashish, Shazeer, Noam, Parmar, Niki, Uszkoreit, Jakob, Jones, Llion, Gomez, Aidan N, Kaiser, Łukasz, and Polosukhin, Illia · 2017
Later among the works it cites.
On the convergence of adam and beyond
Reddi, Sashank J., Kale, Satyen, and Kumar, Sanjiv · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…