Fetching the paper…
Reading the bibliography…
Adaptive optimization methods, which perform local optimization with a metric constructed from the history of iterates, are becoming increasingly popular for training deep neural networks.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D.P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan Salakhutdinov, and Nathan Srebro · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Torch blog
Sergey Zagoruyko · 2015
Cited alongside, same era.
Parsing as language modeling
Do Kook Choe and Eugene Charniak · 2016
Cited alongside, same era.
Span-based constituency parsing with a structure-label system and provably optimal dynamic oracles
James Cross and Liang Huang · 2016
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2016
Train faster, generalize better: Stability of stochastic gradient descent
Benjamin Recht, Moritz Hardt, and Yoram Singer · 2016
Later among the works it cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Later among the works it cites.
A peak at trends in machine learning
Andrej Karparthy · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Closest in time.
Diving into the shallows: a computational perspective on large-scale shallow learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Siyuan Ma and Mikhail Belkin · 2017
Closest in time.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Closest in time.