Fetching the paper…
Reading the bibliography…
Learning rate decay (lrDecay) is a \emph{de facto} technique for training modern neural networks.
Second Order Properties of Error Surfaces: Learning Time and Generalization
Yann LeCun, Ido Kanter, and Sara A. Solla · 1991
Earlier work this paper cites.
Efficient BackProp
Yann A. LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Caltech-256 object category dataset
Gregory Griffin, Alex Holub, and Pietro Perona · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Recognizing indoor scenes
A. Quattoni and A. Torralba · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Practical Recommendations for Gradient-Based Training of Deep Architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
How do humans sketch objects?
Mathias Eitz, James Hays, and Marc Alexa · 2012
Earlier work this paper cites.
ADADELTA: An Adaptive Learning Rate Method
Matthew D. Zeiler · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Learning and Transferring Mid-Level Image Representations using Convolutional Neural Networks
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, and Michael Bernstein · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Wide Residual Networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Accelerated Stochastic Power Iteration
Peng Xu, Bryan He, Christopher De Sa, Ioannis Mitliagkas, and Chris Re · 2018
Later among the works it cites.
Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Closest in time.
Exponential convergence rates for Batch Normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann, Ming Zhou, and Klaus Neymeyr · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N. Smith · 2017
Cited alongside, same era.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Understanding Batch Normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Cited alongside, same era.
An Alternative View: When Does SGD Escape Local Minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Cited alongside, same era.
Do better imagenet models transfer better?
Simon Kornblith, Jonathon Shlens, and Quoc V. Le · 2019
Closest in time.
Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Closest in time.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2019
Closest in time.
Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2019
Closest in time.
Adaptive Gradient Methods with Dynamic Bound of Learning Rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Closest in time.
Do deep neural networks learn shallow learnable examples first?
Karttikeya Mangalam and Vinay Uday Prabhu · 2019
Closest in time.
SGD on Neural Networks Learns Functions of Increasing Complexity
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, and Boaz Barak · 2019
Closest in time.
Transfusion: Understanding Transfer Learning for Medical Imaging
Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio · 2019
Closest in time.