Fetching the paper…
Reading the bibliography…
This paper describes the principle of "General Cyclical Training" in machine learning, where training starts and ends with "easy training" and the "hard training" happens during the middle epochs.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M Pawan Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Gradual dropin of layers to train very deep neural networks
Leslie N Smith, Emily M Hand, and Timothy Doster · 2016
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2018
Cited alongside, same era.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei · 2018
Cited alongside, same era.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2018
Cited alongside, same era.
The early phase of neural network training
Jonathan Frankle, David J Schwab, and Ari S Morcos · 2020
Later among the works it cites.
On the training dynamics of deep networks with l _ 2 l\_2 regularization
Aitor Lewkowycz and Guy Gur-Ari · 2020
Later among the works it cites.
Teacher algorithms for curriculum learning of deep rl in continuously parameterized environments
Rémy Portelas, Cédric Colas, Katja Hofmann, and Pierre-Yves Oudeyer · 2020
Later among the works it cites.
Building one-shot semi-supervised (boss) learning up to fully supervised performance
Leslie N Smith and Adam Conovaloff · 2020
Later among the works it cites.
Dehb: Evolutionary hyberband for scalable, robust and efficient hyperparameter optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Adaptive weight decay for deep neural networks
Kensuke Nakamura and Byung-Woo Hong · 2019
Cited alongside, same era.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Cited alongside, same era.
A new diet plan for weight decay
Leslie N. Smith · 2019
Cited alongside, same era.
Super-convergence: Very fast training of neural networks using large learning rates
Leslie N Smith and Nicholay Topin · 2019
Cited alongside, same era.
Pytorch image models
Ross Wightman · 2019
Cited alongside, same era.
How does learning rate decay help modern neural networks?
Kaichao You, Mingsheng Long, Jianmin Wang, and Michael I Jordan · 2019
Cited alongside, same era.
Noor Awad, Neeratyoy Mallik, and Frank Hutter · 2021
Later among the works it cites.
A survey on data augmentation for text classification
Markus Bayer, Marc-André Kaufhold, and Christian Reuter · 2021
Later among the works it cites.
Hyperparameter optimization: Foundations, algorithms, best practices and open challenges
Bernd Bischl, Martin Binder, Michel Lang, Tobias Pielok, Jakob Richter, Stefan Coors, Janek Thomas, Theresa Ullmann, Marc Becker, Anne-Laure Boulesteix, et al · 2021
Later among the works it cites.
Automated deep learning: Neural architecture search is not the end
Xuanyi Dong, David Jacob Kedziora, Katarzyna Musial, and Bogdan Gabrys · 2021
Later among the works it cites.
Hpobench: A collection of reproducible multi-fidelity benchmark problems for hpo
Katharina Eggensperger, Philipp Müller, Neeratyoy Mallik, Matthias Feurer, René Sass, Aaron Klein, Noor Awad, Marius Lindauer, and Frank Hutter · 2021
Later among the works it cites.
A survey on multi-objective hyperparameter optimization algorithms for machine learning
Alejandro Morales-Hernández, Inneke Van Nieuwenhuyse, and Sebastian Rojas Gonzalez · 2021
Later among the works it cites.
Tresnet: High performance gpu-dedicated architecture
Tal Ridnik, Hussam Lawen, Asaf Noy, Emanuel Ben Baruch, Gilad Sharir, and Itamar Friedman · 2021
Later among the works it cites.
Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks
Lin Wang and Kuk-Jin Yoon · 2021
Later among the works it cites.
Which samples should be learned first: Easy or hard?
Xiaoling Zhou and Ou Wu · 2021
Later among the works it cites.
Leslie N Smith · 2022
Closest in time.