Fetching the paper…
Reading the bibliography…
Benchmarking the tradeoff between neural network accuracy and training time is computationally expensive.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Snapshot ensembles: Train 1, get m for free
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Making convolutional networks shift-invariant again
R. Zhang · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Cited alongside, same era.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
A. S. Golatkar, A. Achille, and S. Soatto · 2019
Cited alongside, same era.
When does label smoothing help?
R. Müller, S. Kornblith, and G. E. Hinton · 2019
Cited alongside, same era.
Demystifying learning rate policies for high accuracy training of deep neural networks
Y. Wu, L. Liu, J. Bae, K.-H. Chow, A. Iyengar, C. Pu, W. Wei, L. Yu, and Q. Zhang · 2019
Cited alongside, same era.
The early phase of neural network training
J. Frankle, D. J. Schwab, and A. S. Morcos · 2020
Cited alongside, same era.
Learning the pareto front with hypernetworks
A. Navon, A. Shamsian, G. Chechik, and E. Fetaya · 2020
Later among the works it cites.
Descending through a crowded valley-benchmarking deep learning optimizers
R. M. Schmidt, F. Schneider, and P. Hennig · 2021
Later among the works it cites.
Reduce, reuse, recycle: Improving training efficiency with distillation
C. Blakeney, J. Z. Forde, J. Frankle, Z. Zong, and M. L. Leavitt · 2022
Closest in time.
Super-acceleration with cyclical step-sizes
B. Goujaud, D. Scieur, A. Dieuleveut, A. B. Taylor, and F. Pedregosa · 2022
Closest in time.
MosaicML ResNet Deep Dive, 2022
M. Leavitt · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fastai: a layered api for deep learning
J. Howard and S. Gugger · 2020
Cited alongside, same era.
Controllable pareto multi-task learning
X. Lin, Z. Yang, Q. Zhang, and S. Kwong · 2020
Cited alongside, same era.
L. N. Smith · 2022
Closest in time.