Fetching the paper…
Reading the bibliography…
The cross-entropy softmax loss is the primary loss function used to train deep neural networks.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Towards open set deep networks
Abhijit Bendale and Terrance E Boult · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Earlier work this paper cites.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2018
Cited alongside, same era.
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Cited alongside, same era.
Large-scale long-tailed recognition in an open world
Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X Yu · 2019
Cited alongside, same era.
Adaptive weight decay for deep neural networks
Kensuke Nakamura and Byung-Woo Hong · 2019
Cited alongside, same era.
Super-convergence: Very fast training of neural networks using large learning rates
Leslie N Smith and Nicholay Topin · 2019
Cited alongside, same era.
Understanding decoupled and early weight decay
Johan Bjorck, Kilian Weinberger, and Carla Gomes · 2020
Later among the works it cites.
The early phase of neural network training
Jonathan Frankle, David J Schwab, and Ari S Morcos · 2020
Later among the works it cites.
On the training dynamics of deep networks with l _ 2 l\_2 regularization
Aitor Lewkowycz and Guy Gur-Ari · 2020
Later among the works it cites.
Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection
Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
Pytorch image models
Ross Wightman · 2019
Cited alongside, same era.
How does learning rate decay help modern neural networks?
Kaichao You, Mingsheng Long, Jianmin Wang, and Michael I Jordan · 2019
Cited alongside, same era.
Focal loss in 3d object detection
Peng Yun, Lei Tai, Yuan Wang, Chengju Liu, and Ming Liu · 2019
Cited alongside, same era.
Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz, Philip HS Torr, and Puneet K Dokania · 2020
Later among the works it cites.
Asymmetric loss for multi-label classification
Tal Ridnik, Emanuel Ben-Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
Tresnet: High performance gpu-dedicated architecture
Tal Ridnik, Hussam Lawen, Asaf Noy, Emanuel Ben Baruch, Gilad Sharir, and Itamar Friedman · 2021
Later among the works it cites.
Contrastive unpaired translation using focal loss for patch classification
Bernard Spiegl · 2021
Later among the works it cites.
General cyclical training of neural networks
Leslie N Smith · 2022
Closest in time.