Fetching the paper…
Reading the bibliography…
On the stability of inverse problems
Andrey Nikolayevich Tikhonov et al · 1943
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E Hoerl and Robert W Kennard · 1970
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John Hertz · 1991
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Remarks on strongly convex functions
Nelson Merentes and Kazimierz Nikodem · 2010
Earlier work this paper cites.
A brief introduction to olympiad inequalities
Evan Chen · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2015
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
On the importance of normalisation layers in deep learning with piecewise linear activation units
Zhibin Liao and Gustavo Carneiro · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
Twan Van Laarhoven · 2017
Cited alongside, same era.
Understanding batch normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Cited alongside, same era.
Towards understanding regularization in batch normalization
Ping Luo, Xinjiang Wang, Wenqi Shao, and Zhanglin Peng · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Cited alongside, same era.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Later among the works it cites.
Randall Balestriero and Richard G Baraniuk · 2022
Later among the works it cites.
Nearest class-center simplification through intermediate layers
Ido Ben-Shaul and Shai Dekel · 2022
Later among the works it cites.
On the emergence of simplex symmetry in the final and penultimate layers of neural network classifiers
Weinan E and Stephan Wojtowytsch · 2022
Later among the works it cites.
Neural collapse under mse loss: Proximity to and dynamics on the central path
Xu Han, Vahe Papyan, and David L Donoho · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exponential convergence rates for batch normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann, Ming Zhou, and Klaus Neymeyr · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2019
Cited alongside, same era.
Why do better loss functions lead to less transferable features?
Simon Kornblith, Ting Chen, Honglak Lee, and Mohammad Norouzi · 2020
Cited alongside, same era.
Neural collapse with unconstrained features, 2020
Dustin G. Mixon, Hans Parshall, and Jianzong Pi · 2020
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, X. Y. Han, and David L. Donoho · 2020
Cited alongside, same era.
Explicit regularization and implicit bias in deep network classifiers trained with the square loss, 2020
Tomaso Poggio and Qianli Liao · 2020
Cited alongside, same era.
Wenlong Ji, Yiping Lu, Yiliang Zhang, Zhun Deng, and Weijie J. Su · 2022
Later among the works it cites.
Extended unconstrained features model for exploring deep neural collapse, 2022
Tom Tirer and Joan Bruna · 2022
Later among the works it cites.
Neural collapse with normalized features: A geometric analysis over the riemannian manifold
Can Yaras, Peng Wang, Zhihui Zhu, Laura Balzano, and Qing Qu · 2022
Later among the works it cites.
Jinxin Zhou, Xiao Li, Tianyu Ding, Chong You, Qing Qu, and Zhihui Zhu · 2022
Later among the works it cites.
Why do we need weight decay in modern deep learning?
Maksym Andriushchenko, Francesco D’Angelo, Aditya Varre, and Nicolas Flammarion · 2023
Closest in time.
Neural collapse: A review on modelling principles and generalization
Vignesh Kothapalli · 2023
Closest in time.
Deep neural collapse is provably optimal for the deep unconstrained features model, 2023
Peter Súkeník, Marco Mondelli, and Christoph Lampert · 2023
Closest in time.