Fetching the paper…
Reading the bibliography…
Neural collapse is a highly symmetric geometric pattern of neural networks that emerges during the terminal phase of training, with profound implications on the generalization performance and robustness of the trained networks.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Characterization of the subdifferential of some matrix norms
G Alistair Watson · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
From symmetry to geometry: Tractable nonconvex problems
Yuqian Zhang, Qing Qu, and John Wright · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
How does mixup help with robustness and generalization?
Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani, and James Zou · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Approximate kkt points and a proximity measure for termination
J. Dutta, K. Deb, Rupesh Tulshyan, and Ramnik Arora · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
Ju Sun, Qing Qu, and John Wright · 2015
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Complete dictionary recovery over the sphere i: Overview and the geometric picture
Ju Sun, Qing Qu, and John Wright · 2016
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Deep learning for classical japanese literature
Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, and David Ha · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Cited alongside, same era.
Pde-net: Learning pdes from data
Zichao Long, Yiping Lu, Xianzhong Ma, and Bin Dong · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Zhun Deng, Frances Ding, Cynthia Dwork, Rachel Hong, Giovanni Parmigiani, Prasad Patil, and Pragya Sur · 2020
Later among the works it cites.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright · 2018
Cited alongside, same era.
On the margin theory of feedforward neural networks
Colin Wei, Jason Lee, Qiang Liu, and Tengyu Ma · 2018
Cited alongside, same era.
Deep potential molecular dynamics: a scalable model with the accuracy of quantum mechanics
Linfeng Zhang, Jiequn Han, Han Wang, Roberto Car, and Weinan E · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
Truth or backpropaganda? an empirical investigation of deep learning theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller, and Tom Goldstein · 2019
Cited alongside, same era.
Jianfeng Lu and Stefan Steinerberger · 2020
Later among the works it cites.
Neural collapse with unconstrained features
Dustin G Mixon, Hans Parshall, and Jianzong Pi · 2020
Later among the works it cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Later among the works it cites.
Explicit regularization and implicit bias in deep network classifiers trained with the square loss
Tomaso Poggio and Qianli Liao · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
To each optimizer a norm, to each norm its generalization
Sharan Vaswani, Reza Babanezhad, Jose Gallego, Aaron Mishkin, Simon Lacoste-Julien, and Nicolas Le Roux · 2020
Later among the works it cites.
Stephan Wojtowytsch and Weinan E · 2020
Later among the works it cites.
How shrinking gradient noise helps the performance of neural networks
Zhun Deng, Jiaoyang Huang, and Kenji Kawaguchi · 2021
Closest in time.
Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training
C. Fang, H. He, Q. Long, and W. Su · 2021
Closest in time.
Neural collapse under mse loss: Proximity to and dynamics on the central path, 2021
X. Y. Han, Vardan Papyan, and David L. Donoho · 2021
Closest in time.
The power of contrast for feature learning: A theoretical analysis
Wenlong Ji, Zhun Deng, Ryumei Nakada, James Zou, and Linjun Zhang · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Closest in time.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Closest in time.
Understanding dynamics of nonlinear representation learning and its application
Kenji Kawaguchi, Linjun Zhang, and Zhun Deng · 2022
Closest in time.