Fetching the paper…
Reading the bibliography…
When training deep neural networks for classification tasks, an intriguing empirical phenomenon has been widely observed in the last-layer classifiers and features, where (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and (ii) cross-example within-class variability of last-layer activations collapses to zero.
Approximation by superposition of sigmoidal functions
G Cybenko · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
Grassmannian frames with applications to coding and communication
Thomas Strohmer and Robert W Heath Jr · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Estimating unknown sparsity in compressed sensing
Miles E. Lopes · 2013
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
Ju Sun, Qing Qu, and John Wright · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Complete dictionary recovery over the sphere i: Overview and the geometric picture
Ju Sun, Qing Qu, and John Wright · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy P. Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2016
Earlier work this paper cites.
The expressive power of neural networks: a view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Earlier work this paper cites.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
Reexamining low rank matrix factorization for trace norm regularization
Carlo Ciliberto, Dimitris Stamos, and Massimiliano Pontil · 2017
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Cited alongside, same era.
Provable approximation properties for deep neural networks
Uri Shaham, Alexander Cloninger, and Ronald R Coifman · 2018
Cited alongside, same era.
A geometric analysis of phase retrieval
From symmetry to geometry: Tractable nonconvex problems
Yuqian Zhang, Qing Qu, and John Wright · 2020
Later among the works it cites.
Exploring the role of loss functions in multiclass classification
Ahmet Demirkaya, Jiasi Chen, and Samet Oymak · 2020
Later among the works it cites.
Geometric analysis of nonconvex optimization landscapes for overcomplete learning
Qing Qu, Yuexiang Zhai, Xiao Li, Yuqian Zhang, and Zhihui Zhu · 2020
Later among the works it cites.
Finding the sparsest vectors in a subspace: Theory, algorithms, and applications
Qing Qu, Zhihui Zhu, Xiao Li, Manolis C. Tsakiris, John Wright, and René Vidal · 2020
Later among the works it cites.
James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ju Sun, Qing Qu, and John Wright · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Cited alongside, same era.
The non-convex geometry of low-rank matrix optimization
Qiuwei Li, Zhihui Zhu, and Gongguo Tang · 2019
Cited alongside, same era.
Symmetry, saddle points, and global optimization landscape of nonconvex matrix factorization
Xingguo Li, Junwei Lu, Raman Arora, Jarvis Haupt, Han Liu, Zhaoran Wang, and Tuo Zhao · 2019
Cited alongside, same era.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M Lu, and Yuxin Chen · 2019
Cited alongside, same era.
Improved protein structure prediction using potentials from deep learning
Andrew W Senior, Richard Evans, John Jumper, James Kirkpatrick, Laurent Sifre, Tim Green, Chongli Qin, Augustin Žídek, Alexander WR Nelson, Alex Bridgland, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2021
Later among the works it cites.
Layer-peeled model: Toward understanding well-trained deep neural networks
Cong Fang, Hangfeng He, Qi Long, and Weijie J Su · 2021
Later among the works it cites.
Dissecting supervised constrastive learning
Florian Graf, Christoph Hofer, Marc Niethammer, and Roland Kwitt · 2021
Later among the works it cites.
Revealing the structure of deep neural networks via convex duality
Tolga Ergen and Mert Pilanci · 2021
Later among the works it cites.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Later among the works it cites.
An unconstrained layer-peeled perspective on neural collapse
Wenlong Ji, Yiping Lu, Yiliang Zhang, Zhun Deng, and Weijie J Su · 2021
Later among the works it cites.
Dynamics and neural collapse in deep classifiers trained with the square loss
Akshay Rangamani, Mengjia Xu, Andrzej Banburski, Qianli Liao, and Tomaso Poggio · 2021
Later among the works it cites.
Neural collapse in deep homogeneous claaifiers and the role of weight decay
Florian Graf, Christoph Hofer, Marc Niethammer, and Roland Kwitt · 2021
Later among the works it cites.
Distribution of classification margins: Are all data equal?
Andrzej Banburski, Fernanda De La Torre, Nishka Pant, Ishana Shastri, and Tomaso Poggio · 2021
Later among the works it cites.
Why do better loss functions lead to less transferable features?
Simon Kornblith, Ting Chen, Honglak Lee, and Mohammad Norouzi · 2021
Later among the works it cites.
For manifold learning, deep neural networks can be locality sensitive hash functions
Nishanth Dikkala, Gal Kaplun, and Rina Panigrahy · 2021
Later among the works it cites.
Neural collapse under MSE loss: Proximity to and dynamics on the central path
X.Y. Han, Vardan Papyan, and David L. Donoho · 2022
Closest in time.
Extended unconstrained features model for exploring deep neural collapse
Tom Tirer and Joan Bruna · 2022
Closest in time.
Limitations of neural collapse for understanding generalization in deep learning
Like Hui, Mikhail Belkin, and Preetum Nakkiran · 2022
Closest in time.
On the role of neural collapse in transfer learning
Tomer Galanti, András György, and Marcus Hutter · 2022
Closest in time.
Nearest class-center simplification through intermediate layers
Ido Ben-Shaul and Shai Dekel · 2022
Closest in time.