Fetching the paper…
Reading the bibliography…
We provide the first global optimization landscape analysis of $Neural\;Collapse$ -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Approximation by superposition of sigmoidal functions
G Cybenko · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Characterization of the subdifferential of some matrix norms
G Alistair Watson · 1992
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Mnist handwritten digit database. at&t labs, 2010
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey Hinton · 2010
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Finding a sparse vector in a subspace: Linear sparsity using alternating directions
Qing Qu, Ju Sun, and John Wright · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
Ju Sun, Qing Qu, and John Wright · 2015
Earlier work this paper cites.
A nonconvex optimization framework for low rank matrix estimation
Tuo Zhao, Zhaoran Wang, and Han Liu · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Complete dictionary recovery over the sphere I: Overview and the geometric picture
Ju Sun, Qing Qu, and John Wright · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2016
Earlier work this paper cites.
Deep residual networks with exponential linear unit
Anish Shah, Eashan Kadam, Hena Shah, Sameer Shinde, and Sandip Shingade · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep neural networks for youtube recommendations
Paul Covington, Jay Adams, and Emre Sargin · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
The expressive power of neural networks: a view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Earlier work this paper cites.
Complete dictionary recovery over the sphere II: Recovery by Riemannian trust
Ju Sun, Qing Qu, and John Wright · 2017
Earlier work this paper cites.
Convolutional phase retrieval
Qing Qu, Yuqian Zhang, Yonina Eldar, and John Wright · 2017
Earlier work this paper cites.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Earlier work this paper cites.
Reexamining low rank matrix factorization for trace norm regularization
Carlo Ciliberto, Dimitris Stamos, and Massimiliano Pontil · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2018
Cited alongside, same era.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Cited alongside, same era.
Learning deep models: Critical points and local openness
Maher Nouiehed and Meisam Razaviyayn · 2018
Cited alongside, same era.
Adding one neuron can eliminate all bad local minima
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Blake E Woodworth, Suriya Gunasekar, Jason D Lee, Nati Srebro, and Daniel Soudry · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
Rethinking bias-variance trade-off for generalization of neural networks
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma · 2020
Later among the works it cites.
Traces of class/cross-class structure pervade deep learning spectra
Vardan Papyan · 2020
Later among the works it cites.
Separation and concentration in deep networks
John Zarka, Florentin Guth, and Stéphane Mallat · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shiyu Liang, Ruoyu Sun, Jason D Lee, and R Srikant · 2018
Cited alongside, same era.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Cited alongside, same era.
Provable approximation properties for deep neural networks
Uri Shaham, Alexander Cloninger, and Ronald R Coifman · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel · 2018
Cited alongside, same era.
The non-convex geometry of low-rank matrix optimization
Qiuwei Li, Zhihui Zhu, and Gongguo Tang · 2018
Cited alongside, same era.
Dustin G Mixon, Hans Parshall, and Jianzong Pi · 2020
Later among the works it cites.
Neural collapse with cross-entropy loss
Jianfeng Lu and Stefan Steinerberger · 2020
Later among the works it cites.
E Weinan and Stephan Wojtowytsch · 2020
Later among the works it cites.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2020
Later among the works it cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Benign overfitting and noisy features
Zhu Li, Weijie Su, and Dino Sejdinovic · 2020
Later among the works it cites.
Yaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song, and Yi Ma · 2020
Later among the works it cites.
Deep networks from the principle of rate reduction
Kwan Ho Ryan Chan, Yaodong Yu, Chong You, Haozhi Qi, John Wright, and Yi Ma · 2020
Later among the works it cites.
From symmetry to geometry: Tractable nonconvex problems
Yuqian Zhang, Qing Qu, and John Wright · 2020
Later among the works it cites.
Another step toward demystifying deep neural networks
Michael Elad, Dror Simon, and Aviad Aberdam · 2020
Later among the works it cites.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Later among the works it cites.
Robust recovery via implicit bias of discrepant learning rates for double over-parameterization
Chong You, Zhihui Zhu, Qing Qu, and Yi Ma · 2020
Later among the works it cites.
Optimization for deep learning: An overview
Ruo-Yu Sun · 2020
Later among the works it cites.
The global landscape of neural networks: An overview
Ruoyu Sun, Dawei Li, Shiyu Liang, Tian Ding, and Rayadurgam Srikant · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Andrea Montanari and Yiqiao Zhong · 2020
Later among the works it cites.
Deep networks and the multiple manifold problem
Sam Buchanan, Dar Gilboa, and John Wright · 2020
Later among the works it cites.
Recent theoretical advances in non-convex optimization
Marina Danilova, Pavel Dvurechensky, Alexander Gasnikov, Eduard Gorbunov, Sergey Guminov, Dmitry Kamzolov, and Innokentiy Shibaev · 2020
Later among the works it cites.
Finding the sparsest vectors in a subspace: Theory, algorithms, and applications
Qing Qu, Zhihui Zhu, Xiao Li, Manolis C. Tsakiris, John Wright, and René Vidal · 2020
Later among the works it cites.
Geometric analysis of nonconvex optimization landscapes for overcomplete learning
Qing Qu, Yuexiang Zhai, Xiao Li, Yuqian Zhang, and Zhihui Zhu · 2020
Later among the works it cites.
Exact recovery of multichannel sparse blind deconvolution via gradient descent
Qing Qu, Xiao Li, and Zhihui Zhu · 2020
Later among the works it cites.
Short and sparse deconvolution — a geometric approach
Yenson Lau, Qing Qu, Han-Wen Kuo, Pengcheng Zhou, Yuqian Zhang, and John Wright · 2020
Later among the works it cites.
Adversarial robustness of supervised sparse coding
Jeremias Sulam, Ramchandran Muthumukar, and Raman Arora · 2020
Later among the works it cites.
Deep isometric learning for visual recognition
Haozhi Qi, Chong You, Xiaolong Wang, Yi Ma, and Jitendra Malik · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Self-supervised visual feature learning with deep neural networks: A survey
Longlong Jing and Yingli Tian · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola · 2020
Later among the works it cites.
Intriguing properties of contrastive losses
Ting Chen and Lala Li · 2020
Later among the works it cites.
Layer-peeled model: Toward understanding well-trained deep neural networks
Cong Fang, Hangfeng He, Qi Long, and Weijie J Su · 2021
Closest in time.
A local convergence theory for mildly over-parameterized two-layer neural network
Mo Zhou, Rong Ge, and Chi Jin · 2021
Closest in time.
Gradient methods never overfit on separable data
Ohad Shamir · 2021
Closest in time.
Inductive bias of multi-channel linear convolutional networks with bounded weight norm
Meena Jagadeesan, Ilya Razenshteyn, and Suriya Gunasekar · 2021
Closest in time.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2021
Closest in time.
Kronecker-factored quasi-newton methods for convolutional neural networks
Yi Ren and Donald Goldfarb · 2021
Closest in time.
Orthogonal over-parameterized training, 2021
Weiyang Liu, Rongmei Lin, Zhen Liu, James M. Rehg, Liam Paull, Li Xiong, Le Song, and Adrian Weller · 2021
Closest in time.
Convolutional normalization: Improving deep convolutional network robustness and training, 2021
Sheng Liu, Xiao Li, Yuexiang Zhai, Chong You, Zhihui Zhu, Carlos Fernandez-Granda, and Qing Qu · 2021
Closest in time.