Fetching the paper…
Reading the bibliography…
Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data.
G. W. Stewart and J.-g. Sun, “Matrix perturbation theory,” 1990
1990
Earlier work this paper cites.
D. L. Donoho and M. Elad, “Optimally sparse representation in general (nonorthogonal) dictionaries via l1 minimization,”
2003
Earlier work this paper cites.
A. Krizhevsky and G. E. Hinton, “Learning multiple layers of features from tiny images,” technical report, Citeseer, 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Earlier work this paper cites.
S. Arora, R. Ge, and A. Moitra, “New algorithms for learning incoherent and overcomplete dictionaries,” in
2014
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?,” in
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,”
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
M. Hardt and T. Ma, “Identity matters in deep learning,” in
2016
Earlier work this paper cites.
K. Kawaguchi, “Deep learning without poor local minima,” in
2016
Earlier work this paper cites.
J. Sun, Q. Qu, and J. Wright, “Complete dictionary recovery over the sphere i: Overview and the geometric picture,”
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
G. Alain and Y. Bengio, “Understanding intermediate layers using linear classifier probes,” in
2017
Earlier work this paper cites.
H. Lu and K. Kawaguchi, “Depth creates no bad local minima,”
2017
Earlier work this paper cites.
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro, “Implicit regularization in matrix factorization,” in
2017
Earlier work this paper cites.
B. Neyshabur, “Implicit regularization in deep learning,”
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L. McInnes, J. Healy, N. Saul, and L. Groberger, “Umap: Uniform manifold approximation and projection,”
2018
Earlier work this paper cites.
S. Arora, N. Cohen, and E. Hazan, “On the optimization of deep networks: Implicit acceleration by overparameterization,” in
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,” in
2018
Earlier work this paper cites.
T. Laurent and J. Brecht, “Deep linear networks with arbitrary loss: All local minima are global,” in
2018
Earlier work this paper cites.
A. K. Lampinen and S. Ganguli, “An analytic theory of generalization dynamics and transfer learning in deep linear networks,” in
2018
Earlier work this paper cites.
S. Arora, N. Cohen, N. Golowich, and W. Hu, “A convergence analysis of gradient descent for deep linear neural networks,” in
2018
Earlier work this paper cites.
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro, “The implicit bias of gradient descent on separable data,”
2018
Earlier work this paper cites.
S. S. Du, W. Hu, and J. D. Lee, “Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced,” in
2018
Earlier work this paper cites.
J. Pennington, S. Schoenholz, and S. Ganguli, “The emergence of spectral universality in deep networks,” in
2018
Earlier work this paper cites.
L. Xiao, Y. Bahri, J. Sohl-Dickstein, S. Schoenholz, and J. Pennington, “Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks,” in
2018
Earlier work this paper cites.
Z. Ji and M. Telgarsky, “Gradient descent aligns the layers of deep linear networks,” in
2018
Earlier work this paper cites.
G. Valle-Perez, C. Q. Camargo, and A. A. Louis, “Deep learning generalizes because the parameter-function map is biased towards simple functions,” in
2018
Earlier work this paper cites.
A. Esteva, A. Robicquet, B. Ramsundar, V. Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, “A guide to deep learning in healthcare,”
2019
Earlier work this paper cites.
A. Ansuini, A. Laio, J. H. Macke, and D. Zoccolan, “Intrinsic dimension of data representations in deep neural networks,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Gidel, F. Bach, and S. Lacoste-Julien, “Implicit regularization of discrete gradient dynamics in linear neural networks,” in
2019
Earlier work this paper cites.
A. M. Saxe, J. L. McClelland, and S. Ganguli, “A mathematical theory of semantic development in deep neural networks,”
2019
Cited alongside, same era.
M. S. Nacson, J. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry, “Convergence of gradient descent on separable data,” in
2019
Cited alongside, same era.
S. Arora, N. Cohen, W. Hu, and Y. Luo, “Implicit regularization in deep matrix factorization,” in
2019
Cited alongside, same era.
Z. Ji and M. Telgarsky, “The implicit bias of gradient descent on nonseparable data,” in
2019
Cited alongside, same era.
W. Hu, L. Xiao, and J. Pennington, “Provable benefit of orthogonal initialization in optimizing deep linear networks,” in
2019
Cited alongside, same era.
T. Galanti, A. György, and M. Hutter, “On the role of neural collapse in transfer learning,” in
2022
Later among the works it cites.
C. Yaras, P. Wang, Z. Zhu, L. Balzano, and Q. Qu, “Neural collapse with normalized features: A geometric analysis over the Riemannian manifold,”
2022
Later among the works it cites.
J. Zhou, X. Li, T. Ding, C. You, Q. Qu, and Z. Zhu, “On the optimization landscape of neural collapse under MSE loss: Global optimality with unconstrained features,” in
2022
Later among the works it cites.
J. Zhou, C. You, X. Li, K. Liu, S. Liu, Q. Qu, and Z. Zhu, “Are all losses created equal: A neural collapse perspective,” in
2022
Later among the works it cites.
N. S. Chatterji, P. M. Long, and P. L. Bartlett, “The interplay between implicit bias and benign overfitting in two-layer linear networks,”
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
M. Belkin, D. Hsu, S. Ma, and S. Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off,”
2019
Cited alongside, same era.
V. Papyan, X. Han, and D. L. Donoho, “Prevalence of neural collapse during the terminal phase of deep learning training,”
2020
Cited alongside, same era.
S. Guo, J. M. Alvarez, and M. Salzmann, “Expandnets: Linear over-parameterization to train compact convolutional networks,” in
2020
Cited alongside, same era.
J. Huang and H.-T. Yau, “Dynamics of deep neural networks and neural tangent hierarchy,” in
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Cited alongside, same era.
L. Hui and M. Belkin, “Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks,” in
2020
Cited alongside, same era.
E. Boursier, L. Pillaud-Vivien, and N. Flammarion, “Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs,” in
2022
Later among the works it cites.
P. Wang, H. Liu, C. Yaras, L. Balzano, and Q. Qu, “Linear convergence analysis of neural collapse with unconstrained features,” in
2022
Later among the works it cites.
K. H. R. Chan, Y. Yu, C. You, H. Qi, J. Wright, and Y. Ma, “Redunet: A white-box deep network from the principle of maximizing rate reduction,”
2022
Later among the works it cites.
Y. Ma, D. Tsao, and H.-Y. Shum, “On the principles of parsimony and self-consistency for the emergence of intelligence,”
2022
Later among the works it cites.
M. Chen, D. Y. Fu, A. Narayan, M. Zhang, Z. Song, K. Fatahalian, and C. Ré, “Perfectly balanced: Improving transfer and robustness of supervised contrastive learning,” in
2022
Later among the works it cites.
Y. Shin, “Effects of depth, width, and initialization: A convergence analysis of layer-wise training for deep linear neural networks,”
2022
Later among the works it cites.
B. Bah, H. Rauhut, U. Terstiege, and M. Westdickenberg, “Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers,”
2022
Later among the works it cites.
D. Kunin, A. Yamamura, C. Ma, and S. Ganguli, “The asymmetric maximum margin bias of quasi-homogeneous neural networks,” in
2022
Later among the works it cites.
Z. Allen-Zhu and Y. Li, “Backward feature correction: How deep learning performs deep (hierarchical) learning,” in
2023
Closest in time.
H. He and W. J. Su, “A law of data separation in deep learning,”
2023
Closest in time.
W. Masarczyk, M. Ostaszewski, E. Imani, R. Pascanu, P. Miłoś, and T. Trzcinski, “The tunnel effect: Building data representations in deep neural networks,” in
2023
Closest in time.
A. Rangamani, M. Lindegaard, T. Galanti, and T. A. Poggio, “Feature learning in deep classifiers through intermediate neural collapse,” in
2023
Closest in time.
M. Huh, H. Mobahi, R. Zhang, B. Cheung, P. Agrawal, and P. Isola, “The low-rank simplicity bias in deep networks,”
2023
Closest in time.
H. Dang, T. T. Huu, S. Osher, H. T. Tran, N. Ho, and T. M. Nguyen, “Neural collapse in deep linear networks: From balanced to imbalanced data,” in
2023
Closest in time.
P. Súkeník, M. Mondelli, and C. H. Lampert, “Deep neural collapse is provably optimal for the deep unconstrained features model,”
2023
Closest in time.
V. Kothapalli, T. Tirer, and J. Bruna, “A neural collapse perspective on feature evolution in graph neural networks,” in
2023
Closest in time.
T. Tirer, H. Huang, and J. Niles-Weed, “Perturbation analysis of neural collapse,” in
2023
Closest in time.
2023
Closest in time.
N. S. Chatterji and P. M. Long, “Deep linear networks can benignly overfit when shallow ones do,”
2023
Closest in time.
S. Frei, G. Vardi, P. Bartlett, N. Srebro, and W. Hu, “Implicit bias in leaky ReLU networks trained on high-dimensional data,” in
2023
Closest in time.
M. Xu, A. Rangamani, Q. Liao, T. Galanti, and T. Poggio, “Dynamics in deep classifiers trained with the square loss: Normalization, low rank, neural collapse, and generalization bounds,”
2023
Closest in time.
Y. Cao, D. Zou, Y. Li, and Q. Gu, “The implicit bias of batch normalization in linear models and two-layer linear convolutional neural networks,” in
2023
Closest in time.
B. Geshkovski, C. Letrouit, Y. Polyanskiy, and P. Rigollet, “The emergence of clusters in self-attention dynamics,” in
2023
Closest in time.
X. Li, S. Liu, J. Zhou, X. Lu, C. Fernandez-Granda, Z. Zhu, and Q. Qu, “Understanding and improving transfer learning of deep models via neural collapse,”
2024
Closest in time.
S. M. Kwon, Z. Zhang, D. Song, L. Balzano, and Q. Qu, “Efficient low-dimensional compression of overparameterized models,” in
2024
Closest in time.
E. M. Achour, F. Malgouyres, and S. Gerchinovitz, “The loss landscape of deep linear neural networks: a second-order analysis,”
2024
Closest in time.
R. Shwartz Ziv and Y. LeCun, “To compress or not to compress—self-supervised learning and information theory: A review,”
2024
Closest in time.
P. Li, X. Li, Y. Wang, and Q. Qu, “Neural collapse in multi-label learning with pick-all-label loss,” in
2024
Closest in time.
2024
Closest in time.
S. Wang, K. Gai, and S. Zhang, “Progressive feedforward collapse of resnet training,”
2024
Closest in time.
C. Yaras, P. Wang, L. Balzano, and Q. Qu, “Compressible dynamics in deep overparameterized low-rank learning & adaptation,” in
2024
Closest in time.
Y. Yu, S. Buchanan, D. Pai, T. Chu, Z. Wu, S. Tong, H. Bai, Y. Zhai, B. D. Haeffele, and Y. Ma, “White-box transformers via sparse rate reduction: Compression is all there is?,”
2024
Closest in time.
A. Jacot, P. Súkeník, Z. Wang, and M. Mondelli, “Wide neural networks trained with weight decay provably exhibit neural collapse,” in
2025
Closest in time.
2025
Closest in time.
V. Kothapalli and T. Tirer, “Can kernel methods explain how the data affects neural collapse?,”
2025
Closest in time.