Fetching the paper…
Reading the bibliography…
In this paper, we introduce the \textit{Layer-Peeled Model}, a nonconvex yet analytically tractable optimization program, in a quest to better understand deep neural networks that are trained for a sufficiently long time.
\JournalTitle
AR Webb, D Lowe, The optimised internal representation of multilayer classifier networks performs nonlinear discriminant analysis · 1990
Earlier work this paper cites.
\JournalTitle
T Strohmer, RW Heath, Grassmannian frames with applications to coding and communication · 2003
Earlier work this paper cites.
\JournalTitle
JF Sturm, S Zhang, On cones of nonnegative quadratic functions · 2003
Earlier work this paper cites.
pp. 57–60 (2006)
E Hovy, M Marcus, M Palmer, L Ramshaw, R Weischedel, Ontonotes: the 90% solution in Proceedings of the human language technology conference of the NAACL, Companion Volume: Short Papers · 2006
Earlier work this paper cites.
F Bach, J Mairal, J Ponce, Convex sparse matrix factorizations · 2008
Earlier work this paper cites.
A Krizhevsky, Master’s thesis (University of Toronto) (2009)
2009
Earlier work this paper cites.
pp. 1532–1543 (2014)
J Pennington, R Socher, CD Manning, Glove: Global vectors for word representation in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) · 2014
Earlier work this paper cites.
\JournalTitle
Y LeCun, Y Bengio, G Hinton, Deep learning · 2015
Earlier work this paper cites.
pp. 448–456 (2015)
S Ioffe, C Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift in International Conference on Machine Learning · 2015
Earlier work this paper cites.
\JournalTitle
D Silver, et al., Mastering the game of go with deep neural networks and tree search · 2016
Earlier work this paper cites.
(IEEE), pp. 4368–4374 (2016)
S Wang, et al., Training deep neural networks on imbalanced data sets in 2016 international joint conference on neural networks (IJCNN) · 2016
Earlier work this paper cites.
pp. 5375–5384 (2016)
C Huang, Y Li, CC Loy, X Tang, Learning deep representation for imbalanced classification in Proceedings of the IEEE conference on computer vision and pattern recognition · 2016
Earlier work this paper cites.
B Zhou, A Khosla, A Lapedriza, A Torralba, A Oliva, Places: An image database for deep scene understanding · 2016
Earlier work this paper cites.
pp. 770–778 (2016)
K He, X Zhang, S Ren, J Sun, Deep residual learning for image recognition in Proceedings of the IEEE conference on computer vision and pattern recognition · 2016
Earlier work this paper cites.
\JournalTitle
A Krizhevsky, I Sutskever, GE Hinton, Imagenet classification with deep convolutional neural networks · 2017
Earlier work this paper cites.
\JournalTitle
D Yarotsky, Error bounds for approximations with deep ReLU networks · 2017
Earlier work this paper cites.
\JournalTitle
P Bartlett, D Foster, M Telgarsky, Spectrally-normalized margin bounds for neural networks · 2017
Earlier work this paper cites.
\JournalTitle
K Madasamy, M Ramaswami, Data imbalance and classifiers: impact and solutions from a big data perspective · 2017
Earlier work this paper cites.
arXiv:1708.07747 (15 Sep 2017)
H Xiao, K Rasul, R Vollgraf, Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms · 2017
Earlier work this paper cites.
\JournalTitle
D Soudry, E Hoffer, MS Nacson, S Gunasekar, N Srebro, The implicit bias of gradient descent on separable data · 2018
Earlier work this paper cites.
\JournalTitle
S Mei, A Montanari, PM Nguyen, A mean field view of the landscape of two-layer neural networks · 2018
Earlier work this paper cites.
(PMLR), pp. 3325–3334 (2018)
S Ma, R Bassily, M Belkin, The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning in International Conference on Machine Learning · 2018
Cited alongside, same era.
\JournalTitle
M Buda, A Maki, MA Mazurowski, A systematic study of the class imbalance problem in convolutional neural networks · 2018
Cited alongside, same era.
pp. 77–91 (2018)
J Buolamwini, T Gebru, Gender shades: Intersectional accuracy disparities in commercial gender classification in Conference on fairness, accountability and transparency · 2018
Cited alongside, same era.
J Zou, L Schiebinger, AI can be sexist and racist—it’s time to make it fair (2018)
2018
Cited alongside, same era.
\JournalTitle
L Bottou, FE Curtis, J Nocedal, Optimization methods for large-scale machine learning · 2018
Cited alongside, same era.
pp. 689–699 (2018)
C Fang, CJ Li, Z Lin, T Zhang, Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator in Advances in Neural Information Processing Systems · 2018
\JournalTitle
S Oymak, M Soltanolkotabi, Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks · 2020
Later among the works it cites.
\JournalTitle
Y Yu, KHR Chan, C You, C Song, Y Ma, Learning diverse and discriminative representations via the principle of maximal coding rate reduction · 2020
Later among the works it cites.
arXiv:2007.00028 (10 Sep 2020)
O Shamir, Gradient methods never overfit on separable data · 2020
Later among the works it cites.
(PMLR), Vol. 119, pp. 1597–1607 (2020)
T Chen, S Kornblith, M Norouzi, G Hinton, A simple framework for contrastive learning of visual representations in Proceedings of the 37th International Conference on Machine Learning · 2020
Later among the works it cites.
arXiv:1904.04326 (21 Feb 2020)
W E, C Ma, L Wu, A comparative analysis of the optimization and generalization property of two-layer neural network and random feature models under gradient descent dynamics · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
\JournalTitle
JM Johnson, TM Khoshgoftaar, Survey on deep learning with class imbalance · 2019
Cited alongside, same era.
pp. 2388–2464 (2019)
Z Allen-Zhu, Y Li, Z Song, A convergence theory for deep learning via over-parameterization in International Conference on Machine Learning · 2019
Cited alongside, same era.
pp. 14601–14610 (2019)
R Kuditipudi, et al., Explaining landscape connectivity of low-cost solutions for multilayer nets in Advances in Neural Information Processing Systems · 2019
Cited alongside, same era.
arXiv:1904.05526 (15 Apr 2019)
J Fan, C Ma, Y Zhong, A selective overview of deep learning · 2019
Cited alongside, same era.
arXiv:1912.08957 (19 Dec 2019)
R Sun, Optimization for deep learning: theory and algorithms · 2019
Cited alongside, same era.
\JournalTitle
M Belkin, D Hsu, S Ma, S Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off · 2019
Cited alongside, same era.
\JournalTitle
T Poggio, A Banburski, Q Liao, Theoretical issues in deep networks · 2020
Later among the works it cites.
\JournalTitle
J Sirignano, K Spiliopoulos, Mean field analysis of neural networks: A central limit theorem · 2020
Later among the works it cites.
arXiv:2007.01452 (3 July 2020)
C Fang, JD Lee, P Yang, T Zhang, Modeling from features: a mean-field framework for over-parameterized deep neural networks · 2020
Later among the works it cites.
arXiv:2004.06977 (15 Apr 2020)
B Shi, WJ Su, MI Jordan, On learning rates and Schrödinger operators · 2020
Later among the works it cites.
arXiv:2012.13982 (27 Dec 2020)
C Fang, H Dong, T Zhang, Mathematical models of overparameterized neural networks · 2020
Later among the works it cites.
arXiv:2012.10931 (20 Dec 2020)
F He, D Tao, Recent advances in deep learning theory · 2020
Later among the works it cites.
\JournalTitle
T Liang, A Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize · 2020
Later among the works it cites.
\JournalTitle
PL Bartlett, PM Long, G Lugosi, A Tsigler, Benign overfitting in linear regression · 2020
Later among the works it cites.
Z Li, W Su, D Sejdinovic, Benign overfitting and noisy features · 2020
Later among the works it cites.
arXiv:2011.11619 (23 Nov 2020)
DG Mixon, H Parshall, J Pi, Neural collapse with unconstrained features · 2020
Later among the works it cites.
arXiv:2012.05420 (19 Dec 2020)
W E, S Wojtowytsch, On the emergence of tetrahedral symmetry in the final and penultimate layers of neural network classifiers · 2020
Later among the works it cites.
arXiv preprint arXiv:2002.09773 (22 Feb 2020)
T Ergen, M Pilanci, Convex duality of deep neural networks · 2020
Later among the works it cites.
arXiv:2101.00072 (31 Dec 2020)
T Poggio, Q Liao, Explicit regularization and implicit bias in deep network classifiers trained with the square loss · 2020
Later among the works it cites.
arXiv:2006.11477 (22 Oct 2020)
A Baevski, H Zhou, A Mohamed, M Auli, wav2vec 2.0: A framework for self-supervised learning of speech representations · 2020
Later among the works it cites.
arXiv:2004.11362 (10 Dec 2020)
P Khosla, et al., Supervised contrastive learning · 2020
Later among the works it cites.
arXiv:2012.08465 (18 Jan 2021)
J Lu, S Steinerberger, Neural collapse with cross-entropy loss · 2021
Closest in time.