Fetching the paper…
Reading the bibliography…
Deep classifier neural networks enter the terminal phase of training (TPT) when training error reaches zero and tend to exhibit intriguing Neural Collapse (NC) properties.
On stochastic differential equations
Kiyosi Itô · 1951
Earlier work this paper cites.
The fritz john necessary optimality conditions in the presence of equality and inequality constraints
Olvi L Mangasarian and Stan Fromovitz · 1967
Earlier work this paper cites.
Regular polytopes
Harold Scott Macdonald Coxeter · 1973
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel · 1974
Earlier work this paper cites.
Stochastic differential equations
Nicolaas G Van Kampen · 1976
Earlier work this paper cites.
Influence functions for proportional hazards regression
Nancy Reid and Helene Crépeau · 1985
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
A general weight matrix formulation using optimal control
Oluseyi Farotimi, Amir Dembo, and Thomas Kailath · 1991
Earlier work this paper cites.
Divergence measures based on the shannon entropy
Jianhua Lin · 1991
Earlier work this paper cites.
Stochastic differential equations
Peter E Kloeden and Eckhard Platen · 1992
Earlier work this paper cites.
Learning many related tasks at the same time with backpropagation
Rich Caruana · 1994
Earlier work this paper cites.
Interior point methods in semidefinite programming with applications to combinatorial optimization
Farid Alizadeh · 1995
Earlier work this paper cites.
An introduction to the kalman filter
Greg Welch, Gary Bishop, et al · 1995
Earlier work this paper cites.
Fokker-Planck Equation , pp. 63–95
Hannes Risken · 1996
Earlier work this paper cites.
Robustness of the black and scholes formula
Nicole El Karoui, Monique Jeanblanc-Picquè, and Steven E. Shreve · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira · 2000
Earlier work this paper cites.
C4. 5, class imbalance, and cost sensitivity: why under-sampling beats over-sampling
Chris Drummond, Robert C Holte, et al · 2003
Earlier work this paper cites.
On cones of nonnegative quadratic functions
Jos F Sturm and Shuzhong Zhang · 2003
Earlier work this paper cites.
Learning with matrix factorizations
Nathan Srebro · 2004
Earlier work this paper cites.
Applied linear regression , volume 528
Sanford Weisberg · 2005
Earlier work this paper cites.
Training cost-sensitive neural networks with methods addressing the class imbalance problem
Zhi-Hua Zhou and Xu-Ying Liu · 2005
Earlier work this paper cites.
Analysis of representations for domain adaptation
Shai Ben-David, John Blitzer, Koby Crammer, and Fernando Pereira · 2006
Earlier work this paper cites.
A kernel method for the two-sample-problem
Arthur Gretton, Karsten Borgwardt, Malte Rasch, Bernhard Schölkopf, and Alex Smola · 2006
Earlier work this paper cites.
Correcting sample selection bias by unlabeled data
Jiayuan Huang, Arthur Gretton, Karsten Borgwardt, Bernhard Schölkopf, and Alex Smola · 2006
Earlier work this paper cites.
Learning bounds for domain adaptation
John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman · 2007
Earlier work this paper cites.
An introduction to frames
Jelena Kovačević, Amina Chebira, et al · 2008
Earlier work this paper cites.
Domain adaptation with multiple sources
Yishay Mansour, Mehryar Mohri, and Afshin Rostamizadeh · 2008
Earlier work this paper cites.
Interior-point methods for optimization
Arkadi S Nemirovski and Michael J Todd · 2008
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
P-A Absil, Robert Mahony, and Rodolphe Sepulchre · 2009
Earlier work this paper cites.
Learning from imbalanced data
Haibo He and Edwardo A Garcia · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2010
Earlier work this paper cites.
Robust statistics
Peter J Huber · 2011
Earlier work this paper cites.
The statistical analysis of failure time data
John D Kalbfleisch and Ross L Prentice · 2011
Earlier work this paper cites.
Deep learning of representations for unsupervised and transfer learning
Yoshua Bengio · 2012
Earlier work this paper cites.
Karush-kuhn-tucker conditions
Geoff Gordon and Ryan Tibshirani · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Group invariant scattering
Stéphane Mallat · 2012
Earlier work this paper cites.
Invariant scattering convolution networks
Joan Bruna and Stéphane Mallat · 2013
Earlier work this paper cites.
Approximate kkt points and a proximity measure for termination
Joydeep Dutta, Kalyanmoy Deb, Rupesh Tulshyan, and Ramnik Arora · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Rotation, scaling and deformation invariant scattering for texture discrimination
Laurent Sifre and Stéphane Mallat · 2013
Earlier work this paper cites.
Domain adaptation under target and conditional shift
Kun Zhang, Bernhard Schölkopf, Krikamol Muandet, and Zhikun Wang · 2013
Earlier work this paper cites.
Stochastic differential equations and diffusion processes
Nobuyuki Ikeda and Shinzo Watanabe · 2014
Earlier work this paper cites.
Flexible transfer learning under support and model shift
Xuezhi Wang and Jeff Schneider · 2014
Earlier work this paper cites.
Mean-normalized stochastic gradient for large-scale deep learning
Simon Wiesler, Alexander Richard, Ralf Schlüter, and Hermann Ney · 2014
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Deep neural networks with random gaussian weights: A universal classification strategy?
Raja Giryes, Guillermo Sapiro, and Alex M Bronstein · 2016
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Learning deep representation for imbalanced classification
Chen Huang, Yining Li, Chen Change Loy, and Xiaoou Tang · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Large-margin softmax loss for convolutional neural networks
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Cited alongside, same era.
A survey of transfer learning
Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang · 2016
Cited alongside, same era.
A discriminative feature learning approach for deep face recognition
Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao · 2016
Cited alongside, same era.
A theoretical analysis of contrastive unsupervised representation learning
Nikunj Saunshi, Orestis Plevrakis, Sanjeev Arora, Mikhail Khodak, and Hrishikesh Khandeparkar · 2019
Later among the works it cites.
Mean-field analysis of batch normalization
Mingwei Wei, James Stokes, and David J Schwab · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised deep embedding for clustering analysis
Junyuan Xie, Ross Girshick, and Ali Farhadi · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Geometric deep learning: going beyond euclidean data
Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst · 2017
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Cited alongside, same era.
On loss functions for deep neural networks in classification
Katarzyna Janocha and Wojciech Marian Czarnecki · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Sphereface: Deep hypersphere embedding for face recognition
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song · 2017
Cited alongside, same era.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Deep networks from the principle of rate reduction
Kwan Ho Ryan Chan, Yaodong Yu, Chong You, Haozhi Qi, John Wright, and Yi Ma · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Exploring the role of loss functions in multiclass classification
Ahmet Demirkaya, Jiasi Chen, and Samet Oymak · 2020
Later among the works it cites.
Does learning require memorization? a short tale about a long tail
Vitaly Feldman · 2020
Later among the works it cites.
Recent advances in deep learning theory
Fengxiang He and Dacheng Tao · 2020
Later among the works it cites.
The surprising simplicity of the early-time learning dynamics of neural networks
Wei Hu, Lechao Xiao, Ben Adlam, and Jeffrey Pennington · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2020
Later among the works it cites.
A survey on contrastive self-supervised learning
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon · 2020
Later among the works it cites.
Self-supervised visual feature learning with deep neural networks: A survey
Longlong Jing and Yingli Tian · 2020
Later among the works it cites.
Supervised contrastive learning
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan · 2020
Later among the works it cites.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2020
Later among the works it cites.
Neural collapse with cross-entropy loss
Jianfeng Lu and Stefan Steinerberger · 2020
Later among the works it cites.
Neural collapse with unconstrained features
Dustin G Mixon, Hans Parshall, and Jianzong Pi · 2020
Later among the works it cites.
Interpretable machine learning
Christoph Molnar · 2020
Later among the works it cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Later among the works it cites.
Explicit regularization and implicit bias in deep network classifiers trained with the square loss
Tomaso Poggio and Qianli Liao · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
Circle loss: A unified perspective of pair similarity optimization
Yifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang, Liang Zheng, Zhongdao Wang, and Yichen Wei · 2020
Later among the works it cites.
On the theory of transfer learning: The importance of task diversity
Nilesh Tripuraneni, Michael Jordan, and Chi Jin · 2020
Later among the works it cites.
Stephan Wojtowytsch et al · 2020
Later among the works it cites.
Learning diverse and discriminative representations via the principle of maximal coding rate reduction
Yaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song, and Yi Ma · 2020
Later among the works it cites.
Separation and concentration in deep networks
John Zarka, Florentin Guth, and Stéphane Mallat · 2020
Later among the works it cites.
Revealing the structure of deep neural networks via convex duality
Tolga Ergen and Mert Pilanci · 2021
Later among the works it cites.
Tolga Ergen, Arda Sahiner, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci · 2021
Later among the works it cites.
Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training
Cong Fang, Hangfeng He, Qi Long, and Weijie J Su · 2021
Later among the works it cites.
On the role of neural collapse in transfer learning
Tomer Galanti, András György, and Marcus Hutter · 2021
Later among the works it cites.
Neural collapse under mse loss: Proximity to and dynamics on the central path
XY Han, Vardan Papyan, and David L Donoho · 2021
Later among the works it cites.
Model complexity of deep learning: A survey
Xia Hu, Lingyang Chu, Jian Pei, Weiqing Liu, and Jiang Bian · 2021
Later among the works it cites.
An unconstrained layer-peeled perspective on neural collapse
Wenlong Ji, Yiping Lu, Yiliang Zhang, Zhun Deng, and Weijie J Su · 2021
Later among the works it cites.
Why do better loss functions lead to less transferable features?
Simon Kornblith, Ting Chen, Honglak Lee, and Mohammad Norouzi · 2021
Later among the works it cites.
On the validity of modeling sgd with stochastic differential equations (sdes)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
Neural collapse under cross-entropy loss
Jianfeng Lu and Stefan Steinerberger · 2021
Later among the works it cites.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
Charles H Martin and Michael W Mahoney · 2021
Later among the works it cites.
The loss surface of deep linear networks viewed through the algebraic geometry lens
Dhagash Mehta, Tianran Chen, Tingting Tang, and Jonathan D Hauenstein · 2021
Later among the works it cites.
Regular polytope networks
Federico Pernici, Matteo Bruni, Claudio Baecchi, and Alberto Del Bimbo · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
Feng Wang and Huaping Liu · 2021
Later among the works it cites.
Incremental learning via rate reduction
Ziyang Wu, Christina Baek, Chong You, and Yi Ma · 2021
Later among the works it cites.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Later among the works it cites.
Nearest class-center simplification through intermediate layers
Ido Ben-Shaul and Shai Dekel · 2022
Closest in time.
Redunet: A white-box deep network from the principle of maximizing rate reduction
Kwan Ho Ryan Chan, Yaodong Yu, Chong You, Haozhi Qi, John Wright, and Yi Ma · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
A note on the implicit bias towards minimal depth of deep neural networks
Tomer Galanti · 2022
Closest in time.
Improved generalization bounds for transfer learning via neural collapse
Tomer Galanti, András György, and Marcus Hutter · 2022
Closest in time.
Limitations of neural collapse for understanding generalization in deep learning
Like Hui, Mikhail Belkin, and Preetum Nakkiran · 2022
Closest in time.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Closest in time.
Extended unconstrained features model for exploring deep neural collapse
Tom Tirer and Joan Bruna · 2022
Closest in time.
Do we really need a learnable classifier at the end of deep neural network?
Yibo Yang, Liang Xie, Shixiang Chen, Xiangtai Li, Zhouchen Lin, and Dacheng Tao · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Closest in time.