Fetching the paper…
Reading the bibliography…
Neural collapse ($\mathcal{NC}$) is a phenomenon observed in classification tasks where top-layer representations collapse into their class means, which become equinorm, equiangular and aligned with the classifiers.
Kawin Ethayarajh · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
The use of multiple measurements in taxonomic problems
Ronald A Fisher · 1936
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
The utilization of multiple measurements in problems of biological classification
C Radhakrishna Rao · 1948
Earlier work this paper cites.
Human behaviour and the principle of least effort
P Sargant Florence · 1950
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Grassmannian frames with applications to coding and communication
Thomas Strohmer and Robert W Heath Jr · 2003
Earlier work this paper cites.
Stable recovery of sparse overcomplete representations in the presence of noise
David L Donoho, Michael Elad, and Vladimir N Temlyakov · 2005
Earlier work this paper cites.
Just relax: Convex programming methods for identifying sparse signals in noise
Joel A Tropp · 2006
Earlier work this paper cites.
Traces of class/cross-class structure pervade deep learning spectra, 2020
Vardan Papyan · 2008
Earlier work this paper cites.
Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition, 2009
Daniel Jurafsky and James H Martin · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Neural collapse with unconstrained features, 2020
Dustin G. Mixon, Hans Parshall, and Jianzong Pi · 2011
Earlier work this paper cites.
Neural collapse with cross-entropy loss, 2021
Jianfeng Lu and Stefan Steinerberger · 2012
Earlier work this paper cites.
Weinan E and Stephan Wojtowytsch · 2012
Earlier work this paper cites.
Separation and concentration in deep networks, 2021
John Zarka, Florentin Guth, and Stéphane Mallat · 2012
Earlier work this paper cites.
Learning with noisy labels
Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari · 2013
Earlier work this paper cites.
Yao Zhu, Zhen Yu, Chao Zhang, Yijia Wu, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Ilya Sutskever, Simon Kornblith, and Nikhil Goyal · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The strange geometry of skip-gram with negative sampling
David Mimno and Laure Thompson · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank rnn language model, 2018
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen · 2018
Earlier work this paper cites.
An introduction to finite tight frames
Shayne FD Waldron · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Bfloat16 processing for neural networks
Neil Burgess, Jelena Milanovic, Nigel Stephens, Konstantinos Monachopoulos, and David Mansell · 2019
Earlier work this paper cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, X. Y. Han, and David L. Donoho · 2020
Earlier work this paper cites.
Tomaso Poggio and Qianli Liao · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Revealing the structure of deep neural networks via convex duality
Tolga Ergen and Mert Pilanci · 2021
Cited alongside, same era.
Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training
Cong Fang, Hangfeng He, Qi Long, and Weijie J Su · 2021
Cited alongside, same era.
A geometric analysis of neural collapse with unconstrained features, 2021
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Cited alongside, same era.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Cited alongside, same era.
Probing neural networks with t-sne, class-specific projections and a guided tour, 2021
Christopher R. Hoyt and Art B. Owen · 2021
Cited alongside, same era.
Perturbation analysis of neural collapse
Tom Tirer, Haoxiang Huang, and Jonathan Niles-Weed · 2023
Later among the works it cites.
Neural collapse in deep linear networks: From balanced to imbalanced data, 2023
Hien Dang, Tho Tran, Stanley Osher, Hung Tran-The, Nhat Ho, and Tan Nguyen · 2023
Later among the works it cites.
Memorization-dilation: Modeling neural collapse under noise
Duc Anh Nguyen, Ron Levie, Julian Lienen, Eyke Hüllermeier, and Gitta Kutyniok · 2023
Later among the works it cites.
Understanding imbalanced semantic segmentation through neural collapse
Zhisheng Zhong, Jiequan Cui, Yibo Yang, Xiaoyang Wu, Xiaojuan Qi, Xiangyu Zhang, and Jiaya Jia · 2023
Later among the works it cites.
A study of neural collapse phenomenon: Grassmannian frame, symmetry and generalization, 2023
Peifeng Gao, Qianqian Xu, Peisong Wen, Huiyang Shao, Zhiyong Yang, and Qingming Huang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Cited alongside, same era.
Tolga Ergen, Arda Sahiner, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci · 2022
Cited alongside, same era.
Feature selection with gradient descent on two-layer networks in low-rotation regimes, 2022
Matus Telgarsky · 2022
Cited alongside, same era.
Limitations of neural collapse for understanding generalization in deep learning, 2022
Like Hui, Mikhail Belkin, and Preetum Nakkiran · 2022
Cited alongside, same era.
Improved generalization bounds for transfer learning via neural collapse
Tomer Galanti, András György, and Marcus Hutter · 2022
Cited alongside, same era.
Neural collapse under MSE loss: Proximity to and dynamics on the central path
X.Y. Han, Vardan Papyan, and David L. Donoho · 2022
Cited alongside, same era.
A study of neural collapse for text classification
Jia Hui Feng, Edmund M-K Lai, and Weihua Li · 2023
Later among the works it cites.
Thomas Laurent, James H. von Brecht, and Xavier Bresson · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
Neural collapse in the intermediate hidden layers of classification neural networks, 2023
Liam Parker, Emre Onal, Anton Stengel, and Jake Intrater · 2023
Later among the works it cites.
The tunnel effect: Building data representations in deep neural networks
Wojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu, Piotr Mił oś, and Tomasz Trzcinski · 2023
Later among the works it cites.
Dynamics in deep classifiers trained with the square loss: Normalization, low rank, neural collapse, and generalization bounds
Mengjia Xu, Akshay Rangamani, Qianli Liao, Tomer Galanti, and Tomaso Poggio · 2023
Later among the works it cites.
Inducing neural collapse to a fixed hierarchy-aware frame for reducing mistake severity
Tong Liang and Jim Davis · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
BIG bench authors · 2023
Later among the works it cites.
Common crawl, 2023
Common Crawl · 2023
Later among the works it cites.
A neural collapse perspective on feature evolution in graph neural networks
Vignesh Kothapalli, Tom Tirer, and Joan Bruna · 2024
Closest in time.
Learning optimal inter-class margin adaptively for few-shot class-incremental learning via neural collapse-based meta-learning
Hang Ran, Weijun Li, Lusi Li, Songsong Tian, Xin Ning, and Prayag Tiwari · 2024
Closest in time.
Neco: Neural collapse based out-of-distribution detection, 2024
Mouïn Ben Ammar, Nacim Belkhir, Sebastian Popescu, Antoine Manzanera, and Gianni Franchi · 2024
Closest in time.
Epa: Neural collapse inspired robust out-of-distribution detector, 2024
Jiawei Zhang, Yufan Chen, Cheng Jin, Lei Zhu, and Yuantao Gu · 2024
Closest in time.
Deep neural collapse is provably optimal for the deep unconstrained features model
Peter Súkeník, Marco Mondelli, and Christoph H Lampert · 2024
Closest in time.
Neural collapse inspired semi-supervised learning with fixed classifier
Zhanxuan Hu, Yichen Wang, Hailong Ning, Yonghang Tai, and Feiping Nie · 2024
Closest in time.
Towards demystifying the generalization behaviors when neural collapse emerges, 2024
Gao Peifeng, Qianqian Xu, Yibo Yang, Peisong Wen, Huiyang Shao, Zhiyong Yang, Bernard Ghanem, and Qingming Huang · 2024
Closest in time.
A. Dubey et al. (101 additional authors) · 2024
Closest in time.
Pushing boundaries: Mixup’s influence on neural collapse, 2024
Quinn Fisher, Haoming Meng, and Vardan Papyan · 2024
Closest in time.
Average gradient outer product as a mechanism for deep neural collapse, 2024
Daniel Beaglehole, Peter Súkeník, Marco Mondelli, and Mikhail Belkin · 2024
Closest in time.
Connall Garrod and Jonathan P. Keating · 2024
Closest in time.
Neural rank collapse: Weight decay and small within-class variability yield low-rank bias, 2024
Emanuele Zangrando, Piero Deidda, Simone Brugiapaglia, Nicola Guglielmi, and Francesco Tudisco · 2024
Closest in time.
Jiachen Jiang, Jinxin Zhou, and Zhihui Zhu · 2024
Closest in time.
Can we understand plasticity through neural collapse?, 2024
Guglielmo Bonifazi, Iason Chalas, Gian Hess, and Jakub Łucki · 2024
Closest in time.
Understanding emergent abilities of language models from the loss perspective, 2024
Zhengxiao Du, Aohan Zeng, Yuxiao Dong, and Jie Tang · 2024
Closest in time.
Compression represents intelligence linearly
Yuzhen Huang, Jinghan Zhang, Zifei Shan, and Junxian He · 2024
Closest in time.
Entropy law: The story behind data compression and llm performance, 2024
Mingjia Yin, Chuhan Wu, Yufei Wang, Hao Wang, Wei Guo, Yasheng Wang, Yong Liu, Ruiming Tang, Defu Lian, and Enhong Chen · 2024
Closest in time.
Cross entropy versus label smoothing: A neural collapse perspective, 2024
Li Guo, Keith Ross, Zifan Zhao, George Andriopoulos, Shuyang Ling, Yufeng Xu, and Zixuan Dong · 2024
Closest in time.
Collapsed language models promote fairness, 2024
Jingxuan Xu, Wuyang Chen, Linyi Li, Yao Zhao, and Yunchao Wei · 2024
Closest in time.
The curse of recursion: Training on generated data makes models forget, 2024
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2024
Closest in time.
Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel A. Roberts, Diyi Yang, David L. Donoho, and Sanmi Koyejo · 2024
Closest in time.