Fetching the paper…
Reading the bibliography…
Within the scaling laws paradigm, which underpins the training of large neural networks like ChatGPT and Llama, we consider a supervised regression setting and establish the existance of a strong form of the model collapse phenomenon, a critical performance degradation due to synthetic data in the training corpus.
Distribution of eigenvalues for some sets of random matrices
V.A. Marčenko and Leonid Pastur · 1967
Earlier work this paper cites.
Priors for infinite networks
Radford M. Neal · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher Williams · 1996
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the loglikelihood function
H. Shimodaira · 2000
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto de Vito · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng · 2012
Earlier work this paper cites.
Subordination for the sum of two random matrices
V. Kargin · 2015
Earlier work this paper cites.
Free Probability and Random Matrices , volume 35 of Fields Institute Monographs
James A. Mingo and Roland Speicher · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Cited alongside, same era.
Asymptotics of ridge(less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Cited alongside, same era.
Covariate shift in high-dimensional random feature regression
Nilesh Tripuraneni, Ben Adlam, and Jeffrey Pennington · 2021
Cited alongside, same era.
Weighted empirical risk minimization: Sample selection bias correction based on importance sampling
Robin Vogel, Mastane Achab, Stéphan Clémençon, and Charles Tillier · 2021
Cited alongside, same era.
Generalization error rates in kernel regression: the crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
The curious decline of linguistic diversity: Training language models on synthetic text, 2023
Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel · 2023
Later among the works it cites.
Will large-scale generative models corrupt future datasets?
Ryuichiro Hataya, Han Bao, and Hiromi Arai · 2023
Later among the works it cites.
Demystifying disagreement-on-the-line in high dimensions
Donghwan Lee, Behrad Moniri, Xinmeng Huang, Edgar Dobriban, and Hamed Hassani · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Cited alongside, same era.
A solvable model of neural scaling laws, 2022
Alexander Maloney, Daniel A. Roberts, and James Sully · 2022
Cited alongside, same era.
Self-consuming generative models go mad
Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G. Baraniuk · 2023
Cited alongside, same era.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2023
Cited alongside, same era.
On the stability of iterative retraining of generative models on their own data
Quentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier Gidel · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Self-improving diffusion models with synthetic data, 2024
Sina Alemohammad, Ahmed Imtiaz Humayun, Shruti Agarwal, John Collomosse, and Richard Baraniuk · 2024
Closest in time.
Abhimanyu Dubey and et al · 2024
Closest in time.
Beyond model collapse: Scaling up with synthesized data requires reinforcement, 2024
Yunzhen Feng, Elvis Dohmatob, Pu Yang, Francois Charton, and Julia Kempe · 2024
Closest in time.
Self-consuming generative models with curated data provably optimize human preferences, 2024
Damien Ferbach, Quentin Bertrand, Avishek Joey Bose, and Gauthier Gidel · 2024
Closest in time.
Scaling laws for learning with real and surrogate data, 2024
Ayush Jain, Andrea Montanari, and Eren Sasoglu · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
How bad is training on synthetic data? a statistical analysis of language model collapse
Mohamed El Amine Seddik, Suei-Wen Chen, Soufiane Hayou, Pierre Youssef, and Merouane Debbah · 2024
Closest in time.
Ai models collapse when trained on recursively generated data
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2024
Closest in time.