Fetching the paper…
Reading the bibliography…
In the era of proliferation of large language and image generation models, the phenomenon of "model collapse" refers to the situation whereby as a model is trained recursively on data generated from previous generations of itself over time, its performance degrades until the model eventually becomes completely useless, i.e the model collapses.
Distribution of eigenvalues for some sets of random matrices
Marčenko, V. and Pastur, L · 1967
Earlier work this paper cites.
Priors for infinite networks
Neal, R. M · 1996
Earlier work this paper cites.
Computing with infinite networks
Williams, C · 1996
Earlier work this paper cites.
Berthier, R., Bach, F. R., and Gaillard, P · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A. and de Vito, E · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices
Bai, Z. and Silverstein, J. W. J. W · 2010
Earlier work this paper cites.
Anisotropic local laws for random matrices
Knowles, A. and Yin, J · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
Rudi, A. and Rosasco, L · 2017
Earlier work this paper cites.
Nonparametric regression using deep neural networks with relu activation function
Schmidt-Hieber, J · 2017
Earlier work this paper cites.
To understand deep learning we need to understand kernel learning
Belkin, M., Ma, S., and Mandal, S · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2018
Earlier work this paper cites.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Pillaud-Vivien, L., Rudi, A., and Bach, F. R · 2018
Earlier work this paper cites.
A continuous-time view of early stopping for least squares regression
Ali, A., Kolter, J. Z., and Tibshirani, R. J · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality
Suzuki, T · 2019
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
Bordelon, B., Canatar, A., and Pehlevan, C · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
More than a toy: Random matrix models predict how real-world neural representations generalize
Wei, A., Hu, W., and Steinhardt, J · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Self-consuming generative models go mad
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., and Baraniuk, R. G · 2023
Later among the works it cites.
High-dimensional analysis of double descent for linear regression with random projections, 2023
Bach, F · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Just interpolate: Kernel “Ridgeless” regression can generalize
Liang, T. and Rakhlin, A · 2020
Cited alongside, same era.
Self-distillation amplifies regularization in Hilbert space
Mobahi, H., Farajtabar, M., and Bartlett, P · 2020
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Spigler, S., Geiger, M., and Wyart, M · 2020
Cited alongside, same era.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Cui, H., Loureiro, B., Krzakala, F., and Zdeborova, L · 2021
Cited alongside, same era.
Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime
Nitanda, A. and Suzuki, T · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Asymptotics of ridge(less) regression under general source condition
Richards, D., Mourtada, J., and Rosasco, L · 2021
Cited alongside, same era.
On the stability of iterative retraining of generative models on their own data
Bertrand, Q., Bose, A. J., Duplessis, A., Jiralerspong, M., and Gidel, G · 2023
Later among the works it cites.
Nepotistically trained generative-ai models collapse, 2023
Bohacek, M. and Farid, H · 2023
Later among the works it cites.
Large language models suffer from their own output: An analysis of the self-consuming training loop, 2023
Briesch, M., Sobania, D., and Rothlauf, F · 2023
Later among the works it cites.
Error scaling laws for kernel classification under source and capacity conditions
Cui, H., Loureiro, B., Krzakala, F., and Zdeborová, L · 2023
Later among the works it cites.
The curious decline of linguistic diversity: Training language models on synthetic text, 2023
Guo, Y., Shang, G., Vazirgiannis, M., and Clavel, C · 2023
Later among the works it cites.
Will large-scale generative models corrupt future datasets?
Hataya, R., Bao, H., and Arai, H · 2023
Later among the works it cites.
Midjourney ai, 2023
Midjourney · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Later among the works it cites.
Data feedback loops: model-driven amplification of dataset biases
Taori, R. and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
A tale of tails: Model collapse as a change of scaling laws, 2024
Dohmatob, E., Feng, Y., Yang, P., Charton, F., and Kempe, J · 2024
Closest in time.