Fetching the paper…
Reading the bibliography…
Researchers in empirical machine learning recently spotlighted their fears of so-called Model Collapse.
Locally asymptotically normal families of distributions. certain approximations to families of distributions and their use in the theory of estimation and testing hypotheses
Lucien Le Cam · 1960
Earlier work this paper cites.
Testing statistical hypotheses , volume 3
Erich Leo Lehmann, Joseph P Romano, and George Casella · 1986
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
One step to efficient synthetic data
Jordan Awan and Zhanrui Cai · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
On the stability of iterative retraining of generative models on their own data
Quentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier Gidel · 2023
Cited alongside, same era.
Will large-scale generative models corrupt future datasets?
Ryuichiro Hataya, Han Bao, and Hiromi Arai · 2023
Cited alongside, same era.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2023
Cited alongside, same era.
On the stability of iterative retraining of generative models on their own data
Quentin Bertrand, Joey Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier Gidel · 2024
Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data
Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Tomasz Korbak, Henry Sleight, Rajashree Agrawal, John Hughes, Dhruv Bhandarkar Pai, Andrey Gromov, Dan Roberts, Diyi Yang, David L. Donoho, and Sanmi Koyejo · 2024
Closest in time.
Scaling laws for learning with real and surrogate data
Ayush Jain, Andrea Montanari, and Eren Sasoglu · 2024
Closest in time.
Collapse or thrive? perils and promises of synthetic data in a self-generating world
Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L Donoho, and Sanmi Koyejo · 2024
Closest in time.
Lightlyssl
Lightly-AI · 2024
Closest in time.
Heat death of generative models in closed-loop learning
Matteo Marchi, Stefano Soatto, Pratik Chaudhari, and Paulo Tabuada · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model collapse demystified: The case of regression
Elvis Dohmatob, Yunzhen Feng, and Julia Kempe · 2024
Cited alongside, same era.
A tale of tails: Model collapse as a change of scaling laws
Yunzhen Feng, Elvis Dohmatob, Pu Yang, Francois Charton, and Julia Kempe · 2024
Cited alongside, same era.
Self-consuming generative models with curated data provably optimize human preferences
Damien Ferbach, Quentin Bertrand, Avishek Joey Bose, and Gauthier Gidel · 2024
Cited alongside, same era.
Self-consuming generative models go MAD
Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard Baraniuk
Cited in the paper.
Self-improving diffusion models with synthetic data
Sina Alemohammad, Ahmed Imtiaz Humayun, Shruti Agarwal, John Collomosse, and Richard Baraniuk
Cited in the paper.
Beyond model collapse: Scaling up with synthesized data requires reinforcement
Yunzhen Feng, Elvis Dohmatob, Pu Yang, Francois Charton, and Julia Kempe
Cited in the paper.
Gonzalo Martínez, Lauren Watson, Pedro Reviriego, José Alberto Hernández, Marc Juarez, and Rik Sarkar
Cited in the paper.
Closest in time.
How bad is training on synthetic data? a statistical analysis of language model collapse
Mohamed El Amine Seddik, Suei-Wen Chen, Soufiane Hayou, Pierre Youssef, and Merouane Debbah · 2024
Closest in time.
Ai models collapse when trained on recursively generated data
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal · 2024
Closest in time.