Fetching the paper…
Reading the bibliography…
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance.
Principles of Mathematical Analysis
Rudin, W · 1976
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Face detection based on receptive field enhanced multi-task cascaded convolutional neural networks
Li, X., Yang, Z., and Wu, H · 2020
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space
Mobahi, H., Farajtabar, M., and Bartlett, P · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2021
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Earlier work this paper cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al · 2022
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., hsin Chi, E. H., Xia, F., Le, Q., and Zhou, D · 2022
Earlier work this paper cites.
Zhu, X., Guan, J., Huang, M., and Liu, J · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Self-consuming generative models go mad
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., and Baraniuk, R. G · 2023
Earlier work this paper cites.
Picor: Multi-task deep reinforcement learning with policy correction
Bai, F., Zhang, H., Tao, T., Wu, Z., Wang, Y., and Xu, B · 2023
Earlier work this paper cites.
On the stability of iterative retraining of generative models on their own data
Bertrand, Q., Bose, A. J., Duplessis, A., Jiralerspong, M., and Gidel, G · 2023
Earlier work this paper cites.
Large language models suffer from their own output: An analysis of the self-consuming training loop
Briesch, M., Sobania, D., and Rothlauf, F · 2023
Earlier work this paper cites.
Tinystories: How small can language models be and still speak coherent english?
Eldan, R. and Li, Y · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Earlier work this paper cites.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al · 2023
Earlier work this paper cites.
Openassistant conversations - democratizing large language model alignment
Kopf, A., Kilcher, Y., von Rutte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., Shahul, E., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., and Mattick, A · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Normalization enhances generalization in visual reinforcement learning
Li, L., Lyu, J., Ma, G., Wang, Z., Yang, Z., Li, X., and Li, Z · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2023
Cited alongside, same era.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Cited alongside, same era.
Paloma: A benchmark for evaluating language model fit
Magnusson, I., Bhagia, A., Hofmann, V., Soldaini, L., Jha, A., Tafjord, O., Schwenk, D., Walsh, P., Elazar, Y., Lo, K., Groeneveld, D., Beltagy, I., Hajishirzi, H., Smith, N. A., Richardson, K., and Dodge, J · 2023
Cited alongside, same era.
Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving
Jia, X., Yang, Z., Li, Q., Zhang, Z., and Yan, J · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Li, Q., Jia, X., Wang, S., and Yan, J · 2024
Closest in time.
Rho-1: Not all tokens are what you need
Lin, Z., Gou, Z., Gong, Y., Liu, X., Shen, Y., Xu, R., Lin, C., Yang, Y., Jiao, J., Duan, N., et al · 2024
Closest in time.
Best practices and lessons learned on synthetic data for language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards understanding the interplay of generative artificial intelligence and the internet
Martínez, G., Watson, L., Reviriego, P., Hernández, J. A., Juarez, M., and Sarkar, R · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P. J., Harrison, J., Lee, J., Xu, K., et al · 2023
Cited alongside, same era.
Data selection for language models via importance resampling
Xie, S. M., Santurkar, S., Ma, T., and Liang, P. S · 2023
Cited alongside, same era.
Llm4drive: A survey of large language models for autonomous driving
Yang, Z., Jia, X., Li, H., and Yan, J · 2023
Cited alongside, same era.
Gobigger: A scalable platform for cooperative-competitive multi-agent interactive simulation
Zhang, M., Zhang, S., Yang, Z., Chen, L., Zheng, J., Yang, C., Li, C., Zhou, H., Niu, Y., and Liu, Y · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta · 2024
Cited alongside, same era.
A survey on data selection for language models
Albalak, A., Elazar, Y., Xie, S. M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., et al · 2024
Cited alongside, same era.
Perplexed by perplexity: Perplexity-based data pruning with small reference models
Ankner, Z., Blakeney, C., Sreenivasan, K., Marion, M., Leavitt, M. L., and Paul, M · 2024
Cited alongside, same era.
Liu, R., Wei, J., Liu, F., Si, C., Zhang, Y., Rao, J., Zheng, S., Peng, D., Yang, D., Zhou, D., et al · 2024
Closest in time.
Rephrasing the web: A recipe for compute and data-efficient language modeling
Maini, P., Seto, S., Bai, R. H., Grangier, D., Zhang, Y., and Jaitly, N · 2024
Closest in time.
Lightzero: A unified benchmark for monte carlo tree search in general sequential decision scenarios
Niu, Y., Pu, Y., Yang, Z., Li, X., Zhou, T., Ren, J., Hu, S., Li, H., and Liu, Y · 2024
Closest in time.
Unizero: Generalized and efficient planning with scalable latent world models
Pu, Y., Niu, Y., Yang, Z., Ren, J., Li, H., and Liu, Y · 2024
Closest in time.
How bad is training on synthetic data? a statistical analysis of language model collapse
Seddik, M. E. A., Chen, S.-W., Hayou, S., Youssef, P., and Debbah, M · 2024
Closest in time.
Ai models collapse when trained on recursively generated data
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y · 2024
Closest in time.
Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2024
Closest in time.
Large language models for data annotation and synthesis: A survey
Tan, Z., Li, D., Wang, S., Beigi, A., Jiang, B., Bhattacharjee, A., Karami, M., Li, J., Cheng, L., and Liu, H · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T., Wu, Y., Le, Q., He, H., and Luong, T · 2024
Closest in time.
Bootstrapping llm-based task-oriented dialogue agents via self-talk
Ulmer, D., Mansimov, E., Lin, K., Sun, J., Gao, X., and Zhang, Y · 2024
Closest in time.
Less: Selecting influential data for targeted instruction tuning
Xia, M., Malladi, S., Gururangan, S., Arora, S., and Chen, D · 2024
Closest in time.
Ultramedical: Building specialized generalists in biomedicine
Zhang, K., Zeng, S., Hua, E., Ding, N., Chen, Z.-R., Ma, Z., Li, H., Cui, G., Qi, B., Zhu, X., et al · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Zheng, Y., Zhang, R., Zhang, J., Ye, Y., Luo, Z., Feng, Z., and Ma, Y · 2024
Closest in time.
PaD: Program-aided distillation can teach small models reasoning better than chain-of-thought fine-tuning
Zhu, X., Qi, B., Zhang, K., Long, X., Lin, Z., and Zhou, B · 2024
Closest in time.
Rat: Adversarial attacks on deep reinforcement agents for targeted behaviors
Bai, F., Liu, R., Du, Y., Wen, Y., and Yang, Y · 2025
Closest in time.
Collapse or thrive? perils and promises of synthetic data in a self-generating world, 2025
Kazdan, J., Schaeffer, R., Dey, A., Gerstgrasser, M., Rafailov, R., Donoho, D. L., and Koyejo, S · 2025
Closest in time.