Fetching the paper…
Reading the bibliography…
As AI model size grows, neural scaling laws have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A. and de Vito, E · 2007
Earlier work this paper cites.
On the convergence of the empirical distribution
Berend, D. and Kontorovich, A · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Nonparametric regression using deep neural networks with relu activation function
Schmidt-Hieber, J · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality
Suzuki, T · 2019
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Bordelon, B., Canatar, A., and Pehlevan, C · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Self-distillation amplifies regularization in Hilbert space
Mobahi, H., Farajtabar, M., and Bartlett, P · 2020
Earlier work this paper cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Papyan, V., Han, X., and Donoho, D. L · 2020
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 2020
Earlier work this paper cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Spigler, S., Geiger, M., and Wyart, M · 2020
Earlier work this paper cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Cui, H., Loureiro, B., Krzakala, F., and Zdeborova, L · 2021
Earlier work this paper cites.
Data and parameter scaling laws for neural machine translation
Gordon, M. A., Duh, K., and Kaplan, J · 2021
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2021
Earlier work this paper cites.
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Earlier work this paper cites.
Hutter, M · 2021
Earlier work this paper cites.
Evaluating distributional distortion in neural language modeling
LeBrun, B., Sordoni, A., and O’Donnell, T. J · 2021
Earlier work this paper cites.
Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime
Nitanda, A. and Suzuki, T · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Generalization error rates in kernel regression: the crossover from the noiseless to noisy regime
Cui, H., Loureiro, B., Krzakala, F., and Zdeborová, L · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
Large language models can self-improve, 2022
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Cited alongside, same era.
Error scaling laws for kernel classification under source and capacity conditions
Cui, H., Loureiro, B., Krzakala, F., and Zdeborová, L · 2023
Later among the works it cites.
Auggpt: Leveraging chatgpt for text data augmentation, 2023
Dai, H., Liu, Z., Liao, W., Huang, X., Cao, Y., Wu, Z., Zhao, L., Xu, S., Liu, W., Liu, N., Li, S., Zhu, D., Cai, H., Sun, L., Li, Q., Shen, D., Liu, T., and Li, X · 2023
Later among the works it cites.
A simplistic model of neural scaling laws: Multiperiodic Santa Fe processes, 2023
Debowski, L · 2023
Later among the works it cites.
Scaling laws of synthetic images for model training… for now
Fan, L., Chen, K., Krishnan, D., Katabi, D., Isola, P., and Tian, Y · 2023
Later among the works it cites.
The curious decline of linguistic diversity: Training language models on synthetic text, 2023
Guo, Y., Shang, G., Vazirgiannis, M., and Clavel, C · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A solvable model of neural scaling laws, 2022
Maloney, A., Roberts, D. A., and Sully, J · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Scaling laws for generative mixed-modal language models, 2023
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L · 2023
Cited alongside, same era.
Self-consuming generative models go mad
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., and Baraniuk, R. G · 2023
Cited alongside, same era.
A theory for emergence of complex skills in language models
Arora, S. and Goyal, A · 2023
Cited alongside, same era.
Will large-scale generative models corrupt future datasets?
Hataya, R., Bao, H., and Arai, H · 2023
Later among the works it cites.
Is synthetic data from generative models ready for image recognition?
He, R., Sun, S., Yu, X., Xue, C., Zhang, W., Torr, P., Bai, S., and QI, X · 2023
Later among the works it cites.
Explore the power of synthetic data on few-shot object detection
Lin, S., Wang, K., Zeng, X., and Zhao, R · 2023
Later among the works it cites.
Inverse scaling: When bigger isn’t better
McKenzie, I. R., Lyzhov, A., Pieler, M. M., Parrish, A., Mueller, A., Prabhu, A., McLean, E., Shen, X., Cavanagh, J., Gritsevskiy, A. G., Kauffman, D., Kirtland, A. T., Zhou, Z., Zhang, Y., Huang, S., Wurgaft, D., Weiss, M., Ross, A., Recchia, G., Liu, A., Liu, J., Tseng, T., Korbak, T., Kim, N., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
The quantization model of neural scaling
Michaud, E. J., Liu, Z., Girit, U., and Tegmark, M · 2023
Later among the works it cites.
Midjourney ai, 2023
Midjourney · 2023
Later among the works it cites.
Diversity is definitely needed: Improving model-agnostic zero-shot classification via stable diffusion, 2023
Shipard, J., Wiliem, A., Thanh, K. N., Xiang, W., and Fookes, C · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Veselovsky, V., Ribeiro, M. H., and West, R · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Xu, C., Guo, D., Duan, N., and McAuley, J · 2023
Later among the works it cites.
The psycho-biology of language: an introduction to dynamic philology
Zipf, G · 2023
Later among the works it cites.
openai now generates about 100 billion words per day. all people on earth generate about 100 trillion words per day
Altman, S · 2024
Closest in time.
Model collapse demystified: The case of regression
Dohmatob, E., Feng, Y., and Kempe, J · 2024
Closest in time.