Fetching the paper…
Reading the bibliography…
Progress in machine learning has been driven in large part by massive increases in data.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2019
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 1910
Earlier work this paper cites.
On the resemblance and containment of documents
A. Broder · 1997
Earlier work this paper cites.
YFCC100M: the new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Earlier work this paper cites.
Coresets and sketches
J. M. Phillips · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. Patwary, M. Ali, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
M. Toneva, A. Sordoni, R. T. des Combes, A. Trischler, Y. Bengio, and G. J. Gordon · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. Lipton, and E. P. Xing · 2019
Earlier work this paper cites.
Do ImageNet classifiers generalize to ImageNet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz · 2019
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al · 2020
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
J. S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
V. Feldman and C. Zhang · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation
M. A. Gordon, K. Duh, and J. Kaplan · 2021
Cited alongside, same era.
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish · 2021
Cited alongside, same era.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Cited alongside, same era.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. v. d. Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. S. Morcos · 2022
Later among the works it cites.
Dataset Deduplication with Datamodels
Y. Liao · 2022
Later among the works it cites.
Deduplicating training data mitigates privacy risks in language models, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Openclip, July 2021
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt · 2021
Cited alongside, same era.
Deduplicating training data makes language models better, 2021
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis and insights from training gopher, 2021
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. v. d. Driessche, L. A. Hendricks, M. Rauh, P.-S. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J.-B. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. d. M. d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. d. L. Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. Hechtman, L. Weidinger, I. Gabriel, W. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
M. Paul, S. Ganguli, and G. K. Dziugaite · 2021
Cited alongside, same era.
Training data subset search with ensemble active learning
K. Chitta, J. M. Álvarez, E. Haussmann, and C. Farabet · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
N. Kandpal, E. Wallace, and C. Raffel · 2022
Later among the works it cites.
Noise-Robust De-Duplication at scale
E. Silcock, L. D’Amico-Wong, J. Yang, and M. Dell · 2022
Later among the works it cites.
DUEL: Adaptive duplicate elimination on working memory for Self-Supervised learning
W.-S. Choi, D.-S. Han, H. Lee, J. Park, and B.-T. Zhang · 2022
Later among the works it cites.
DeepCore: A comprehensive library for coreset selection in deep learning
C. Guo, B. Zhao, and Y. Bai · 2022
Later among the works it cites.
Trivial or impossible—dichotomous data difficulty masks model differences (on ImageNet and beyond)
K. Meding, L. M. S. Buschoff, R. Geirhos, and F. A. Wichmann · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Later among the works it cites.
Opt-iml: Scaling language model instruction meta learning through the lens of generalization, 2022
S. Iyer, X. V. Lin, R. Pasunuru, T. Mihaylov, D. Simig, P. Yu, K. Shuster, T. Wang, Q. Liu, P. S. Koura, X. Li, B. O’Horo, G. Pereyra, J. Wang, C. Dewan, A. Celikyilmaz, L. Zettlemoyer, and V. Stoyanov · 2022
Later among the works it cites.
Scaling laws for generative mixed-modal language models, 2023
A. Aghajanyan, L. Yu, A. Conneau, W.-N. Hsu, K. Hambardzumyan, S. Zhang, S. Roller, N. Goyal, O. Levy, and L. Zettlemoyer · 2023
Closest in time.
Filtering, distillation, and hard negatives for vision-language pre-training
F. Radenovic, A. Dubey, A. Kadian, T. Mihaylov, S. Vandenhende, Y. Patel, Y. Wen, V. Ramanathan, and D. Mahajan · 2023
Closest in time.