Fetching the paper…
Reading the bibliography…
Recent works have shown that machine learning models improve at a predictable rate with the total amount of training data, leading to scaling laws that describe the relationship between error and dataset size.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N. (2019) · 1909
Earlier work this paper cites.
Characterizations of an empirical influence function for detecting influential cases in regression
Cook, R. D. and Weisberg, S. (1980) · 1980
Earlier work this paper cites.
Asymptotic statistics
Van der Vaart, A. W. (2000) · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Boosted decision trees as an alternative to artificial neural networks for particle identification
Roe, B. P., Yang, H.-J., Zhu, J., Liu, Y., Stancu, I., and McGregor, G. (2005) · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
UCI machine learning repository
Dua, D. and Graff, C. (2017) · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. (2017) · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P. (2017) · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Deterministic inequalities for smooth M-estimators
Kuchibhotla, A. K. (2018) · 2018
Earlier work this paper cites.
Data Shapley: Equitable valuation of data for machine learning
Ghorbani, A. and Zou, J. (2019) · 2019
Cited alongside, same era.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C. (2020) · 2020
Cited alongside, same era.
A distributional framework for data valuation
Ghorbani, A., Kim, M., and Zou, J. (2020) · 2020
Cited alongside, same era.
Model performance scaling with multiple data sources
Hashimoto, T. (2021) · 2021
Cited alongside, same era.
Hutter, M. (2021) · 2021
Cited alongside, same era.
Beta Shapley: a unified and noise-reduced data valuation framework for machine learning
Evaluating self-supervised learning via risk decomposition
Dubois, Y., Hashimoto, T., and Liang, P. (2023) · 2023
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al. (2023) · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. (2023) · 2023
Later among the works it cites.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kwon, Y. and Zou, J. (2021) · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Cited alongside, same era.
Representation matters: Assessing the importance of subgroup allocations in training data
Rolf, E., Worledge, T. T., Recht, B., and Jordan, M. (2021) · 2021
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Cited alongside, same era.
Datamodels: Predicting predictions from training data
Ilyas, A., Park, S. M., Engstrom, L., Leclerc, G., and Madry, A. (2022) · 2022
Cited alongside, same era.
Data Banzhaf: A data valuation framework with maximal robustness to learning stochasticity
Wang, T. and Jia, R. (2022) · 2022
Cited alongside, same era.
Tutorial on amortized optimization
Amos, B. et al. (2023) · 2023
Cited alongside, same era.
Jiang, K. F., Liang, W., Zou, J., and Kwon, Y. (2023) · 2023
Later among the works it cites.
Lava: Data valuation without pre-specified learning algorithms
Just, H. A., Kang, F., Wang, J. T., Zeng, Y., Ko, M., Jin, M., and Jia, R. (2023) · 2023
Later among the works it cites.
DataInf: Efficiently estimating data influence in lora-tuned llms and diffusion models
Kwon, Y., Wu, E., Wu, K., and Zou, J. (2023) · 2023
Later among the works it cites.
Data-OOB: Out-of-bag estimate as a simple and efficient data value
Kwon, Y. and Zou, J. (2023) · 2023
Later among the works it cites.
Robust data valuation with weighted banzhaf values
Li, W. and Yu, Y. (2023) · 2023
Later among the works it cites.
OpenAI (2023) · 2023
Later among the works it cites.
TRAK: Attributing model behavior at scale
Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G., and Madry, A. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
Variance reduced Shapley value estimation for trustworthy data valuation
Wu, M., Jia, R., Lin, C., Huang, W., and Chang, X. (2023) · 2023
Later among the works it cites.