Fetching the paper…
Reading the bibliography…
Modern machine learning pipelines are increasingly combining and mixing data from diverse and disparate sources, e.g., pre-training large language models.
An extended kuhn–tucker approach for linear bilevel programming
Shi, C., Lu, J., and Zhang, G · 2005
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W · 2010
Earlier work this paper cites.
Extended-connectivity fingerprints
Rogers, D. and Hahn, M · 2010
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Earlier work this paper cites.
Pubchem substance and compound databases
Kim, S., Thiessen, P. A., Bolton, E. E., Chen, J., Fu, G., Gindulyte, A., Han, L., He, J., He, S., Shoemaker, B. A., et al · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Pedregosa, F · 2016
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Introductory lectures on stochastic optimization
Duchi, J. C · 2018
Earlier work this paper cites.
An integrated transfer learning and multitask learning approach for pharmacokinetic parameter prediction
Ye, Z., Yang, Y., Li, X., Cao, D., and Ouyang, D · 2018
Earlier work this paper cites.
Analyzing learned molecular representations for property prediction
Yang, K., Swanson, K., Jin, W., Coley, C., Eiden, P., Gao, H., Guzman-Perez, A., Hopper, T., Kelley, B., Mathea, M., et al · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Transcreen: transfer learning on graph-based anti-cancer virtual screening model
Salem, M., Khormali, A., Arshadi, A. K., Webb, J., and Yuan, J.-S · 2020
Earlier work this paper cites.
Regmix: Data mixing augmentation for regression
Hwang, S.-H. and Whang, S. E · 2021
Cited alongside, same era.
Bilevel optimization: Convergence analysis and enhanced design
Ji, K., Yang, J., and Liang, Y · 2021
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Improving molecular property prediction through a task similarity enhanced transfer learning strategy
Li, H., Zhao, X., Li, S., Wan, F., Zhao, D., and Zeng, J · 2022
Cited alongside, same era.
A feature transferring workflow between data-poor compounds in various tasks
Sun, X., Zhu, J., Chen, B., You, H., and Xu, H · 2022
Cited alongside, same era.
Data selection for language models via importance resampling
Xie, S. M., Santurkar, S., Ma, T., and Liang, P · 2023
Later among the works it cites.
Does your data spark joy? performance gains from domain upsampling at the end of training
Blakeney, C., Paul, M., Larsen, B. W., Owen, S., and Frankle, J · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Gio: Gradient information optimization for training dataset selection
Everaert, D. and Potts, C · 2024
Later among the works it cites.
Scaling laws for downstream task performance of large language models
Isik, B., Ponomareva, N., Hazimeh, H., Paparas, D., Vassilvitskii, S., and Koyejo, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama, 2023
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama · 2023
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Efficient online data mixing for language model pre-training
Albalak, A., Pan, L., Raffel, C., and Wang, W. Y · 2023
Cited alongside, same era.
Towards foundational models for molecular learning on large-scale multi-task datasets
Beaini, D., Huang, S., Cunha, J. A., Li, Z., Moisescu-Pareja, G., Dymov, O., Maddrell-Mander, S., McLean, C., Wenkel, F., Müller, L., et al · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Cited alongside, same era.
Doge: Domain reweighting with generalization estimation
Fan, S., Pagliardini, M., and Jaggi, M · 2023
Cited alongside, same era.
Openwebmath: An open dataset of high-quality mathematical web text, 2023
Paster, K., Santos, M. D., Azerbayev, Z., and Ba, J · 2023
Cited alongside, same era.
Later among the works it cites.
Scaling laws for learning with real and surrogate data
Jain, A., Montanari, A., and Sasoglu, E · 2024
Later among the works it cites.
Adaptive data optimization: Dynamic sample selection with scaling laws
Jiang, Y., Zhou, A., Feng, Z., Malladi, S., and Kolter, J. Z · 2024
Later among the works it cites.
Regmix: Data mixture as regression for language model pre-training
Liu, Q., Zheng, X., Muennighoff, N., Zeng, G., Dou, L., Pang, T., Jiang, J., and Lin, M · 2024
Later among the works it cites.
No filter: Cultural and socioeconomic diversityin contrastive vision-language models
Pouget, A., Beyer, L., Bugliarello, E., Wang, X., Steiner, A. P., Zhai, X., and Alabdulmohsin, I · 2024
Later among the works it cites.
Improving pretraining data using perplexity correlations
Thrush, T., Potts, C., and Hashimoto, T · 2024
Later among the works it cites.
Finding optimally robust data mixtures via concave maximization
Thudi, A. and Maddison, C. J · 2024
Later among the works it cites.
C-pack: Packed resources for general chinese embeddings
Xiao, S., Liu, Z., Zhang, P., Muennighoff, N., Lian, D., and Nie, J.-Y · 2024
Later among the works it cites.
Doremi: Optimizing data mixtures speeds up language model pretraining
Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P. S., Le, Q. V., Ma, T., and Yu, A. W · 2024
Later among the works it cites.
Optimizing pretraining data mixtures with llm-estimated utility
Held, W., Paranjape, B., Koura, P. S., Lewis, M., Zhang, F., and Mihaylov, T · 2025
Closest in time.