Fetching the paper…
Reading the bibliography…
In this paper, we introduce Group-MATES, an efficient group-level data selection approach to optimize the speed-quality frontier of language model pretraining.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Least squares quantization in pcm
Stuart Lloyd · 1982
Earlier work this paper cites.
Residuals and influence in regression
Sanford Weisberg and R Dennis Cook · 1982
Earlier work this paper cites.
Multi-armed bandit algorithms and empirical evaluation
Joannes Vermorel and Mehryar Mohri · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
On the accuracy of influence functions for measuring group effects
Pang Wei W Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang · 2019
Earlier work this paper cites.
On second-order group influence functions for black-box predictions
Samyadeep Basu, Xuchen You, and Soheil Feizi · 2020
Earlier work this paper cites.
An automatic finite-sample robustness metric: When can dropping a little data make a big difference?
Tamara Broderick, Ryan Giordano, and Rachael Meager · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, and Laurent Sifre · 2022
Earlier work this paper cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Earlier work this paper cites.
Fair infinitesimal jackknife: Mitigating the influence of biased training data points without refitting
Prasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre Dognin, and Kush R Varshney · 2022
Earlier work this paper cites.
Understanding instance-level impact of fairness constraints
Jialu Wang, Xin Eric Wang, and Yang Liu · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2022
Cited alongside, same era.
SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Amro Kamal Mohamed Abbas, Kushal Tirumala, Daniel Simig, Surya Ganguli, and Ari S. Morcos · 2023
Cited alongside, same era.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Cited alongside, same era.
The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data only
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Cited alongside, same era.
DsDm: Model-aware dataset selection with datamodels
Logan Engstrom, Axel Feldmann, and Aleksander Madry · 2024
Later among the works it cites.
Approximations to worst-case data dropping: unmasking failure modes
Jenny Y Huang, David R Burt, Tin D Nguyen, Yunyi Shen, and Tamara Broderick · 2024
Later among the works it cites.
Generalized group data attribution
Dan Ley, Suraj Srinivas, Shichang Zhang, Gili Rusak, and Himabindu Lakkaraju · 2024
Later among the works it cites.
DataComp-LM: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, et al · 2024
Later among the works it cites.
Rephrasing the web: A recipe for compute and data-efficient language modeling
Pratyush Maini, Skyler Seto, Richard Bai, David Grangier, Yizhe Zhang, and Navdeep Jaitly · 2024
Later among the works it cites.
The FineWeb datasets: Decanting the web for the finest text data at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding influence functions and datamodels via harmonic analysis
Nikunj Saunshi, Arushi Gupta, Mark Braverman, and Sanjeev Arora · 2023
Cited alongside, same era.
D4: Improving LLM pretraining via document de-duplication and diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos · 2023
Cited alongside, same era.
Farewell to aimless large-scale pretraining: Influential subset selection for language model
Xiao Wang, Weikang Zhou, Qi Zhang, Jie Zhou, Songyang Gao, Junzhe Wang, Menghan Zhang, Xiang Gao, Yun Wen Chen, and Tao Gui · 2023
Cited alongside, same era.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang · 2023
Cited alongside, same era.
Dataset pruning: Reducing training data by examining generalization influence
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li · 2023
Cited alongside, same era.
Phi-4 technical report
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al · 2024
Cited alongside, same era.
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al · 2024
Cited alongside, same era.
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Later among the works it cites.
RedPajama: an open dataset for training large language models
Maurice Weber, Daniel Y Fu, Quentin Gregory Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Re, Irina Rish, and Ce Zhang · 2024
Later among the works it cites.
QuRating: Selecting high-quality data for training language models
Alexander Wettig, Aatmik Gupta, Saumya Malik, and Danqi Chen · 2024
Later among the works it cites.
C-Pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie · 2024
Later among the works it cites.
Synthetic continued pretraining
Zitong Yang, Neil Band, Shuangping Li, Emmanuel Candes, and Tatsunori Hashimoto · 2024
Later among the works it cites.
MATES: Model-aware data selection for efficient pretraining with data influence models
Zichun Yu, Spandan Das, and Chenyan Xiong · 2024
Later among the works it cites.
Programming every example: Lifting pre-training data quality like experts at scale
Fan Zhou, Zengzhi Wang, Qian Liu, Junlong Li, and Pengfei Liu · 2024
Later among the works it cites.
Data selection via optimal control for language models
Yuxian Gu, Li Dong, Hongning Wang, Yaru Hao, Qingxiu Dong, Furu Wei, and Minlie Huang · 2025
Closest in time.
Regmix: Data mixture as regression for language model pre-training
Qian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng, Longxu Dou, Tianyu Pang, Jing Jiang, and Min Lin · 2025
Closest in time.
Organize the web: Constructing domains enhances pre-training data curation
Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini · 2025
Closest in time.
Harnessing diversity for important data selection in pretraining large language models
Chi Zhang, Huaping Zhong, Kuan Zhang, Chengliang Chai, Rui Wang, Xinlin Zhuang, Tianyi Bai, Jiantao Qiu, Lei Cao, Ju Fan, et al · 2025
Closest in time.