Fetching the paper…
Reading the bibliography…
The quality of foundation models depends heavily on their training data.
A value for n-person games
Lloyd S Shapley et al · 1953
Earlier work this paper cites.
Evolutionsstrategie : Optimierung technischer Systeme nach Prinzipien der biologischen Evolution
Ingo Rechenberg · 1973
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel · 1974
Earlier work this paper cites.
Detection of influential observation in linear regression
R Dennis Cook · 1977
Earlier work this paper cites.
Optimization and nonsmooth analysis
Frank H Clarke · 1990
Earlier work this paper cites.
Sparse approximate solutions to linear systems
B. K. Natarajan · 1995
Earlier work this paper cites.
Adaptive greedy approximations
G Davis, S Mallat, and M Avellaneda · 1997
Earlier work this paper cites.
Mapreduce: Simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat · 2004
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning, 2017
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Jax: composable transformations of python+ numpy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, et al · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Earlier work this paper cites.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado P van Hasselt, and David Silver · 2018
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
The DeepMind JAX Ecosystem, 2020
DeepMind, Igor Babuschkin, Kate Baumli, Alison Bell, Surya Bhupatiraju, Jake Bruce, Peter Buchlovsky, David Budden, Trevor Cai, Aidan Clark, Ivo Danihelka, Antoine Dedieu, Claudio Fantacci, Jonathan Godwin, Chris Jones, Ross Hemsley, Tom Hennigan, Matteo Hessel, Shaobo Hou, Steven Kapturowski, Thomas Keck, Iurii Kemaev, Michael King, Markus Kunesch, Lena Martens, Hamza Merzic, Vladimir Mikulik, Tamara Norman, George Papamakarios, John Quan, Roman Ring, Francisco Ruiz, Alvaro Sanchez, Laurent Sartran, Rosalia Schneider, Eren Sezener, Stephen Spencer, Srivatsan Srinivasan, Miloš Stanojević, Wojciech Stokowiec, Luyu Wang, Guangyao Zhou, and Fabio Viola · 2020
Lima: less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, Colin Raffel, Shiyu Chang, Tatsunori Hashimoto, and William Yang Wang · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology, 2024
Gemma Team · 2024
Later among the works it cites.
Bilevel optimization to learn training distributions for language modeling under domain shift
David Grangier, Pierre Ablin, and Awni Hannun · 2024
Later among the works it cites.
Not all tokens are what you need for pretraining
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, yelong shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, and Weizhu Chen · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Safe deep semi-supervised learning for unseen-class unlabeled data
Lan-Zhe Guo, Zhen-Yu Zhang, Yuan Jiang, Yu-Feng Li, and Zhi-Hua Zhou · 2020
Cited alongside, same era.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Optimizing data usage via differentiable rewards
Xinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos, Jaime Carbonell, and Graham Neubig · 2020
Cited alongside, same era.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2020
Cited alongside, same era.
Podracer architectures for scalable reinforcement learning
Matteo Hessel, Manuel Kroiss, Aidan Clark, Iurii Kemaev, John Quan, Thomas Keck, Fabio Viola, and Hado van Hasselt · 2021
Cited alongside, same era.
The llama 3 herd of models, 2024
Llama 3 Authors · 2024
Later among the works it cites.
Scalebio: Scalable bilevel optimization for llm data reweighting, 2024
Rui Pan, Jipeng Zhang, Xingyuan Pan, Renjie Pi, Xiaoyu Wang, and Tong Zhang · 2024
Later among the works it cites.
Data, data everywhere: A guide for pretraining dataset construction
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu, Aastha Jhunjhunwala, Zhilin Wang, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Later among the works it cites.
The fineweb datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydlícek, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin A. Raffel, Leandro von Werra, and Thomas Wolf · 2024
Later among the works it cites.
Dolma: an open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2024
Later among the works it cites.
Greats: Online selection of high-quality data for llm training in every iteration
Jiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal, and Ruoxi Jia · 2024
Later among the works it cites.
Perplexed by perplexity: Perplexity-based data pruning with small reference models
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L Leavitt, and Mansheej Paul · 2025
Closest in time.
Optimizing ml training with metagradient descent, 2025
Logan Engstrom, Andrew Ilyas, Benjamin Chen, Axel Feldmann, William Moses, and Aleksander Madry · 2025
Closest in time.
Data selection via optimal control for language models
Yuxian Gu, Li Dong, Hongning Wang, Yaru Hao, Qingxiu Dong, Furu Wei, and Minlie Huang · 2025
Closest in time.
Advancing data selection for foundation models: From heuristics to principled methods
Ruoxi Jia Jiachen Wang, Ludwig Schmidt · 2025
Closest in time.
Scalable meta-learning via mixed-mode differentiation
Iurii Kemaev, Dan A. Calian, Luisa M Zintgraf, Gregory Farquhar, and Hado Van Hasselt · 2025
Closest in time.
SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection
Han Shen, Pin-Yu Chen, Payel Das, and Tianyi Chen · 2025
Closest in time.
Predictive data selection: The data that predicts is the data that teaches
Kashun Shum, Yuzhen Huang, Hongjian Zou, Ding Qi, Yixuan Liao, Xiaoxin Chen, Qian Liu, and Junxian He · 2025
Closest in time.
Dynamic loss-based sample reweighting for improved large language model pretraining
Daouda Sow, Herbert Woisetschläger, Saikiran Bulusu, Shiqiang Wang, Hans Arno Jacobsen, and Yingbin Liang · 2025
Closest in time.
Data shapley in one training run
Jiachen T. Wang, Prateek Mittal, Dawn Song, and Ruoxi Jia · 2025
Closest in time.
Just select twice: Leveraging low quality data to improve data selection
Yifei Zhang, Yusen Jiao, Jiayi Chen, Jieyu Zhang, and Frederic Sala · 2025
Closest in time.