Fetching the paper…
Reading the bibliography…
The goal of predictive data attribution is to estimate how adding or removing a given set of training datapoints will affect model predictions.
“The Influence Curve and Its Role in Robust Estimation”
Frank. Hampel · 1947
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images”
Alex Krizhevsky · 2009
Earlier work this paper cites.
“Robust statistics: the approach based on influence functions”
Frank Hampel, Elvezio Ronchetti, Peter Rousseeuw and Werner Stahel · 2011
Earlier work this paper cites.
“Gradient-based hyperparameter optimization through reversible learning”
Dougal Maclaurin, David Duvenaud and Ryan Adams · 2015
Earlier work this paper cites.
“Optimizing Neural Networks with Kronecker-factored Approximate Curvature”
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
“Forward and reverse gradient-based hyperparameter optimization”
Luca Franceschi, Michele Donini, Paolo Frasconi and Massimiliano Pontil · 2017
Earlier work this paper cites.
“Understanding Black-box Predictions via Influence Functions”
Pang Koh and Percy Liang · 2017
Earlier work this paper cites.
“Fast Approximate Natural Gradient Descent in a Kronecker-factored Eigenbasis”
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas and Pascal Vincent · 2018
Earlier work this paper cites.
“Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)”
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler and Fernanda Viegas · 2018
Earlier work this paper cites.
“A scalable estimate of the extra-sample prediction error via approximate leave-one-out”
Kamiar Rad and Arian Maleki · 2018
Earlier work this paper cites.
“A swiss army infinitesimal jackknife”
Ryan Giordano, William Stephenson, Runjing Liu, Michael Jordan and Tamara Broderick · 2019
Earlier work this paper cites.
“On the accuracy of influence functions for measuring group effects”
Pang Koh, Kai-Siang Ang, Hubert Teo and Percy Liang · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Earlier work this paper cites.
“Optimizing millions of hyperparameters by implicit differentiation”
Jonathan Lorraine, Paul Vicol and David Duvenaud · 2020
Earlier work this paper cites.
“Predictive multiplicity in classification”
Charles Marx, Flavio Calmon and Berk Ustun · 2020
Earlier work this paper cites.
“Influence Functions in Deep Learning Are Fragile”
Samyadeep Basu, Phillip Pope and Soheil Feizi · 2021
Earlier work this paper cites.
“Model performance scaling with multiple data sources”
Tatsunori Hashimoto · 2021
Cited alongside, same era.
“Measuring Mathematical Problem Solving With the MATH Dataset”
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song and Jacob Steinhardt · 2021
Cited alongside, same era.
“LoRA: Low-Rank Adaptation of Large Language Models”, 2021
Edward. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision”
Alec Radford, Jong Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin and Jack Clark · 2021
Cited alongside, same era.
“If Influence Functions are the Answer, Then What is the Question?”
“The flan collection: Designing data and methods for effective instruction tuning”
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Chung, Yi Tay, Denny Zhou, Quoc Le, Barret Zoph and Jason Wei · 2023
Later among the works it cites.
“Scaling Data-Constrained Language Models”
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf and Colin Raffel · 2023
Later among the works it cites.
“A Bayesian Perspective On Training Data Attribution”
Elisa Nguyen, Minjoon Seo and Seong Oh · 2023
Later among the works it cites.
“TRAK: Attributing Model Behavior at Scale”
Sung Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc and Aleksander Madry · 2023
Later among the works it cites.
“Understanding Influence Functions and Datamodels via Harmonic Analysis”
Nikunj Saunshi, Arushi Gupta, Mark Braverman and Sanjeev Arora · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi and Roger Grosse · 2022
Cited alongside, same era.
“Model multiplicity: Opportunities, concerns, and solutions”
Emily Black, Manish Raghavan and Solon Barocas · 2022
Cited alongside, same era.
“English Wikipedia”, https://huggingface.co/datasets/wikipedia , 2022
Wikimedia Foundation · 2022
Cited alongside, same era.
“Training Data Influence Analysis and Estimation: A Survey”
Zayd Hammoudeh and Daniel Lowd · 2022
Cited alongside, same era.
“Datamodels: Predicting Predictions from Training Data”
Andrew Ilyas, Sung Park, Logan Engstrom, Guillaume Leclerc and Aleksander Madry · 2022
Cited alongside, same era.
“Tracing and Removing Data Errors in Natural Language Generation Datasets”
Faisal Ladhak, Esin Durmus and Tatsunori Hashimoto · 2022
Cited alongside, same era.
“ModelDiff: A Framework for Comparing Learning Algorithms”
Harshay Shah, Sung Park, Andrew Ilyas and Aleksander Madry · 2022
Cited alongside, same era.
“Scaling up influence functions”
Andrea Schioppa, Polina Zablotskaia, David Vilar and Artem Sokolov · 2022
Cited alongside, same era.
Juhan Bae, Wu Lin, Jonathan Lorraine and Roger Grosse · 2024
Later among the works it cites.
“Attribute-to-Delete: Machine Unlearning via Datamodel Matching”
Kristian Georgiev, Roy Rinberg, Sung Park, Shivam Garg, Andrew Ilyas, Aleksander Madry and Seth Neel · 2024
Later among the works it cites.
“Data Attribution at Scale”, Tutorial at ICML 2024, 2024
Andrew Ilyas, Kristian Georgiev, Logan Engstrom and Sung Park · 2024
Later among the works it cites.
“94 percent on CIFAR-10 in 3.29 Seconds on a Single GPU”
Keller Jordan · 2024
Later among the works it cites.
“On the Variance of Neural Network Training with respect to Test Sets and Distributions”
Keller Jordan · 2024
Later among the works it cites.
“Openassistant conversations-democratizing large language model alignment”
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley and Richárd Nagyfi · 2024
Later among the works it cites.
“Generalized Group Data Attribution”
Dan Ley, Suraj Srinivas, Shichang Zhang, Gili Rusak and Himabindu Lakkaraju · 2024
Later among the works it cites.
“RandALO: Out-of-sample risk estimation in no time flat”
Parth Nobel, Daniel LeJeune and Emmanuel Candès · 2024
Later among the works it cites.
“Gemma: Open models based on gemini research and technology”
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Kale and Juliette Love · 2024
Later among the works it cites.
“Less: Selecting influential data for targeted instruction tuning”
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora and Danqi Chen · 2024
Later among the works it cites.
“Optimizing ML Training with Metagradient Descent”
Logan Engstrom, Andrew Ilyas, Benjamin Chen, Axel Feldmann, William Moses and Aleksander Madry · 2025
Closest in time.