Fetching the paper…
Reading the bibliography…
Data Shapley provides a principled framework for attributing data's contribution within machine learning contexts.
A value for n-person games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Characterizations of an empirical influence function for detecting influential cases in regression
R Dennis Cook and Sanford Weisberg · 1980
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Gradient similarity: An explainable approach to detect adversarial attacks against deep learning
Jasjeet Dhaliwal and Saurabh Shintre · 2018
Earlier work this paper cites.
Deep batch active learning by diverse, uncertain gradient lower bounds
Jordan T Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal · 2019
Earlier work this paper cites.
Emergent properties of the local geometry of neural loss landscapes
Stanislav Fort and Surya Ganguli · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
Estimation of the shapley value by ergodic sampling
Ferenc Illés and Péter Kerényi · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
A distributional framework for data valuation
Amirata Ghorbani, Michael Kim, and James Zou · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Efficient per-example gradient computations in convolutional neural networks
Gaspar Rochette, Andre Manoel, and Eric W Tramel · 2020
Earlier work this paper cites.
A principled approach to data valuation for federated learning
Tianhao Wang, Johannes Rausch, Ce Zhang, Ruoxi Jia, and Dawn Song · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Earlier work this paper cites.
Influence functions in deep learning are fragile
S Basu, P Pope, and S Feizi · 2021
Cited alongside, same era.
Approximating the shapley value using stratified empirical bernstein sampling
Mark Alexander Burgess and Archie C Chapman · 2021
Cited alongside, same era.
Efficient computation and analysis of distributional shapley values
Yongchan Kwon, Manuel A Rivas, and James Zou · 2021
Cited alongside, same era.
Scaling up differentially private deep learning with fast per-example gradient clipping
Jaewoo Lee and Daniel Kifer · 2021
Cited alongside, same era.
Large language models can be strong differentially private learners
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto · 2021
Cited alongside, same era.
A multilinear sampling algorithm to estimate shapley values
Ramin Okhrati and Aldo Lipani · 2021
Studying large language model generalization with influence functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al · 2023
Later among the works it cites.
The times sues openai and microsoft
Michael M Grynbaum and Ryan Mac · 2023
Later among the works it cites.
This new tool could give artists an edge over ai
Melissa Heikkilä · 2023
Later among the works it cites.
Foundation models and fair use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang · 2023
Later among the works it cites.
Faster approximation of probabilistic and distributional values via least squares
Weida Li and Yaoliang Yu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Revisiting methods for finding influential examples
Anders Søgaard et al · 2021
Cited alongside, same era.
If influence functions are the answer, then what is the question?
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger B Grosse · 2022
Cited alongside, same era.
Scalable and efficient training of large convolutional neural networks with differential privacy
Zhiqi Bu, Jialin Mao, and Shiyun Xu · 2022
Cited alongside, same era.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Cited alongside, same era.
Beta shapley: a unified and noise-reduced data valuation framework for machine learning
Yongchan Kwon and James Zou · 2022
Cited alongside, same era.
Measuring the effect of training data on deep learning predictions via randomized experiments
Jinkun Lin, Anqi Zhang, Mathias Lécuyer, Jinyang Li, Aurojit Panda, and Siddhartha Sen · 2022
Cited alongside, same era.
A bayesian approach to analysing training data attribution in deep learning
Elisa Nguyen, Minjoon Seo, and Seong Joon Oh · 2023
Later among the works it cites.
Trak: attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Threshold knn-shapley: A linear-time and privacy-friendly approach to data valuation
Jiachen T Wang, Yuqing Zhu, Yu-Xiang Wang, Ruoxi Jia, and Prateek Mittal · 2023
Later among the works it cites.
What is your data worth to gpt? llm-scale data valuation with influence functions
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, et al · 2024
Closest in time.
Stochastic amortization: A unified approach to accelerate feature and data attribution
Ian Covert, Chanwoo Kim, Su-In Lee, James Zou, and Tatsunori Hashimoto · 2024
Closest in time.
A unified fast gradient clipping framework for dp-sgd
Weiwei Kong and Andres Munoz Medina · 2024
Closest in time.
Robust data valuation with weighted banzhaf values
Weida Li and Yaoliang Yu · 2024
Closest in time.
Generative ai’s end run around copyright
Caitlin Mulligan and James Li · 2024
Closest in time.
Gradient sketches for training data attribution and studying the loss landscape
Andrea Schioppa · 2024
Closest in time.
Less: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen · 2024
Closest in time.
On the inflation of knn-shapley value
Ziao Yang, Han Yue, Jian Chen, and Hongfu Liu · 2024
Closest in time.