Fetching the paper…
Reading the bibliography…
Many training data attribution (TDA) methods aim to estimate how a model's behavior would change if one or more data points were removed from the training set.
The principle of minimized iterations in the solution of the matrix eigenvalue problem
Walter Edwin Arnoldi · 1951
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A value for n n -person games
Lloyd Shapley · 1953
Earlier work this paper cites.
Weighted voting doesn’t work: A mathematical analysis
John F Banzhaf III · 1964
Earlier work this paper cites.
The infinitesimal jackknife
Louis A Jaeckel · 1972
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel · 1974
Earlier work this paper cites.
Residuals and influence in regression
Sanford Weisberg and R Dennis Cook · 1982
Earlier work this paper cites.
Extensions of lipschitz maps into banach spaces
William B Johnson, Joram Lindenstrauss, and Gideon Schechtman · 1986
Earlier work this paper cites.
The proof and measurement of association between two things
Charles Spearman · 1987
Earlier work this paper cites.
Using the adap learning algorithm to forecast the onset of diabetes mellitus
Jack W Smith, James E Everhart, WC Dickson, William C Knowler, and Robert Scott Johannes · 1988
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Okapi at TREC-3
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, and Mike Gatford · 1995
Earlier work this paper cites.
Case-based explanation of non-case-based learning methods
Rich Caruana, Hooshang Kangarloo, John David Dionisio, Usha Sinha, and David Johnson · 1999
Earlier work this paper cites.
The Implicit Function Theorem: History, theory, and applications
Steven George Krantz and Harold R Parks · 2002
Earlier work this paper cites.
Concrete Compressive Strength
I-Cheng Yeh · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Parkinsons Telemonitoring
Athanasios Tsanas and Max Little · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Domain generalization for object recognition with multi-task autoencoders
Muhammad Ghifary, W Bastiaan Kleijn, Mengjie Zhang, and David Balduzzi · 2015
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks, 2015
Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Wide residual networks, 2017
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K Ravikumar · 2018
Influence selection for active learning
Zhuoming Liu, Hao Ding, Huaping Zhong, Weijia Li, Jifeng Dai, and Conghui He · 2021
Later among the works it cites.
Towards tracing knowledge in language models back to the training data
Ekin Akyürek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Later among the works it cites.
Beta Shapley: A unified and noise-reduced data valuation framework for machine learning
Yongchan Kwon and James Zou · 2022
Later among the works it cites.
Rank list sensitivity of recommender systems to interaction perturbations
Sejoon Oh, Berk Ustun, Julian McAuley, and Srijan Kumar · 2022
Later among the works it cites.
Scaling up influence functions
Andrea Schioppa, Polina Zablotskaia, David Vilar, and Artem Sokolov · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data Shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Cited alongside, same era.
Data cleansing for models trained with sgd
Satoshi Hara, Atsushi Nitanda, and Takanori Maehara · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Towards efficient data valuation based on the shapley value
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos · 2019
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Interpreting black box predictions using fisher kernels
Rajiv Khanna, Been Kim, Joydeep Ghosh, and Sanmi Koyejo · 2019
Cited alongside, same era.
Later among the works it cites.
First is better than last for language data influence
Chih-Kuan Yeh, Ankur Taly, Mukund Sundararajan, Frederick Liu, and Pradeep Ravikumar · 2022
Later among the works it cites.
Adapting and evaluating influence-estimation methods for gradient-boosted decision trees
Jonathan Brophy, Zayd Hammoudeh, and Daniel Lowd · 2023
Later among the works it cites.
Revisiting the fragility of influence functions
Jacob R Epifano, Ravi P Ramachandran, Aaron J Masino, and Ghulam Rasool · 2023
Later among the works it cites.
The journey, not the destination: How data guides diffusion models, 2023
Kristian Georgiev, Joshua Vendrow, Hadi Salman, Sung Min Park, and Aleksander Madry · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs, 2023
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Later among the works it cites.
Opendataval: A unified benchmark for data valuation, 2023
Kevin Fu Jiang, Weixin Liang, James Zou, and Yongchan Kwon · 2023
Later among the works it cites.
The UCI machine learning repository, 2023
Markelle Kelly, Rachel Longjohn, and Kolby Nottingham · 2023
Later among the works it cites.
Attributing learned concepts in neural networks to training data, 2023
Nicholas Konz, Charles Godfrey, Madelyn Shapiro, Jonathan Tu, Henry Kvinge, and Davis Brown · 2023
Later among the works it cites.
DataInf: Efficiently estimating data influence in lora-tuned llms and diffusion models
Yongchan Kwon, Eric Wu, Kevin Wu, and James Zou · 2023
Later among the works it cites.
Contrastive error attribution for finetuned language models
Faisal Ladhak, Esin Durmus, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Trustworthy machine learning, 2023
Bálint Mucsányi, Michael Kirchhof, Elisa Nguyen, Alexander Rubinstein, and Seong Joon Oh · 2023
Later among the works it cites.
TRAK: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
A simple and efficient baseline for data attribution on images
Vasu Singla, Pedro Sandoval-Segura, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Later among the works it cites.
Data Banzhaf: A robust data valuation framework for machine learning
Jiachen T Wang and Ruoxi Jia · 2023
Later among the works it cites.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini · 2023
Later among the works it cites.
Intriguing properties of data attribution on diffusion models
Xiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang, and Min Lin · 2023
Later among the works it cites.
DsDm: Model-aware dataset selection with datamodels, 2024
Logan Engstrom, Axel Feldmann, and Aleksander Madry · 2024
Closest in time.
Kronecker-factored approximate curvature for modern neural network architectures
Runa Eschenhagen, Alexander Immer, Richard Turner, Frank Schneider, and Philipp Hennig · 2024
Closest in time.
Training data influence analysis and estimation: A survey
Zayd Hammoudeh and Daniel Lowd · 2024
Closest in time.
GEX: A flexible method for approximating influence via geometric ensemble
SungYub Kim, Kyungsu Kim, and Eunho Yang · 2024
Closest in time.
A bayesian approach to analysing training data attribution in deep learning
Elisa Nguyen, Minjoon Seo, and Seong Joon Oh · 2024
Closest in time.
The memory-perturbation equation: Understanding model’s sensitivity to data
Peter Nickl, Lu Xu, Dharmesh Tailor, Thomas Möllenhoff, and Mohammad Emtiyaz E Khan · 2024
Closest in time.
Theoretical and practical perspectives on what influence functions do
Andrea Schioppa, Katja Filippova, Ivan Titov, and Polina Zablotskaia · 2024
Closest in time.
A privacy-friendly approach to data valuation
Jiachen Tianhao Wang, Yuqing Zhu, Yu-Xiang Wang, Ruoxi Jia, and Prateek Mittal · 2024
Closest in time.
LESS: Selecting influential data for targeted instruction tuning, 2024
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen · 2024
Closest in time.