Fetching the paper…
Reading the bibliography…
Training data attribution (TDA) provides insights into which training data is responsible for a learned model behavior.
The proof and measurement of association between two things
C Spearman · 1904
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems
Magnus R Hestenes, Eduard Stiefel, et al · 1952
Earlier work this paper cites.
A rapidly convergent descent method for minimization
R. Fletcher and M. J. D. Powell · 1963
Earlier work this paper cites.
Function minimization by conjugate gradients
Reeves Fletcher and Colin M Reeves · 1964
Earlier work this paper cites.
A class of methods for solving nonlinear simultaneous equations
Charles G Broyden · 1965
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel · 1974
Earlier work this paper cites.
Influential observations in linear regression
R Dennis Cook · 1979
Earlier work this paper cites.
Updating quasi-newton matrices with limited storage
Jorge Nocedal · 1980
Earlier work this paper cites.
Statistics and causal inference
Paul W Holland · 1986
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Iterative methods for sparse linear systems
Yousef Saad · 2003
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi · 2006
Earlier work this paper cites.
Spectral algorithms for supervised learning
L Lo Gerfo, Lorenzo Rosasco, Francesca Odone, E De Vito, and Alessandro Verri · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens et al · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
A stochastic quasi-newton method for large-scale optimization
Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Second-order stochastic optimization for machine learning in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2017
Cited alongside, same era.
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms
Revisiting methods for finding influential examples
Anders Søgaard et al · 2021
Later among the works it cites.
Tracing knowledge in language models back to the training data
Ekin Akyürek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Later among the works it cites.
On implicit bias in overparameterized bilevel optimization
Paul Vicol, Jonathan P Lorraine, Fabian Pedregosa, David Duvenaud, and Roger B Grosse · 2022
Later among the works it cites.
Scaling up influence functions
Andrea Schioppa, Polina Zablotskaia, David Vilar, and Artem Sokolov · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K Ravikumar · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
A scalable laplace approximation for neural networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Cited alongside, same era.
Noisy natural gradient as variational inference
Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
If influence functions are the answer, then what is the question?
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger B Grosse · 2022
Later among the works it cites.
Combining feature and instance attribution to detect artifacts
Pouya Pezeshkpour and Sarthak Jain · 2022
Later among the works it cites.
Operationalizing machine learning: An interview study
Shreya Shankar, Rolando Garcia, Joseph M Hellerstein, and Aditya G Parameswaran · 2022
Later among the works it cites.
More than a toy: Random matrix models predict how real-world neural representations generalize
Alexander Wei, Wei Hu, and Jacob Steinhardt · 2022
Later among the works it cites.
Understanding influence functions and datamodels via harmonic analysis
Nikunj Saunshi, Arushi Gupta, Mark Braverman, and Sanjeev Arora · 2022
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Later among the works it cites.
Error discovery by clustering influence embeddings
Fulton Wang, Julius Adebayo, Sarah Tan, Diego Garcia-Olano, and Narine Kokhlikyan · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Revisiting the fragility of influence functions
Jacob R Epifano, Ravi P Ramachandran, Aaron J Masino, and Ghulam Rasool · 2023
Later among the works it cites.
Datainf: Efficiently estimating data influence in lora-tuned llms and diffusion models
Yongchan Kwon, Eric Wu, Kevin Wu, and James Zou · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Later among the works it cites.
How many and which training points would need to be removed to flip this prediction?
Jinghan Yang, Sarthak Jain, and Byron C Wallace · 2023
Later among the works it cites.
What is your data worth to gpt? llm-scale data valuation with influence functions
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, et al · 2024
Later among the works it cites.
Improving subgroup robustness via data selection
Saachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas, Marzyeh Ghassemi, and Aleksander Madry · 2024
Later among the works it cites.
Influence functions for scalable data attribution in diffusion models
Bruno Mlodozeniec, Runa Eschenhagen, Juhan Bae, Alexander Immer, David Krueger, and Richard Turner · 2024
Later among the works it cites.
Training data attribution via approximate unrolled differentation
Juhan Bae, Wu Lin, Jonathan Lorraine, and Roger Grosse · 2024
Later among the works it cites.
Revisiting inverse hessian vector products for calculating influence functions
Yegor Klochkov and Yang Liu · 2024
Later among the works it cites.
Theoretical and practical perspectives on what influence functions do
Andrea Schioppa, Katja Filippova, Ivan Titov, and Polina Zablotskaia · 2024
Later among the works it cites.
The mirrored influence hypothesis: Efficient data influence estimation by harnessing forward passes
Myeongseob Ko, Feiyang Kang, Weiyan Shi, Ming Jin, Zhou Yu, and Ruoxi Jia · 2024
Later among the works it cites.
Training data influence analysis and estimation: A survey
Zayd Hammoudeh and Daniel Lowd · 2024
Later among the works it cites.
dattri: A library for efficient data attribution
Junwei Deng, Ting-Wei Li, Shiyuan Zhang, Shixuan Liu, Yijun Pan, Hao Huang, Xinhe Wang, Pingbang Hu, Xingjian Zhang, and Jiaqi Ma · 2024
Later among the works it cites.
Do influence functions work on large language models?
Zhe Li, Wei Zhao, Yige Li, and Jun Sun · 2024
Later among the works it cites.
Most influential subset selection: Challenges, promises, and beyond
Yuzheng Hu, Pingbang Hu, Han Zhao, and Jiaqi W Ma · 2024
Later among the works it cites.
Measuring stochastic data complexity with boltzmann influence functions
Nathan Ng, Roger Grosse, and Marzyeh Ghassemi · 2024
Later among the works it cites.
Who owns the output? bridging law and technology in llms attribution
Emanuele Mezzi, Asimina Mertzani, Michael P Manis, Siyanna Lilova, Nicholas Vadivoulis, Stamatis Gatirdakis, Styliani Roussou, and Rodayna Hmede · 2025
Closest in time.
Magic: Near-optimal data attribution for deep learning
Andrew Ilyas and Logan Engstrom · 2025
Closest in time.
Beyond Gradients: Using Curvature Information for Deep Learning
Juhan Bae · 2025
Closest in time.