Fetching the paper…
Reading the bibliography…
How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often fails to capture the more complex and pronounced collective influence of a set of samples.
A higher-order swiss army infinitesimal jackknife
R. Giordano, M. I. Jordan, and T. Broderick · 1907
Earlier work this paper cites.
The distribution of an arbitrary studentized residual and the effects of updating in multiple regression
R. Beckman and H. Trussell · 1974
Earlier work this paper cites.
The influence curve and its role in robust estimation
F. R. Hampel · 1974
Earlier work this paper cites.
D-optimality for regression designs: a review
R. S. John and N. R. Draper · 1975
Earlier work this paper cites.
Detection of influential observation in linear regression
R. D. Cook · 1977
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
G. L. Nemhauser, L. A. Wolsey, and M. L. Fisher · 1978
Earlier work this paper cites.
Regression Diagnostics: Identifying Influential Data and Sources of Collinearity
D. Belsley, E. Kuh, and R. Welsch · 1980
Earlier work this paper cites.
K-clustering as a detection tool for influential subsets in regression
J. B. Gray and R. F. Ling · 1984
Earlier work this paper cites.
K-clustering and the detection of influential subsets
A. S. Hadi · 1985
Earlier work this paper cites.
Influential observations, high leverage points, and outliers in linear regression
S. Chatterjee and A. S. Hadi · 1986
Earlier work this paper cites.
Assessment of local influence
R. D. Cook · 1986
Earlier work this paper cites.
Robust regression and outlier detection, 1987
P. Rousseeuw and A. Leroy · 1987
Earlier work this paper cites.
Waveform Database Generator (Version 1)
L. Breiman and C. Stone · 1988
Earlier work this paper cites.
Procedures for the identification of multiple outliers in linear models
A. S. Hadi and J. S. Simonoff · 1993
Earlier work this paper cites.
The detection of influential subsets in linear regression by using an influence matrix
D. Peña and V. J. Yohai · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
The implicit function theorem: history, theory, and applications
S. G. Krantz and H. R. Parks · 2002
Earlier work this paper cites.
Robust Statistics: The Approach Based on Influence Functions
F. R. Hampel, E. M. Ronchetti, P. J. Rousseeuw, and W. A. Stahel · 2005
Earlier work this paper cites.
Modern Methods for Robust Regression
R. Andersen · 2007
Earlier work this paper cites.
Concrete Compressive Strength
I.-C. Yeh · 2007
Earlier work this paper cites.
The matrix cookbook
K. B. Petersen, M. S. Pedersen, et al · 2008
Earlier work this paper cites.
Sensitivity analysis in linear regression
S. Chatterjee and A. S. Hadi · 2009
Earlier work this paper cites.
Outlier detection using nonconvex penalized regression
Y. She and A. B. Owen · 2011
Earlier work this paper cites.
Evasion attacks against machine learning at test time
B. Biggio, I. Corona, D. Maiorca, B. Nelson, N. Šrndić, P. Laskov, G. Giacinto, and F. Roli · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Microcredit impacts: Evidence from a randomized microcredit program placement experiment by compartamos banco
M. Angelucci, D. Karlan, and J. Zinman · 2015
Earlier work this paper cites.
The impacts of microfinance: Evidence from joint-liability lending in mongolia
O. Attanasio, B. Augsburg, R. De Haas, E. Fitzsimons, and H. Harmgart · 2015
Cited alongside, same era.
An adaptive, automatic multiple-case deletion technique for detecting influence in regression
S. Roberts, M. A. Martin, and L. Zheng · 2015
Cited alongside, same era.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Cited alongside, same era.
“influence sketching”: Finding influential samples in large-scale regressions
M. Wojnowicz, B. Cruz, X. Zhao, B. Wallace, M. Wolff, J. Luan, and C. Crable · 2016
Cited alongside, same era.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Cited alongside, same era.
If influence functions are the answer, then what is the question?
J. Bae, N. H. Ng, A. Lo, M. Ghassemi, and R. B. Grosse · 2022
Later among the works it cites.
Enfranchisement and incarceration after the 1965 voting rights act
N. Eubank and A. Fresh · 2022
Later among the works it cites.
The adoption of pesticide-free wheat production and farmers’ perceptions of its environmental and health effects
R. Finger and N. Möhring · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data
A. Ilyas, S. M. Park, L. Engstrom, G. Leclerc, and A. Madry · 2022
Later among the works it cites.
How much should we trust the dictator’s gdp growth estimates?
L. R. Martinez · 2022
Later among the works it cites.
Hardness and algorithms for robust and sparse optimization
E. Price, S. Silwal, and S. Zhou · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. F. Agarap · 2018
Cited alongside, same era.
Should we treat data as labor? moving beyond “free”
I. Arrieta-Ibarra, L. Goff, D. Jiménez-Hernández, J. Lanier, and E. G. Weyl · 2018
Cited alongside, same era.
Why is my classifier discriminatory?
I. Chen, F. D. Johansson, and D. Sontag · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
T. George, C. Laurent, X. Bouthillier, N. Ballas, and P. Vincent · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
C.-K. Yeh, J. Kim, I. E.-H. Yen, and P. K. Ravikumar · 2018
Cited alongside, same era.
Machine learning explainability in finance: an application to default risk analysis
P. Bracke, A. Datta, C. Jung, and S. Sen · 2019
Cited alongside, same era.
Regression diagnostics: An introduction
J. Fox · 2019
Cited alongside, same era.
Later among the works it cites.
Fair infinitesimal jackknife: Mitigating the influence of biased training data points without refitting
P. Sattigeri, S. Ghosh, I. Padhi, P. Dognin, and K. R. Varshney · 2022
Later among the works it cites.
Scaling up influence functions
A. Schioppa, P. Zablotskaia, D. Vilar, and A. Sokolov · 2022
Later among the works it cites.
Understanding instance-level impact of fairness constraints
J. Wang, X. E. Wang, and Y. Liu · 2022
Later among the works it cites.
Explainable machine learning for public policy: Use cases, gaps, and research directions
K. Amarasinghe, K. T. Rodolfa, H. Lamba, and R. Ghani · 2023
Later among the works it cites.
Influence diagnostics under self-concordance
J. Fisher, L. Liu, K. Pillutla, Y. Choi, and Z. Harchaoui · 2023
Later among the works it cites.
Towards practical robustness auditing for linear regression
D. Freund and S. B. Hopkins · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
R. Grosse, J. Bae, C. Anil, N. Elhage, A. Tamkin, A. Tajdini, B. Steiner, D. Li, E. Durmus, E. Perez, E. Hubinger, K. Lukošiūtė, K. Nguyen, N. Joseph, S. McCandlish, J. Kaplan, and S. R. Bowman · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
K. Guu, A. Webson, E. Pavlick, L. Dixon, I. Tenney, and T. Bolukbasi · 2023
Later among the works it cites.
Towards a statistical theory of data selection under weak supervision
G. Kolossov, A. Montanari, and P. Tandon · 2023
Later among the works it cites.
Provably auditing ordinary least squares in low dimensions
A. Moitra and D. Rohatgi · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
S. M. Park, K. Georgiev, A. Ilyas, G. Leclerc, and A. Madry · 2023
Later among the works it cites.
Understanding influence functions and datamodels via harmonic analysis
N. Saunshi, A. Gupta, M. Braverman, and S. Arora · 2023
Later among the works it cites.
Farewell to aimless large-scale pretraining: Influential subset selection for language model
X. Wang, W. Zhou, Q. Zhang, J. Zhou, S. Gao, J. Wang, M. Zhang, X. Gao, Y. W. Chen, and T. Gui · 2023
Later among the works it cites.
Training data attribution via approximate unrolled differentation
J. Bae, W. Lin, J. Lorraine, and R. Grosse · 2024
Closest in time.
”what data benefits my classifier?” enhancing model performance and interpretability through influence-based data selection
A. Chhabra, P. Li, P. Mohapatra, and H. Liu · 2024
Closest in time.
Training data influence analysis and estimation: A survey
Z. Hammoudeh and D. Lowd · 2024
Closest in time.
Approximations to worst-case data dropping: unmasking failure modes
J. Y. Huang, D. R. Burt, T. D. Nguyen, Y. Shen, and T. Broderick · 2024
Closest in time.
Rethinking data shapley for data selection tasks: Misleads and merits
J. T. Wang, T. Yang, J. Zou, Y. Kwon, and R. Jia · 2024
Closest in time.
Intriguing properties of data attribution on diffusion models
X. Zheng, T. Pang, C. Du, J. Jiang, and M. Lin · 2024
Closest in time.